Sitelet
https://cursor.com/cursorbench
CursorBench 4.0 We evaluate agents on ambiguous, multi-file tasks from real Cursor sessions. Higher scores are better.
More about CursorBench↗ A scatter and line chart comparing Fable 5.1, Opus 5.5, Opus 5, Grok 4.7, Grok 4.6, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, Sonnet 5.5, Sonnet 5, Gemini 3.8 Flash, Muse Spark 1.3, GLM 5.3, GLM 5.3 Flash, and Composer 2.5 scores against average cost per task. 60% CursorBench 4.0 score 55% 50% 45% 40% 35% 30% 25% 20% $18 $15 $12 $9 $6 $3 $0 Average cost per task Opus 5.5 Fable 5.1 Sonnet 5.5 Gemini 3.8 Flash GPT-5.6 Sol Grok 4.7 Cost Tokens Steps
Model Score Cost Cost / task Tokens Tokens / task Steps Steps / task 1 Opus 5.5 Max 57.8% $13.43 218,363 185 2 Opus 5.5 Extra High 56.0% $6.98 101,083 109 3 Opus 5.5 High 56.0% $3.97 53,078 68 4 Sonnet 5.5 Max 55.5% $9.67 271,920 170 5 Sonnet 5.5 Extra High 53.1% $3.88 100,158 78 6 Opus 5.5 Medium 52.5% $2.91 37,954 54 7 Fable 5.1 Max 51.8% $17.28 117,236 128 8 Fable 5.1 Extra High 51.6% $13.01 87,294 101 9 Fable 5.1 High 49.2% $9.08 58,438 77 10 Sonnet 5.5 High 47.8% $1.67 37,391 41 11 Fable 5.1 Medium 46.8% $7.05 45,411 63 12 Opus 5 Max 46.6% $11.95 85,384 106 13 Grok 4.7 Extra High 46.3% $6.01 70,141 88 14 Opus 5 Extra High 46.1% $11.43 80,094 103 15 Fable 5.1 Low 45.1% $5.44 34,795 51 16 Opus 5 High 44.7% $9.00 61,405 86 17 Grok 4.7 High 43.9% $4.69 56,382 71 18 Opus 5.5 Low 43.7% $1.17 15,811 28 19 Opus 5 Medium 43.3% $6.94 45,272 72 20 GLM 5.3 Max 42.6% $5.05 96,387 166 21 GPT-5.6 Sol Max 41.7% $8.23 42,944 99 22 Grok 4.7 Medium 41.6% $3.49 36,683 60 23 Muse Spark 1.3 Max 41.6% $2.64 52,005 98 24 Grok 4.6 Extra High 41.4% $6.10 49,814 56 25 GPT-5.6 Terra Max 41.3% $5.14 60,814 107 26 Opus 5 Low 40.7% $4.87 31,995 57 27 Grok 4.6 High 40.4% $5.20 41,387 48 28 Gemini 3.8 Flash High 39.6% $4.70 162,565 324 29 Sonnet 5.5 Medium 39.2% $0.70 16,036 22 30 GLM 5.3 High 38.0% $3.24 60,031 114 31 GPT-5.6 Sol Extra High 37.7% $4.40 24,729 55 32 Muse Spark 1.3 Extra High 37.5% $2.10 40,891 83 33 Gemini 3.8 Flash Medium 37.3% $4.06 128,364 290 34 GLM 5.3 Flash Max 36.8% $0.39 56,410 118 35 Grok 4.6 Medium 36.1% $3.48 24,893 40 36 GPT-5.6 Luna Max 35.9% $1.03 87,284 208 37 Sonnet 5.5 Low 35.8% $0.50 11,668 18 38 GPT-5.6 Sol High 35.7% $2.85 16,174 41 39 Sonnet 5 Max 34.1% $7.17 149,257 140 40 GPT-5.6 Terra Extra High 33.6% $1.81 23,436 43 41 Grok 4.6 Low 33.4% $2.25 16,307 32 42 Muse Spark 1.3 High 33.4% $1.66 30,654 69 43 GLM 5.3 Low 33.3% $2.04 31,983 81 44 Grok 4.7 Low 33.1% $1.58 15,677 40 45 GPT-5.6 Luna Extra High 33.0% $0.44 40,598 98 46 Muse Spark 1.3 Medium 32.6% $1.49 27,255 64 47 Sonnet 5 Extra High 32.0% $4.55 83,373 102 48 GPT-5.6 Sol Medium 31.1% $1.77 10,111 32 49 GLM 5.3 Flash High 31.1% $0.25 35,104 84 50 Sonnet 5 High 30.8% $3.48 61,146 85 51 GPT-5.6 Terra High 30.7% $1.11 13,162 33 52 GPT-5.6 Luna High 29.4% $0.25 23,368 64 53 Muse Spark 1.3 Low 29.3% $0.93 17,483 47 54 Sonnet 5 Medium 28.0% $2.31 39,114 65 55 Composer 2.5 27.7% $0.68 17,347 41 56 GPT-5.6 Terra Medium 27.6% $0.64 7,307 25 57 GLM 5.3 Flash Low 26.9% $0.15 17,831 58 58 GPT-5.6 Terra Low 25.2% $0.52 5,914 23 59 GPT-5.6 Sol Low 24.6% $0.87 4,885 21 60 Muse Spark 1.3 Minimal 24.3% $0.56 10,620 34 61 Sonnet 5 Low 24.1% $1.39 23,772 46 62 GPT-5.6 Luna Medium 22.2% $0.08 7,642 32 63 GPT-5.6 Luna Low 16.0% $0.03 3,288 18
Changelog Sep 10, 2026 Tasks CursorBench 4.0Introduced new long-horizon problems focused on edit, refactor, investigation, intent understanding, managing jobs, and design adherence. Aug 11, 2026 Reporting Updated Sonnet 5 results to account for adjusted pricing. Jul 30, 2026 Reporting Updated GPT-5.6 Terra and Luna results to account for adjusted pricing. Jul 9, 2026 Reporting Updated GPT-5.6 Sol, Terra, and Luna results to account for cache write costs. Jul 8, 2026 Tasks CursorBench 3.2Introduced instruction following and advanced tool use problems. May 19, 2026 Tasks CursorBench 3.1Introduced problems focused on codebase understanding, bugfinding, planning, and code review. Improved grading criteria for some edit tasks. Mar 11, 2026 Tasks CursorBench 3.0Initial set of tasks focused on edit, refactor, and bugfix problems. Avg cost / task is computed by applying each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task. Results are subject to variance; small differences in scores may not be statistically meaningful.
CursorBench 4.0 We evaluate agents on ambiguous, multi-file tasks from real Cursor sessions. Higher scores are better.
More about CursorBench↗ A scatter and line chart comparing Fable 5.1, Opus 5.5, Opus 5, Grok 4.7, Grok 4.6, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, Sonnet 5.5, Sonnet 5, Gemini 3.8 Flash, Muse Spark 1.3, GLM 5.3, GLM 5.3 Flash, and Composer 2.5 scores against average cost per task. 60% CursorBench 4.0 score 55% 50% 45% 40% 35% 30% 25% 20% $18 $15 $12 $9 $6 $3 $0 Average cost per task Opus 5.5 Fable 5.1 Sonnet 5.5 Gemini 3.8 Flash GPT-5.6 Sol Grok 4.7 Cost Tokens Steps
Model Score Cost Cost / task Tokens Tokens / task Steps Steps / task 1 Opus 5.5 Max 57.8% $13.43 218,363 185 2 Opus 5.5 Extra High 56.0% $6.98 101,083 109 3 Opus 5.5 High 56.0% $3.97 53,078 68 4 Sonnet 5.5 Max 55.5% $9.67 271,920 170 5 Sonnet 5.5 Extra High 53.1% $3.88 100,158 78 6 Opus 5.5 Medium 52.5% $2.91 37,954 54 7 Fable 5.1 Max 51.8% $17.28 117,236 128 8 Fable 5.1 Extra High 51.6% $13.01 87,294 101 9 Fable 5.1 High 49.2% $9.08 58,438 77 10 Sonnet 5.5 High 47.8% $1.67 37,391 41 11 Fable 5.1 Medium 46.8% $7.05 45,411 63 12 Opus 5 Max 46.6% $11.95 85,384 106 13 Grok 4.7 Extra High 46.3% $6.01 70,141 88 14 Opus 5 Extra High 46.1% $11.43 80,094 103 15 Fable 5.1 Low 45.1% $5.44 34,795 51 16 Opus 5 High 44.7% $9.00 61,405 86 17 Grok 4.7 High 43.9% $4.69 56,382 71 18 Opus 5.5 Low 43.7% $1.17 15,811 28 19 Opus 5 Medium 43.3% $6.94 45,272 72 20 GLM 5.3 Max 42.6% $5.05 96,387 166 21 GPT-5.6 Sol Max 41.7% $8.23 42,944 99 22 Grok 4.7 Medium 41.6% $3.49 36,683 60 23 Muse Spark 1.3 Max 41.6% $2.64 52,005 98 24 Grok 4.6 Extra High 41.4% $6.10 49,814 56 25 GPT-5.6 Terra Max 41.3% $5.14 60,814 107 26 Opus 5 Low 40.7% $4.87 31,995 57 27 Grok 4.6 High 40.4% $5.20 41,387 48 28 Gemini 3.8 Flash High 39.6% $4.70 162,565 324 29 Sonnet 5.5 Medium 39.2% $0.70 16,036 22 30 GLM 5.3 High 38.0% $3.24 60,031 114 31 GPT-5.6 Sol Extra High 37.7% $4.40 24,729 55 32 Muse Spark 1.3 Extra High 37.5% $2.10 40,891 83 33 Gemini 3.8 Flash Medium 37.3% $4.06 128,364 290 34 GLM 5.3 Flash Max 36.8% $0.39 56,410 118 35 Grok 4.6 Medium 36.1% $3.48 24,893 40 36 GPT-5.6 Luna Max 35.9% $1.03 87,284 208 37 Sonnet 5.5 Low 35.8% $0.50 11,668 18 38 GPT-5.6 Sol High 35.7% $2.85 16,174 41 39 Sonnet 5 Max 34.1% $7.17 149,257 140 40 GPT-5.6 Terra Extra High 33.6% $1.81 23,436 43 41 Grok 4.6 Low 33.4% $2.25 16,307 32 42 Muse Spark 1.3 High 33.4% $1.66 30,654 69 43 GLM 5.3 Low 33.3% $2.04 31,983 81 44 Grok 4.7 Low 33.1% $1.58 15,677 40 45 GPT-5.6 Luna Extra High 33.0% $0.44 40,598 98 46 Muse Spark 1.3 Medium 32.6% $1.49 27,255 64 47 Sonnet 5 Extra High 32.0% $4.55 83,373 102 48 GPT-5.6 Sol Medium 31.1% $1.77 10,111 32 49 GLM 5.3 Flash High 31.1% $0.25 35,104 84 50 Sonnet 5 High 30.8% $3.48 61,146 85 51 GPT-5.6 Terra High 30.7% $1.11 13,162 33 52 GPT-5.6 Luna High 29.4% $0.25 23,368 64 53 Muse Spark 1.3 Low 29.3% $0.93 17,483 47 54 Sonnet 5 Medium 28.0% $2.31 39,114 65 55 Composer 2.5 27.7% $0.68 17,347 41 56 GPT-5.6 Terra Medium 27.6% $0.64 7,307 25 57 GLM 5.3 Flash Low 26.9% $0.15 17,831 58 58 GPT-5.6 Terra Low 25.2% $0.52 5,914 23 59 GPT-5.6 Sol Low 24.6% $0.87 4,885 21 60 Muse Spark 1.3 Minimal 24.3% $0.56 10,620 34 61 Sonnet 5 Low 24.1% $1.39 23,772 46 62 GPT-5.6 Luna Medium 22.2% $0.08 7,642 32 63 GPT-5.6 Luna Low 16.0% $0.03 3,288 18
Changelog Sep 10, 2026 Tasks CursorBench 4.0Introduced new long-horizon problems focused on edit, refactor, investigation, intent understanding, managing jobs, and design adherence. Aug 11, 2026 Reporting Updated Sonnet 5 results to account for adjusted pricing. Jul 30, 2026 Reporting Updated GPT-5.6 Terra and Luna results to account for adjusted pricing. Jul 9, 2026 Reporting Updated GPT-5.6 Sol, Terra, and Luna results to account for cache write costs. Jul 8, 2026 Tasks CursorBench 3.2Introduced instruction following and advanced tool use problems. May 19, 2026 Tasks CursorBench 3.1Introduced problems focused on codebase understanding, bugfinding, planning, and code review. Improved grading criteria for some edit tasks. Mar 11, 2026 Tasks CursorBench 3.0Initial set of tasks focused on edit, refactor, and bugfix problems. Avg cost / task is computed by applying each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task. Results are subject to variance; small differences in scores may not be statistically meaningful.