Claude Sonnet 4.6
Comprehensive benchmark performance across 11 evaluation categories
Composite Score
0.0/100Rank
Token Benchmark
70.3Lower burn, higher score
Total Tokens
~11.0k/test
Category Radar
The full radar chart is shown on wider screens. On mobile, the category breakdown below provides the same values in a readable stacked layout.
Historical Composite Score
Category Breakdown
| # | Category | Score | Tests | 7-Day Trend | Weight |
|---|---|---|---|---|---|
| 1 | Coding Tasks | 100.0 | 3 | 1.0x | |
| 2 | Instruction Following | 100.0 | 3 | 1.0x | |
| 3 | Feature Implementation | 97.7 | 3 | 1.0x | |
| 4 | Bug Introduction Rate | 97.0 | 3 | 1.0x | |
| 5 | Code Quality | 96.7 | 3 | 1.0x | |
| 6 | Bug Fixes | 96.5 | 3 | 1.0x | |
| 7 | Security Awareness | 96.3 | 3 | 1.0x | |
| 8 | Code Thoroughness | 92.3 | 3 | 1.0x | |
| 9 | Token Efficiency | 70.3 | 29 | 1.0x | |
| 10 | Long Reasoning | 68.6 | 3 | 1.0x | |
| 11 | Performance & Efficiency | 33.3 | 3 | 1.0x |
Individual Test Results
Token Efficiency is computed from every successful task in the run. The model with the lowest average token burn receives 100, and heavier token usage is penalized proportionally.
Avg Tokens/Test
11.0k
Total Tokens
318.7k
Regression History
9Score dropped -7.9% from 95.6 to 88.1
Score dropped -3.0% from 95.2 to 92.3
Score dropped -9.7% from 88.4 to 79.8
Score dropped -17.8% from 74.9 to 61.6
Score dropped -55.7% from 88.7 to 39.3
Score dropped -18.7% from 81.5 to 66.3
Score dropped -14.8% from 67.9 to 57.9
Score dropped -24.8% from 98.8 to 74.3
Score dropped -4.7% from 100.0 to 95.3