GPT-5.5
Comprehensive benchmark performance across 11 evaluation categories
Composite Score
0.0/100Rank
Token Benchmark
51.7Lower burn, higher score
Total Tokens
~15.0k/test
Category Radar
The full radar chart is shown on wider screens. On mobile, the category breakdown below provides the same values in a readable stacked layout.
Historical Composite Score
Category Breakdown
| # | Category | Score | Tests | 7-Day Trend | Weight |
|---|---|---|---|---|---|
| 1 | Instruction Following | 100.0 | 3 | 1.0x | |
| 2 | Coding Tasks | 98.3 | 3 | 1.0x | |
| 3 | Feature Implementation | 98.3 | 3 | 1.0x | |
| 4 | Bug Introduction Rate | 96.0 | 3 | 1.0x | |
| 5 | Security Awareness | 95.3 | 3 | 1.0x | |
| 6 | Code Quality | 94.3 | 3 | 1.0x | |
| 7 | Bug Fixes | 93.3 | 3 | 1.0x | |
| 8 | Code Thoroughness | 92.7 | 3 | 1.0x | |
| 9 | Performance & Efficiency | 92.3 | 3 | 1.0x | |
| 10 | Long Reasoning | 69.6 | 3 | 1.0x | |
| 11 | Token Efficiency | 51.7 | 30 | 1.0x |
Individual Test Results
Token Efficiency is computed from every successful task in the run. The model with the lowest average token burn receives 100, and heavier token usage is penalized proportionally.
Avg Tokens/Test
15.0k
Total Tokens
448.9k
Regression History
7Score dropped -3.1% from 90.9 to 88.1
Score dropped -4.2% from 92.6 to 88.7
Score dropped -5.1% from 92.4 to 87.7
Score dropped -3.4% from 96.6 to 93.3
Score dropped -9.0% from 53.8 to 49.0
Score dropped -4.3% from 92.3 to 88.3
Score dropped -5.7% from 65.0 to 61.3