Back to Dashboard
CategoryWeight: 1.0x
Code Thoroughness
Evaluates completeness of generated code: edge case handling, input validation, error paths, and test coverage.
Best Score
0.0Avg Score
0.0Tests
3Performance Over Time — All Models
Model Rankings
Test Breakdown
Edge Case Coverage
Generate code handling null, empty, unicode, and overflow inputs
GPT-5.5
92.7Claude Sonnet 4.6
92.3Grok 4.5
91.8Claude Opus 4.8
89.9Error Path Completeness
Ensure all failure modes have proper error handling and logging
GPT-5.5
92.7Claude Sonnet 4.6
92.3Grok 4.5
91.8Claude Opus 4.8
89.9Test Suite Completeness
Generate tests covering happy path, edge cases, and integration
GPT-5.5
92.7Claude Sonnet 4.6
92.3Grok 4.5
91.8Claude Opus 4.8
89.9