Updated
Which AI model writes the most error-handling bugs?
As of Sep 27, 2026, Claude Sonnet 5 had the highest rate of error-handling bugs, 123% above the all-model rate (2.23). GPT-5.6 Sol had the lowest, 38% below average (0.62), statistically tied with Claude Fable 5 and GPT-5.5.
| Rank | Model | Rate |
|---|---|---|
| 1= (statistically tied) | GPT-5.6 Sol | 0.62 |
| 2= (statistically tied) | Claude Fable 5 | 0.67 |
| 3= (statistically tied) | GPT-5.5 | 0.70 |
| 4= (statistically tied) | Claude Fable 5.1 | 0.88 |
| 5= (statistically tied) | GPT-6 Astra | 0.89 |
| 6= (statistically tied) | Claude Opus 4.8 | 1.14 |
| 7= (statistically tied) | GPT-5.6 Luna | 1.16 |
| 8= (statistically tied) | Claude Opus 5 | 1.45 |
| 9 | Claude Sonnet 5 | 2.23 |
1.00 = all-model rate · lower is better · = statistically tied