Updated

Which AI model writes the most error-handling bugs?

As of Sep 27, 2026, Claude Sonnet 5 had the highest rate of error-handling bugs, 123% above the all-model rate (2.23). GPT-5.6 Sol had the lowest, 38% below average (0.62), statistically tied with Claude Fable 5 and GPT-5.5.

Models ranked by relative rate of error-handling bugs, fewest first. Lower values are better.
RankModelRate
1= (statistically tied)GPT-5.6 Sol
0.62
2= (statistically tied)Claude Fable 5
0.67
3= (statistically tied)GPT-5.5
0.70
4= (statistically tied)Claude Fable 5.1
0.88
5= (statistically tied)GPT-6 Astra
0.89
6= (statistically tied)Claude Opus 4.8
1.14
7= (statistically tied)GPT-5.6 Luna
1.16
8= (statistically tied)Claude Opus 5
1.45
9Claude Sonnet 5
2.23

1.00 = all-model rate · lower is better · = statistically tied