Updated
Which AI model writes the most business-logic bugs?
As of Sep 27, 2026, Claude Opus 5 had the highest rate of business-logic bugs, 44% above the all-model rate (1.44), statistically tied with Claude Opus 4.8, Claude Fable 5.1, and GPT-6 Astra. GPT-5.6 Sol had the lowest, 36% below average (0.64), statistically tied with Claude Fable 5, Claude Fable 5.1, and GPT-6 Astra.
| Rank | Model | Rate |
|---|---|---|
| 1= (statistically tied) | GPT-5.6 Sol | 0.64 |
| 2= (statistically tied) | Claude Fable 5 | 0.78 |
| 3= (statistically tied) | Claude Opus 4.8 | 1.07 |
| 4= (statistically tied) | Claude Fable 5.1 | 1.22 |
| 5= (statistically tied) | GPT-6 Astra | 1.24 |
| 6= (statistically tied) | Claude Opus 5 | 1.44 |
1.00 = all-model rate · lower is better · = statistically tied