Updated

Which AI model writes the most business-logic bugs?

As of Sep 27, 2026, Claude Opus 5 had the highest rate of business-logic bugs, 44% above the all-model rate (1.44), statistically tied with Claude Opus 4.8, Claude Fable 5.1, and GPT-6 Astra. GPT-5.6 Sol had the lowest, 36% below average (0.64), statistically tied with Claude Fable 5, Claude Fable 5.1, and GPT-6 Astra.

Models ranked by relative rate of business-logic bugs, fewest first. Lower values are better.
RankModelRate
1= (statistically tied)GPT-5.6 Sol
0.64
2= (statistically tied)Claude Fable 5
0.78
3= (statistically tied)Claude Opus 4.8
1.07
4= (statistically tied)Claude Fable 5.1
1.22
5= (statistically tied)GPT-6 Astra
1.24
6= (statistically tied)Claude Opus 5
1.44

1.00 = all-model rate · lower is better · = statistically tied