Updated

Which AI model writes the most testing and documentation issues?

As of Sep 27, 2026, Claude Sonnet 5 had the highest rate of testing and documentation issues, 147% above the all-model rate (2.47). GPT-5.5 had the lowest, 81% below average (0.19).

Models ranked by relative rate of testing and documentation issues, fewest first. Lower values are better.
RankModelRate
1GPT-5.5
0.19
2= (statistically tied)Claude Fable 5
0.58
3= (statistically tied)Claude Opus 4.8
0.61
4GPT-5.6 Sol
0.80
5GPT-5.6 Luna
1.12
6= (statistically tied)Claude Fable 5.1
1.59
7= (statistically tied)Claude Opus 5
1.79
8= (statistically tied)GPT-6 Astra
1.80
9Claude Sonnet 5
2.47

1.00 = all-model rate · lower is better · = statistically tied