# cubic Bug Bench: which AI coding model writes the fewest bugs?

> Bug Bench ranks AI coding models by relative bug rate: counted bugs per AI-attributed line in code reviewed by cubic, compared with the rate across all AI-attributed lines. 1.00 is average; lower is better.

Updated Sep 26, 2026 · Ranking window: Jun 1 to Sep 20, 2026 · Bug Bench v0.1

As of Sep 26, 2026, GPT-5.5 had the lowest relative bug rate on Bug Bench (0.47, 53% fewer counted bugs than average), among 9 ranked AI coding models.

## Ranking

| Rank | Model | Relative bug rate | Reading | Statistically tied with |
| ---: | :--- | ---: | :--- | :--- |
| 1 | [GPT-5.5](https://www.cubic.dev/bugbench/models/gpt-5-5) | 0.47 | 53% fewer counted bugs than average | — |
| 2 | [Claude Fable 5](https://www.cubic.dev/bugbench/models/claude-fable-5) | 0.70 | 30% fewer counted bugs than average | GPT-5.6 Sol |
| 3 | [GPT-5.6 Sol](https://www.cubic.dev/bugbench/models/gpt-5-6-sol) | 0.73 | 27% fewer counted bugs than average | Claude Fable 5 |
| 4 | [Claude Opus 4.8](https://www.cubic.dev/bugbench/models/claude-opus-4-8) | 0.88 | 12% fewer counted bugs than average | — |
| 5 | [GPT-6 Astra](https://www.cubic.dev/bugbench/models/gpt-6-astra) | 1.16 | 16% more counted bugs than average | Claude Fable 5.1 and GPT-5.6 Luna |
| 6 | [Claude Fable 5.1](https://www.cubic.dev/bugbench/models/claude-fable-5-1) | 1.17 | 17% more counted bugs than average | GPT-6 Astra and GPT-5.6 Luna |
| 7 | [GPT-5.6 Luna](https://www.cubic.dev/bugbench/models/gpt-5-6-luna) | 1.28 | 28% more counted bugs than average | GPT-6 Astra and Claude Fable 5.1 |
| 8 | [Claude Opus 5](https://www.cubic.dev/bugbench/models/claude-opus-5) | 1.58 | 58% more counted bugs than average | — |
| 9 | [Claude Sonnet 5](https://www.cubic.dev/bugbench/models/claude-sonnet-5) | 2.10 | 110% more counted bugs than average | — |

## Not yet ranked

- [Claude Opus 5.5](https://www.cubic.dev/bugbench/models/claude-opus-5-5) (launch week): first seen the week of Sep 21, 2026; it will be ranked once enough of its code has been reviewed across many organizations.
- [GPT-6 Sol](https://www.cubic.dev/bugbench/models/gpt-6-sol) (launch week): first seen the week of Sep 21, 2026; it will be ranked once enough of its code has been reviewed across many organizations.

## Best model by coding agent

Coding-agent rates compare each model inside one harness against the overall average. [All coding agents](https://www.cubic.dev/bugbench/harnesses).

| Coding agent | Best model | Relative bug rate | Statistically tied with | Models ranked |
| :--- | :--- | ---: | :--- | ---: |
| [Claude Code](https://www.cubic.dev/bugbench/harnesses/claude-code) | [Claude Fable 5](https://www.cubic.dev/bugbench/models/claude-fable-5) | 0.72 | — | 5 |
| [Codex](https://www.cubic.dev/bugbench/harnesses/codex) | [GPT-5.5](https://www.cubic.dev/bugbench/models/gpt-5-5) | 0.42 | — | 4 |
| [Cursor](https://www.cubic.dev/bugbench/harnesses/cursor) | [Claude Fable 5](https://www.cubic.dev/bugbench/models/claude-fable-5) | 0.17 | — | 5 |

## Bug rate vs list price

List prices are USD per 1M tokens from each model's own provider. [Bug rate vs price](https://www.cubic.dev/bugbench/price).

| Rank | Model | Relative bug rate | Input price | Output price |
| ---: | :--- | ---: | ---: | ---: |
| 1 | [GPT-5.5](https://www.cubic.dev/bugbench/models/gpt-5-5) | 0.47 | $5 | $30 |
| 2 | [Claude Fable 5](https://www.cubic.dev/bugbench/models/claude-fable-5) | 0.70 | $10 | $50 |
| 3 | [GPT-5.6 Sol](https://www.cubic.dev/bugbench/models/gpt-5-6-sol) | 0.73 | $4 | $20 |
| 4 | [Claude Opus 4.8](https://www.cubic.dev/bugbench/models/claude-opus-4-8) | 0.88 | $5 | $25 |
| 5 | [GPT-6 Astra](https://www.cubic.dev/bugbench/models/gpt-6-astra) | 1.16 | $10 | $50 |
| 6 | [Claude Fable 5.1](https://www.cubic.dev/bugbench/models/claude-fable-5-1) | 1.17 | $10 | $50 |
| 7 | [GPT-5.6 Luna](https://www.cubic.dev/bugbench/models/gpt-5-6-luna) | 1.28 | $0.20 | $1.20 |
| 8 | [Claude Opus 5](https://www.cubic.dev/bugbench/models/claude-opus-5) | 1.58 | $5 | $25 |
| 9 | [Claude Sonnet 5](https://www.cubic.dev/bugbench/models/claude-sonnet-5) | 2.10 | $2 | $10 |

## Methodology

- A counted bug is a cubic code review finding on AI-written code that a developer upvoted or an AI coding agent addressed.
- Relative bug rate is counted bugs per AI-attributed line divided by the rate across all AI-attributed lines in the window, so 1.00 is average and lower is better. The overall rate includes unranked and unspecified models, so it is not an average of the models shown.
- "Statistically tied" means the two models' 95% confidence intervals overlap, so neither is clearly ahead.
- A model is ranked once enough of its code has been reviewed across many organizations. Launch-week models are listed but not ranked.
- Bug Bench is observational, not causal: teams choose their models, so a rate reflects the code each model was used for as well as the model.
- Full methodology: https://www.cubic.dev/bugbench#methodology

## Pages

- [Leaderboard](https://www.cubic.dev/bugbench)
- [All models](https://www.cubic.dev/bugbench/models)
- [Compare two models](https://www.cubic.dev/bugbench/compare)
- [Best model by language](https://www.cubic.dev/bugbench/languages)
- [Bugs by type](https://www.cubic.dev/bugbench/bugs)
- [Coding agents](https://www.cubic.dev/bugbench/harnesses)
- [Bug rate vs price](https://www.cubic.dev/bugbench/price)
- [Trends](https://www.cubic.dev/bugbench/trends)
- [New model launches](https://www.cubic.dev/bugbench/launches)
- [Monthly editions](https://www.cubic.dev/bugbench/editions)

## Data

- [JSON](https://www.cubic.dev/bugbench/data.json): every published model, coding-agent, language, and bug-type rate.
- [CSV](https://www.cubic.dev/bugbench/data.csv): the same rankings in long format.
- [Atom feed](https://www.cubic.dev/bugbench/feed.xml): ranking changes and monthly editions.
- SVG badge: https://www.cubic.dev/bugbench/badge/<model-slug>

## License

CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) — cite as: cubic Bug Bench, Sep 26, 2026, https://www.cubic.dev/bugbench
