Four AI Coding Agents, One Bug Fix: Where They Actually Differ

Last week I ran a small, deliberately boring experiment. I gave four AI coding agents the same task on the same real codebase, with the same written brief and the same time limit. The agents were Claude Opus 5.5, GPT 5.6 Sol, GPT 5.6 Terra and GPT 6 Astra, each set to its highest reasoning level.
The codebase is an ordinary business web application, the kind a lot of teams run every day. It happens to be written in PHP, but the review questions apply in other languages too. Whether the results would be the same is a separate experiment.
I reviewed the four results myself, the way I would review work from four contractors I had never met. I did not run anything. I read what each agent changed, the checks it wrote to prove the change worked, and the supporting files those checks depend on. Then I scored each delivery against a scorecard I wrote before anyone started.
The outcome surprised me, though not in the direction the marketing suggests. All four agents addressed the core bug in the code I reviewed. What separated the best delivery from the worst was not the main fix. It was the safety net around it, and how accurately each agent described what it had actually verified.
Which models, and why the distinction matters
There are four models in this comparison. Their providers position them differently:
| Model | Public positioning |
|---|---|
| Claude Opus 5.5 | Anthropic emphasizes sustained coding work, including large migrations and audits, with improved token efficiency over Opus 5. |
| GPT 5.6 Sol | The flagship tier of the GPT-5.6 family, aimed at complex professional work. |
| GPT 5.6 Terra | A lower-cost tier balancing capability and price; OpenAI compares its place in the family to earlier mini models. |
| GPT 6 Astra | OpenAI’s most capable tier for demanding work, including complex reasoning and coding. |
These are provider descriptions, checked on September 29, 2026, not conclusions from this experiment. A model is also only one part of a coding agent: the surrounding application supplies instructions, repository access, tools and context management. Setting each model to its highest available reasoning level does not give them equal compute budgets. Read the results as a comparison of the four deliveries under this brief, not an isolated measurement of model intelligence.
The task
The bug is a very common one. In this application, many kinds of records (offices, properties, users and more) can have files attached, such as a logo or a photo. Every time the app needed one file, for example a thumbnail, it went to the database and asked for it. Every single time, even when the data had already been loaded.
For one office, that is fine. For a page listing fifty properties, it means fifty extra trips to the database just to show fifty thumbnails. Think of a shop assistant who fetches each item from the warehouse one at a time, when the whole order was already sitting on the counter. Developers call this the “N+1 problem”, and it slows applications down quietly until someone notices.
Someone had already patched one part of the app by copying the logic and adding a workaround. So the real job was to fix it once, in the shared place, and remove the copy.
The written brief fit on two pages. The rules that mattered:
- If the data is already loaded, use only that data. Never go back to the database to “complete” it, and never return something the caller deliberately left out.
- Keep the existing order. The item marked as the default wins; otherwise the first matching item wins.
- If nothing was loaded in advance, ask the database for exactly one item, not for everything and then pick one.
- Fix a second, smaller defect: one function could crash when a file did not exist. It should return “nothing” instead, while the property pages keep their placeholder image. Private files must stay private.
- Write automated checks that count the actual database requests, cover both the “already loaded” and “not loaded” cases, use several records at once, and include at least one check that fails on the old code and passes on the new.
- Three hours. Do not run anything. Say clearly what you could not verify.
That last rule matters. This was a test of reading and reasoning, not of tooling. Nobody was allowed to run the code, including me.
How I judged
Before the first agent started, I fixed ten scenarios to check in every delivery, from the simplest case to many records at once. I also fixed the weights: correctness 40, database efficiency 20, tests 20, cleanliness and scope 15, the agent’s report 5.
Two rules kept me honest. An agent’s own report counts as evidence that it understood the task, not as proof that the work is right. And I looked at how long each agent took only after the technical scores were final.
One disclosure: the review was not blind. I knew which agent produced which delivery while I read it. I tried to keep that out of the scores. The findings below explain the defects behind the deductions, but the patches and full review records are not included here, so readers cannot independently reproduce the scoring from this article alone.
How this relates to public benchmarks
SWE-bench Verified uses 500 human-filtered repository issues, with patches assessed by executing tests. SWE-Bench Pro extends repository-level evaluation to longer, more complex tasks, often involving multiple files. Both ask whether an agent can resolve an issue across a collection of problems.
Current evaluations also look beyond a single bug fix. Terminal-Bench 4.0 evaluates agents working in terminal environments and explicitly controls resources and task timeouts. Anthropic’s Opus 5.5 evaluation also reports FrontierCode, which assesses whether changes would be merged, and CursorBench, which uses ambiguous tasks drawn from real coding sessions. Quality of a proposed change is already part of the evaluation landscape.
Our exercise adds a close reading of the tests an agent writes and the claims it makes about its own work. It cannot provide the runtime evidence or breadth of those benchmarks. A 100 here means no deduction under this review’s rubric; it is not a 100% task success rate. Benchmark versions, agent tools, reasoning settings and repeat runs all matter when comparing published scores.
The results
Static review scores out of 100. No code or tests were executed; test failures below were inferred from the source.
| Agent | Review score | One-line reason |
|---|---|---|
| Claude Opus 5.5 | 100 | Core fix looks correct; no test setup defect found; report matches the source |
| GPT 6 Astra | 100 | Core fix looks correct; the most thorough tests of the four |
| GPT 5.6 Sol | 94 | Core fix looks correct; one small edge-case slip and one test that could not work |
| GPT 5.6 Terra | 84 | Core fix looks correct; seven of eight tests would fail during setup |
Two perfect scores and a tie is not what I expected. I expected a spread in the quality of the fix. There was almost none.
Finding 1: the fixes converged
Four results, nearly the same shape. One shared piece of logic. Use what is already loaded. Otherwise ask for one item. The copied workaround shrank from about eighty lines to one.
The differences were matters of taste, not correctness. The one real slip came from Sol: for an empty label, it could return a file from the wrong category instead of nothing. No part of the app does that today. Two points off and a note in the review.
On this task, evaluating only whether the agents could write the main fix would have missed most of what told them apart.
Finding 2: the safety nets are where they diverge
Beyond Sol’s small edge-case slip, the defects that separated the deliveries lived in the tests, and each could be found by reading. Three patterns.
Test data that cannot be created. Terra built its test records sensibly, but skipped one step the real application performs automatically: filling in a required field. In this app that field is normally set behind the scenes when a record is saved. The tests turned that behaviour off, so the database would refuse the records before the first check could run. Seven of eight tests would never reach the behaviour they were meant to test. The repair is a single line. The lesson: check test data against the actual database design, not against what usually works.
Relying on how the tool “usually” behaves. One test needed a web address to exist. Sol and Terra registered it on the fly, which looks right and matches what most developers would expect. But the framework only refreshes its list of addresses at startup, so the new one was invisible and the test would have failed. Opus and Astra both knew that and refreshed the list explicitly. The documentation does not say this. The framework’s own source code does.
Reproducing what production really does. The real listing page asks for “one file per property” in a way that looks, on paper, like it would return one file for the whole page. The framework quietly handles it per property. Opus copied that exact real-world setup into a test with two properties, read the framework’s code to confirm the behaviour, and said so in its report. That is the difference between a test that documents production and a test that merely happens to pass.
The two strong test suites shared habits. They used real records rather than fakes. They counted database requests only around the step being measured, not during setup. They included traps: deleted files, files belonging to someone else, even a file with a matching ID but the wrong type of owner. And each key test carried a comment explaining why it would have failed on the old code.
Finding 3: reports are claims, not results
The best report came from Opus. It was the one where I could verify every sentence. It listed two existing problems nobody had asked about, both real: a stale backup copy of a file still in version control, and two pages requesting a thumbnail size that does not exist. It had checked every place affected by the change as a side effect. And it said plainly that nothing had been run.
The reports from Sol and Terra described one test as covering something it could not reach, because the test would have failed earlier. That does not establish dishonesty. The claims were unverified, and they were written in the same confident tone as the claims supported by the source. That is the part to remember.
Length was not the signal. Opus wrote the longest report and Astra a much shorter one; both checked out completely. What mattered was whether I could verify it.
Finding 4: the reviewer is fallible too
Halfway through, one of my own searches came back empty. It looked like there was nothing to find. In fact my search pattern contained a special character that made it match nothing. I noticed only when the same search returned an impossible empty result somewhere else.
Fixing it did not change any verdict. It could have. If you are going to grade an agent on whether it verified its claims, grade yourself by the same standard.
Finding 5: speed bought nothing
Self-reported times: Opus about eleven minutes, Sol nineteen, Terra about ten, and Astra did not report one. The fastest delivery, Terra, had the weakest tests. The slowest, Sol, landed in the middle. Time is a tiebreaker between equally correct results, and a time an agent reports about itself is not a measurement.
Finding 6: the same review score at different prices
Cost matters to whoever signs the invoice. Here, it helps to separate the price of tokens from the cost of a completed task.
Claude Opus 5.5 and GPT 6 Astra both scored 100. Both addressed the core bug, wrote strong tests, and delivered reports I could check against the source. Their published base API prices differ by a factor of 2.5: Opus 5.5 lists $4 per million input tokens and $20 per million output tokens; Astra lists $10 and $50 respectively. Those are standard, uncached rates checked on September 29, 2026, excluding fast mode, long-context surcharges and other pricing adjustments. Sources: Anthropic and OpenAI.
Astra’s test suite was the most thorough of the four, which is a real strength, but it did not translate into a higher review score. The price ratio alone does not establish that its run cost 2.5 times as much. That requires the actual input, cached and output token counts, any billed reasoning tokens, and the applicable billing mode. Those usage records are not included here.
That does not make the pricier model a bad choice in general. A harder task might use what this one did not. It does mean that price is not a proxy for quality, and that you should measure both on a task like yours before you commit to either. A more expensive model that produces the same result as a cheaper one is a cost, not an upgrade.
A checklist if you want to run this yourself
- Write the brief and the scenarios before the first agent starts. Fix the weights and do not move them.
- Pin the starting version of the code. Review new files too, since they do not show up in a simple comparison of changes.
- Record the exact model version, agent application, reasoning setting, instructions and tool permissions. For a broader comparison, repeat runs and vary the tasks.
- For every piece of test data, check it against the real database design. For every claim about how a tool behaves, read its source. Remembering how it “usually” works is not evidence.
- Count database requests on a real connection, only around the step under test.
- Treat every sentence in an agent’s report as a hypothesis. Verify it or mark it unverified.
- Score correctness first, then efficiency, then tests, then cleanliness, then the report. Never lines changed. Never report length.
- Look at time and cost only after the technical scores are final, then compare what each one paid for the same result.
- Say whether the review was blind. Mine was not.
What reading cannot tell you
Whether the tests pass. I ranked four deliveries by reading, and the ranking was clear enough to decide which two test suites to run first. That is useful, and it is limited. Nothing above proves runtime behaviour. This was one task, one run per agent, on one codebase. It is not a leaderboard, and it says nothing about how these agents would do on a different task.
What it does say is simple. On this task, the main fixes were roughly interchangeable. The safety nets around them, and the accuracy of the reports, were not. If you accept an agent’s work on the strength of its own summary, you are accepting a claim. Read what changed. Read the tests. Then check the foundations they stand on.
This is the kind of human review we run at Flashback before an agent’s output goes anywhere near a main branch. AI does the work; a person decides whether it is good.

