If you are picking a model for a structured critique job, the obvious question is whether the more expensive option earns its price. We ran that comparison on our critique.pdf_guide job across 288 calls over 90 days, and the numbers tell a clear story.
Sonnet (anthropic/claude-sonnet-4-6) handled 241 of those calls. Opus (anthropic/claude-opus-4-8) handled 47. Both ran on the same job, against the same prompts. Neither produced a single failure on our side.
We stopped using Sonnet on 2026-07-25, the same date Opus first appeared in the record. The record does not say why we switched.
What the job asks each model to do
The critique.pdf_guide job sends a prompt to the model and gets back a short critique. Output was 27 tokens per call on Sonnet and 31 on Opus, so the two models were producing roughly the same volume of text.
Input is where the picture is more complicated. Sonnet received 2,849 total input tokens per call. Opus received 3,752. That is a different job size, not just a different model, and it matters when you read the cost figures.
The record does not say whether the prompts changed when we switched models. That uncertainty sits behind every cost comparison in this piece.
Cost per call: Opus is 2.26 times more expensive
Sonnet cost us $0.00473 per call. Opus cost us $0.01067 per call. Opus is 2.26 times the cost of Sonnet, or $0.00594 more per call on our volume.
Before reading those figures as a quality gap, look at what each model was actually billed for. Sonnet had 2,849 total input tokens per call but was billed on only 1,026, because 56.8 percent came from the prompt cache. Opus had 3,752 total input tokens per call and was billed on 1,319, with 55.2 percent cached.
The billed token counts are closer than the total counts suggest, and the cache hit rates are nearly identical at 56.8 percent and 55.2 percent. The cost difference is real, but it reflects a larger job going to Opus as much as it reflects Opus being a pricier model. We cannot separate those two factors from this data alone.
What we can say is that neither model was paying heavily for cache misses. Both were caching more than half their input on every call. The cost gap comes from the larger total input on Opus calls and from whatever rate difference applies at our volume, not from one model failing to cache.
Speed: practically the same
Sonnet had a median response time of 2.1 seconds. Opus came in at 2.2 seconds, which is 1.05 times slower and 0.1 seconds behind. For a critique job, that gap is not a practical concern.
Speed is not a reason to prefer one model over the other here. The two figures are close enough that other factors will drive any real decision.
Failures: none on either side
Both models returned zero failures across all calls in this window. Sonnet ran 241 calls with a 0.0 percent failure rate. Opus ran 47 calls with a 0.0 percent failure rate.
A failure rate is what came back as a failure on our side. It is not a statement about either vendor's reliability in general.
The fact that Sonnet is listed as the model with the most failures in the underlying data is technically accurate, and it is also zero. Both entries are zero.
The limits of this comparison
These numbers come from one job, our prompts, our input sizes, our call volumes, over 90 days ending 2026-09-08. Sonnet's calls ran between 2026-06-11 and 2026-07-25. Opus's calls ran from 2026-07-25 onward. The two windows do not overlap.
The input token counts differ between the two models. We do not know from this data whether that is because the prompts changed when we switched, or for some other reason. That means the per-call cost comparison carries a caveat: we are not confident both models were asked exactly the same question.
This is not a benchmark. A reader running a different job, with different prompt lengths or different cache conditions, should expect different figures. The cost per call figures here are what we paid on our volume, not a rate card.
Opus ran 47 calls against Sonnet's 241. The smaller sample on Opus means its figures carry more variance, and a single unusual call moves its averages more than the same call would move Sonnet's.
Which one we kept and why
Opus is the model still in use as of 2026-09-08. Sonnet's last call was 2026-07-25, which is the day Opus appeared. The record does not say why we switched.
The cost difference is 2.26 times per call. Whether that is worth it depends on what the job produces and what you need from it. The record does not give us output quality figures, so we cannot make that argument here.
What the record does show is that we moved to Opus and have stayed there across 47 calls. That is the plainest summary of what happened.
We would pick Opus today, because it is the tool still in use. If cost per call became a constraint, the gap is $0.00473 for Sonnet against $0.01067 for Opus, and the caveat is that the two models were not handling identical input sizes in this record. We would not switch back on speed alone: 0.1 seconds is not a reason to change anything.
Originally published at neuragrowth.co. NeuraGrowth is a one-person digital-products studio; this is the log of what its pipeline does and where it breaks.












