Claude Opus 5.5 Costs 20% Less. Here Is How to Actually Use That.
Anthropic's new model is cheaper, faster and, it says, as capable as its premium tier. Six practical things to do before you switch.
Anthropic released Claude Opus 5.5 on September 22. The headline numbers: input costs $4 per million tokens (down from $5), output costs $20 per million (down from $25), and cached reads fall 60%, to $0.20. Anthropic says the model performs at the level of Claude Fable 5.1, its more expensive model released earlier in September, and runs about 40% cheaper than Opus 5 on typical workloads.
Price cuts are easy to announce and easy to misread. Here is what the release changes in practice, and what to do about it.
What a "token" price actually means for you
A token is a chunk of text, roughly three-quarters of a word. You pay separately for what you send in (your instructions, documents and conversation history) and what comes back. Output is five times the price of input, so a model that writes long answers costs far more than one that reads long documents.
A worked example: a job that sends 1 million tokens and gets 200,000 back cost $10 on Opus 5 ($5 plus $5). On Opus 5.5 it costs $8 ($4 plus $4). That is the 20%. The larger saving, and the one most teams will miss, is in caching.
Six things to do
1. Turn on prompt caching. If your application sends the same long context repeatedly, such as a product catalogue, a policy manual or a codebase, caching lets the model reuse it. Cached reads now cost $0.20 per million instead of $4. For a support bot that sends the same 50-page handbook with every question, this is the single biggest cost lever you have.
2. Use the effort setting as your cost dial. Opus 5.5 offers five effort levels, from low to max, and defaults to medium. Higher effort means more reasoning, better answers on hard problems and a bigger bill. Start at the default and raise it only for the tasks that measurably need it, rather than setting everything to maximum.
3. Use fast mode only where someone is waiting. Fast mode runs up to 2.5 times quicker, at double the price ($8 in, $40 out). It is worth it for a live chat a customer is watching. It is wasted on an overnight report.
4. Re-run your numbers, not the benchmarks. Anthropic reports strong scores on its chosen tests, including 66.4% on Terminal-Bench 4.0. Those are the vendor's results. Take 20 to 50 real tasks from your own work, run them on your current model and on Opus 5.5, and compare quality and cost side by side before switching anything in production.
5. Update the model name deliberately. The identifier is claude-opus-5-5. Change it behind a setting you can flip back, not in scattered code, so a regression is a one-line rollback.
6. Expect the rest of the family. Anthropic has said Sonnet 5.5 and Haiku 5.5 are coming. For high-volume, simpler tasks, a smaller model is often the better choice, so avoid locking every workload to Opus.
Questions You Should Be Asking
- What share of our AI bill is output rather than input, and would shorter answers save more than a new model?
- Are we resending the same context on every request without caching it?
- Which of our tasks genuinely need maximum effort, and which are paying for it by default?
- Have we tested this model on our own work, or only read the launch post?
What To Watch Next
The release of Sonnet 5.5 and Haiku 5.5. If the smaller models land with similar price cuts, the right move for most businesses will be to route simple work to them and keep Opus for the hard cases.
Sources
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
