Why Is Opus 5.5 Such a Big Deal? The Scary Part Is Not “Smarter AI” but “AI That Actually Finishes the Job”

Right after release, people posted anecdotes about running High effort for hours while barely denting their usage allowance, or running several creative and…

How reading tools work

Listen reads the article aloud. Speed read shows phrases in sequence at your chosen pace. Language practice compares available translations. Save keeps a bookmark in this browser; find it in the player’s bookmarks.

Share this article
Advertisement
Advertisement

Five-second takeaway: Claude Opus 5.5 is not just about higher one-shot accuracy. The bigger shift is that long coding jobs, giant repositories, tool-heavy workflows, self-checking, and multi-step completion appear to require fewer turns and fewer tokens than Opus 5. API prices also fell. The frontier is moving from “Who wins the IQ test?” to “Who can keep working until the task is actually done?” [1]

1. Social media screamed “Opus is insane,” but the usage bar is not the real story

Right after release, people posted anecdotes about running High effort for hours while barely denting their usage allowance, or running several creative and coding tasks in parallel.

Those reports are entertaining, but a usage bar is not a scientific benchmark. Plan tier, cache hit rate, effort level, workload, and rate-limit resets can all change what “8% used” actually means.

This launch is different because the anecdotes are backed by a substantial efficiency claim. Anthropic says Opus 5.5 costs about 40% less than Opus 5 on typical workloads, generates output more than 30% faster, and is priced at $4 per million input tokens, $20 per million output tokens, and $0.20 per million cache-read tokens. Opus 5 was $5, $25, and $0.50 respectively. [1]

The important point is not “it answers more briefly.” It is that the model appears to waste fewer steps getting from assignment to completion.

2. The jump from 5 to 5.5 is as much about work habits as raw intelligence

The examples are unusually concrete.

Anthropic reports that an early tester completed a 680,000-line code migration in less than a day. In another case, Opus 5.5 audited and fixed a 200,000-line codebase in under three hours; Opus 5 took more than 20 hours and used roughly 2.5 times as many tokens. [1]

In a web-app optimization test, Opus 5.5 succeeded in 39 of 40 attempts to reduce load times across the application. Early testers from GitHub, Lovable, Kiro, and others also described fewer retries, more complete edits after gathering context, and fewer steps or tool calls to solve tasks. [1]

The classic agent failure mode is painfully familiar:

  1. investigate the problem;
  2. fix half of it;
  3. discover another issue;
  4. proudly announce the discovery;
  5. wait for the human to say, “Great. Now actually fix it.”

Opus 5.5 seems designed to reduce this “intermediate-results delivery service” and connect investigation, modification, testing, correction, and completion.

3. Is it better at exploration than GPT-6 Astra? That depends on what “exploration” means

Treating exploration as one scalar is a good way to get confused.

In Anthropic’s comparison table, Opus 5.5 scores 66.4% on Terminal-Bench 4.0 versus Astra’s 57.9%. On FrontierCode Main it scores 54.4% versus 53.3%. On GDPval-AA knowledge work it reaches 1846 Elo versus 1542. [1]

But Terminal-Bench Science 0.1 goes the other way: 58.7% for Opus 5.5 versus 64.6% for Astra. AutomationBench is also slightly higher for Astra, 41.4% versus 40.0%. [1]

OpenAI separately reports Astra at 99.9% on ARC-AGI-3 and 64.6% on Terminal-Bench Science, emphasizing adaptation to novel environments and scientific workflows. [2]

A rough practical split looks like this:

Type of exploration Current tendency
Read a giant repo, isolate causes, and keep patching Strong case for Opus 5.5
Multi-file edits, tests, and repeated correction Strong case for Opus 5.5
Hold the same objective through a very long coding session Major Opus 5.5 improvement
Scientific software, simulations, fitting, and research workflows Astra scores higher
Learning the rules of a genuinely novel environment Astra is exceptionally strong
Browser, computer, and cross-app agent work Harness and task matter heavily

So this is neither “Astra is obsolete” nor “Opus wins everything.”

Exploration is not only having a brilliant idea. It is creating branches, trying them, discarding failures, preserving useful context, and continuing. Token efficiency and tool discipline therefore become part of effective intelligence.

4. GPT-6 Sol got cheaper too, so AI labor suddenly has lower overtime rates

OpenAI is playing the same economics game.

The official GPT-6 Sol model page lists $2 per million input tokens, $10 per million output tokens, and $0.20 per million cached-input tokens. It also lists a 1.05-million-token context window and up to 128,000 output tokens, with the model positioned for complex coding and agentic workflows. [3]

That changes the competitive question.

It is no longer only:

“Which model is smartest?”

It is increasingly:

“Which model can make 100 tool calls, run for hours, and still leave both the budget and rate limit alive?”

Cheaper inference buys more experiments, more verification passes, more branches, and more subagents.

A price cut in intelligence is also a price cut in trial and error.

5. What do Claude plans cost? Starting at $20 does not mean “everything is unlimited”

Anthropic currently lists Free at $0, Pro at $20 per month or $200 per year, Max 5x at $100 per month, and Max 20x at $200 per month. Max provides roughly five or twenty times the Pro session capacity and includes Claude Code. [4][5][6]

The important part is not to treat the subscription, interactive Claude Code, the API, and non-interactive execution as one wallet.

A Pro or Max subscription does not make the Anthropic API unlimited. Normal API-key usage remains metered, and Opus 5.5’s base API pricing is $4 input and $20 output per million tokens. [1][4]

Interactive Claude Code, meanwhile, uses the plan’s included allowance.

Since June 15, 2026, eligible Pro, Max, Team, and Enterprise users have had a separate arrangement for the Agent SDK and non-interactive Claude Code via claude -p. That usage no longer counts against the normal Claude plan limits, and eligible users can claim a separate monthly Agent SDK credit: $20 on Pro, $100 on Max 5x, and $200 on Max 20x. Pay-as-you-go Claude Platform accounts using an API key do not receive that credit. [7]

So the lesson is simple: “I pay a subscription, therefore everything is unlimited” is wrong, but splitting interactive work from automation can make the economics surprisingly attractive. AI pricing has somehow become a telecom plan with more footnotes.

6. Can you put Opus 5.5 inside OpenCode? Yes through the API; subscription piggybacking is different

OpenCode supports many model providers, so using an Anthropic API key with Claude models is a normal architecture. [8]

The tempting idea is: “If I already pay for Claude Pro or Max, can OpenCode just consume that flat-rate allowance?”

OpenCode’s current provider documentation says plugins exist for using Claude Pro/Max models through OpenCode, but Anthropic explicitly prohibits that use, and OpenCode stopped bundling those plugins as of version 1.3.0. [8]

So the clean routes are:

  • OpenCode + Anthropic API: flexible, but metered.
  • Claude Code + Pro/Max: Anthropic’s supported route, with interactive usage consuming subscription limits.

This also makes a mixed architecture plausible: GPT-6 Sol for cheap high-volume execution, Opus 5.5 for long repository-scale completion work, and Astra for the hardest scientific or novel-environment exploration.

AI has developed job specialization. At least it still does not schedule a status meeting.

7. The real tectonic shift is that delegation itself changes

Older frontier models could be brilliant advisers yet awkward workers.

They often ended with:

“I investigated the issue.”
“I found the root cause.”
“The next step would be…”

Then the human had to reply, “Yes. Do the next step.”

What matters in Opus 5.5 is the reported ability to keep going: one early tester described more than 18 hours of unattended work across six repositories, while other examples include huge migrations, audits, and stronger self-verification loops. [1]

If that behavior generalizes, prompts change from:

“Fix this function.”

to:

“Achieve this outcome. Investigate the cause, make the necessary changes, test them, repair failures, and keep going until the completion criteria are satisfied.”

And model evaluation changes too. Instead of staring at a single benchmark crown, practical teams can measure:

  • completion rate;
  • number of human follow-up prompts;
  • total tokens;
  • total tool calls;
  • elapsed time;
  • self-recovery after failures.

There may never be one universal “best model.”

But Opus 5.5 is a useful symbol of the new phase: the frontier is no longer only about how smart the model looks in an interview. It is increasingly about whether it shows up, stays on task, and finishes the shift.

The AI-agent era has apparently reached the point where attendance records are becoming as interesting as aptitude tests.

8. So how fast does the quota actually disappear? Astra is heavy even in OpenAI’s own table

Now we reach the internet’s favorite benchmark: watching the usage bar move.

OpenAI’s current help documentation says Work and Codex share an allowance, with both five-hour and weekly limits on some plans. Its estimated local messages per five-hour period are 5–45 for Astra on Plus, 25–225 on Pro 5x, and 100–900 on Pro 20x. For GPT-5.6 Sol, the corresponding estimates are 10–100, 50–500, and 200–2,000. [9]

These are not fixed message caps. Actual usage changes with task complexity, reasoning settings, and context. But the official table already frames Astra as a model that does fewer requests per included allowance than Sol.

Community measurements point in the same direction. One user tracked low-cost Claude and Codex plans for twenty days and estimated that Claude provided roughly three to five times the effective weekly API-equivalent allowance, depending on model. [10] Another long-time Codex user reported exhausting a $200-tier weekly allowance in about twenty hours of Astra High coding. [11]

On Hacker News, one user reported burning about 70% of a $100 Astra weekly allowance in roughly five hours, while another reported only 5% weekly usage after eight hours of deep algorithmic work with Opus 5.5. [12]

Do not turn that into “Claude is always fourteen times cheaper.” These are not controlled experiments. Workload, effort, cache behavior, subagents, and reset timing differ. Anthropic also raised five-hour limits when Opus 5.5 launched. [1]

The defensible takeaway is narrower:

Astra is powerful but expensive in quota terms. Opus 5.5 currently looks competitive not only on capability, but on how many hours of serious work fit into a week.

AI comparison has moved from horsepower to fuel economy.

9. Claude Code starts at $20 per month; you do not have to jump straight to $100

The subscription entry point for Claude Code is Pro: $20 month-to-month, or $200 per year, equivalent to about $17 per month. Max 5x is $100 per month and Max 20x is $200. [6][5]

So if the goal is simply to test Claude Code, the minimum is $20, not $100 or $300.

Max is mainly about capacity. Max 5x provides roughly five times Pro usage per session, and Max 20x roughly twenty times. [5]

A sensible comparison therefore uses the same repository and similar tasks, then records:

  • time until the five-hour limit matters;
  • weekly percentage consumed;
  • number of human interventions;
  • whether tests were completed;
  • number of unnecessary edits.

A higher tier is less like “buying a smarter brain” and more like giving a good worker a longer shift after proving they are useful.

10. What should move from Codex to Claude Code? Not the “brain” — the Git history and the handoff

When switching coding agents, it is tempting to think every old chat needs to be imported and somehow “learned.”

In practice, a giant transcript is usually less useful than a compressed set of decisions stored in the repository.

Claude Code automatically reads a root-level CLAUDE.md at the start of each session. Anthropic recommends using it for build commands, test commands, directory structure, coding conventions, and durable constraints, and suggests keeping it under roughly 200 lines. [13]

A practical migration structure is:

  • Git repository: source of truth for code and history;
  • AGENTS.md: reference for rules used by the previous agent;
  • CLAUDE.md: durable rules Claude should read every session;
  • docs/AI-HANDOFF.md: current state, unresolved problems, and recent decisions;
  • Git history: evidence for why the system looks the way it does.

Claude Code can also draft a CLAUDE.md with /init. [13]

Do not turn CLAUDE.md into “all human knowledge.txt.” Because it loads repeatedly, excessive detail dilutes the rules that matter. Do not place secrets, tokens, credentials, or connection strings in it either. [13]

Migration is therefore not model training.

You are not copying one model’s brain. You are putting the project constitution and the handoff notes into Git.

That also makes the next model switch far less dramatic.

11. The first Claude prompt should be “reconstruct the world from the repo” before “start changing things”

For the first session, asking Claude to restore context before editing reduces accidental redesign.

Example handoff prompt

Take over this repository from another coding agent.
Before changing code, inspect the repository, README, AGENTS.md, docs,
package scripts, CI/deployment configuration, and Git history.
Identify the sources of truth, build/test/deploy paths, completion criteria,
forbidden actions, and unresolved issues.
Then place stable rules in CLAUDE.md and mutable current-state information
in docs/AI-HANDOFF.md.
Never display, record, or commit secrets, tokens, or credentials.
Resolve uncertainty from repository evidence and history before asking me.
Do not call the task complete while tests or reachable production checks remain.

The advantage is simple: the human does not have to retell the entire history of the project.

And because only active decisions are compressed into the handoff, context stays cleaner than dumping months of raw chat logs into a new agent.

Apparently the most important part of AI employee onboarding is still documentation.

12. What is the feature that keeps working after the laptop closes? Cloud execution, not logout magic

The terminology matters.

Claude Code on the web clones a GitHub repository into an Anthropic-managed remote environment and works inside an isolated virtual machine. Once a task starts, you can leave the page and it continues. When finished, it can push changes to a branch and prepare a pull request for review. [14]

Claude Code Desktop also has “Continue with Claude Code on the web,” which moves a local desktop session into the cloud so it can be continued from the web or mobile. [15]

Recurring work has another split:

  • /loop: recurring local execution, intended for local jobs that can run for up to about three days; it depends on the local machine.
  • /schedule: Cloud Jobs; the work continues in the cloud even when the laptop is closed. [16]

Routines go further: a prompt, repository, and connectors can be packaged into an automation that runs on Claude Code’s web infrastructure on a schedule, via API call, or in response to an event. [17]

So the trick is not keeping a local process alive through sheer determination. It is moving the job itself to cloud infrastructure.

Closing the browser no longer means the AI clocks out. The human can leave first.

Local-only files or actions that require the physical computer are a separate matter: if the cloud environment cannot see them, it cannot retrieve them by telepathy.

By 2026, picking a coding AI is no longer just “which model is smartest?”

It is capability × completion rate × weekly fuel economy × handoff quality × cloud execution.

That multiplication explains why Opus 5.5 plus Claude Code suddenly looks much more interesting than a benchmark table alone would suggest.


References (17)

  1. Anthropic — “Introducing Claude Opus 5.5” (2026-09-22) Used for release claims, pricing, speed, Opus 5 comparisons, benchmark table, long-running coding examples, early-tester reports, and efficiency claims anthropic.com
  2. OpenAI — “GPT-6 Astra: A new generation of intelligence” Used for Astra benchmark results, ARC-AGI-3, Terminal-Bench Science, coding comparisons, and positioning around novel-environment adaptation openai.com
  3. OpenAI Developers — GPT-6 Sol model page Used for GPT-6 Sol model positioning, context/output limits, and advertised token pricing developers.openai.com
  4. Anthropic Help Center — “Choose a Claude plan” Used for Free, Pro, Max 5x, and Max 20x consumer pricing support.claude.com
  5. Anthropic Help Center — “What is the Max plan?” Used for Max usage multipliers and Claude Code access support.claude.com
  6. Anthropic Help Center — “What is the Pro plan?” Used for Pro pricing and Claude Code inclusion support.claude.com
  7. Anthropic Help Center — “Use the Claude Agent SDK with your Claude plan” Used for the June 15, 2026 separation of Agent SDK / claude -p usage from normal interactive plan limits, and the separate monthly credits for eligible Pro, Max, Team, and Enterprise users support.claude.com
  8. OpenCode — “Providers” Used for Anthropic provider setup and OpenCode’s current warning that using Claude Pro/Max through OpenCode is explicitly prohibited by Anthropic and that such plugins stopped being bundled from OpenCode 1.3.0 opencode.ai
  9. OpenAI Help Center — “Managing usage with GPT-6 Astra in Work and Codex” Used for shared Work/Codex allowance mechanics, five-hour and weekly-limit caveats, and estimated local messages per five-hour period for Astra, Sol, and other models help.openai.com
  10. Reddit r/codex — “Measured Claude usage limits are 3–5× those of Codex” (2026-09-18) Community measurement based on roughly twenty days of usage logging. Treated as anecdotal observational evidence, not a controlled benchmark reddit.com
  11. Reddit r/codex — “Codex is in a really bad spot right now…” (2026-09-23) Used only for community reports about heavy Astra weekly-quota consumption and Opus 5.5 usage experience. Not treated as vendor-independent benchmark proof reddit.com
  12. Hacker News — “Claude Opus 5.5” discussion (2026-09-23/24) Used for contrasting real-user usage anecdotes, including heavy Astra quota consumption and low Opus 5.5 weekly consumption in long algorithmic work. These reports are uncontrolled news.ycombinator.com
  13. Anthropic Help Center — “Give Claude context: CLAUDE.md and better prompts” Used for CLAUDE.md automatic loading, /init, recommended content, and keeping project instructions concise support.claude.com
  14. Anthropic Help Center — “Claude Code on the web” Used for remote GitHub execution, isolated environments, leaving the page while work continues, and branch/PR workflows support.claude.com
  15. Anthropic — “Preview, review, and merge with Claude Code” (2026-02-20) Used for “Continue with Claude Code on the web” and moving a desktop session into the cloud claude.com
  16. Anthropic Help Center — “Claude Code power user tips” Used for /loop as local recurring execution and /schedule as Cloud Jobs that continue when the laptop is closed support.claude.com
  17. Anthropic — “Introducing routines in Claude Code” (2026-04-14) Used for cloud routines triggered by schedules, API calls, or events claude.com

AdFind what this article discusses

This article contains affiliate links (ads). About advertising As an Amazon Associate I earn from qualifying purchases.

Advertisement

Find other articles

All articles

Mendoi-chan

Written by

Mendoi-chan

She turns friction at work and in everyday life into clear structure and practical next steps.

About
Advertisement

Latest articles

  1. 1The Black Knights Should Have Retreated When Zero Left|Todo and the Limits of an Organization Built Around One Person
  2. 2The Hell of Watching Code Geass in Real Time: Waiting from Season 1 Episode 25 to R2
  3. 3Early June Summer Events to Enjoy Before It Gets Too Hot
  4. 4A Blue Moon Is Not a Blue-Colored Moon
  5. 5“You Never Reply” — Even Though You Do: What Happens When One Person Outsources the Conversation Engine

You may also like

Advertisement