Anthropic · Claude Opus 4.8 · May 2026

Claude Opus 4.8 — more reliable coding and tool collaboration

  • Million-token context
  • Extra-long single replies
  • Usually high depth
  • Stronger alignment and honesty
Claude Opus 4.8

What is Claude Opus 4.8?

Claude Opus 4.8 is Anthropic’s Opus released in May 2026, with noticeable upgrades versus Claude Opus 4.7 on collaboration judgment, agent reliability, and honesty. It usually works at high depth; harder jobs can go deeper. New long-horizon coding and knowledge work more often pick Claude Opus 5; if workflows are already validated on 4.8—or you need a steady Opus fallback—you can keep using it. On iMini Agent, pick Claude Opus 4.8 to use it.

Vendor
Anthropic
Released
May 2026
Context
1M tokens
Max output
128K tokens
Deep thinking
Usually high depth · tunable
Best as
Primary for validated Opus workflows

What’s new versus Claude Opus 4.7

Honesty, agent judgment, and computer use—what to check before moving from 4.7 to 4.8.

More proactive about uncertainty

Answers more often flag what’s uncertain; missed defects in self-written code drop sharply—suited as a collaboration primary.

More reliable agent judgment

Fewer extra steps in tool calls and multi-step jobs—cleaner paths to the same intelligence outcomes.

Stronger computer and browser agents

Clear lifts on official computer-use evals such as Online-Mind2Web—suited to end-to-end web and desktop flows.

Better alignment audits

Deception and assisting-misuse rates drop sharply versus 4.7, approaching Mythos Preview safety levels at the time.

Official evaluations and safety audits

Results Anthropic published with Claude Opus 4.8, compared with Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro—covering software engineering, multidisciplinary reasoning, computer use, and behavioral alignment.

Claude Opus 4.8 vs Opus 4.7, GPT-5.5, and Gemini 3.1 Pro multi-task evaluation table
Multi-task overview. SWE-Bench Pro 69.2%, OSWorld-Verified 83.4%, GDPval-AA 1890, Finance Agent v2 53.9%—all above Claude Opus 4.7 and on-screen peers; Terminal-Bench 2.1 at 74.6% sits slightly below GPT-5.5 at 78.2%.
Behavioral alignment audit: Claude Opus 4.8 misalignment score
Behavioral alignment audits count deception, induced misuse, and related misalignment—lower is better. Claude Opus 4.8 is about 1.83, below Claude Opus 4.7 and Claude Sonnet 4.6; read this alongside capability scores for production deploy.

Three common workflows

Claude Opus 4.8 still fits validated Opus-tier work; new flagship loads can compare Opus 5.

Coding collaboration & review

For feature work, defect hunting, and evidence-based PR review. Suited to teams whose prompts and evals already bind to 4.8.

Computer use & browser flows

For multi-step jobs that operate a web or desktop environment. Write completion criteria and abort conditions clearly.

Domain professional agents

For legal, finance, and other flows that need steady judgment. New ceiling long-horizon jobs can re-evaluate Opus 5 or Fable 5.

How to choose vs Claude Opus 5 and Claude Sonnet 5

All three are Claude workhorses. The differences are generation ceiling, high-frequency cost, and whether your workflow is already validated.

DimensionClaude Opus 4.8Claude Opus 5Claude Sonnet 5
Context1M tokens1M tokens1M tokens
ReasoningUsually high depth, tunableDeep thinking built in, tunableTunable depth tiers
Coding / agentsOpus coding & collaborationCurrent long-horizon coding flagship primaryScaled agents & everyday coding
Long tool runsSteady Opus-tier collaborationStronger goal holding and completionBetter as a high-frequency primary
Speed postureQuality and experience in balanceCompletion quality firstThroughput first
Prefer whenWorkflows already validated on 4.8New long-horizon coding and knowledge workNew high-frequency agent primary

Using Claude Opus 4.8 on iMini

Free quota, model naming, key specs, and how it differs from Opus 5.

Use Claude Opus 4.8 free on iMini Agent

Million-token context, usually high depth, suited to validated Opus workflows.

Open Claude Opus 4.8

No Anthropic API key required.