OpenAI · GPT-5.5 · April 2026

GPT-5.5 — general-purpose flagship for coding, analysis, and tool use

  • Million-token context
  • Extra-long single replies
  • Tunable depth tiers
  • Agent-coding workhorse generation
GPT-5.5

What is GPT-5.5?

GPT-5.5 is OpenAI’s flagship released in April 2026 for complex professional work: agent coding, computer use, knowledge work, and early research. Versus GPT-5.4, OpenAI emphasizes stronger agent capability and lower token use. With GPT-5.6 available, the heaviest new jobs can compare Sol and everyday balanced work can compare Terra; if your workflows already run stably on GPT-5.5, it remains a solid primary. On iMini Agent, pick GPT-5.5 to use it.

Vendor
OpenAI
Released
April 2026
Context
1M tokens
Max output
128K tokens
Deep reasoning
Tunable tiers
Best as
Flagship primary for validated workflows

What’s new versus GPT-5.4

Agent coding, intelligence at matched latency, and knowledge work—what to check before moving from 5.4 to 5.5.

Stronger agent coding

OpenAI reports higher completion on agent coding evals such as Terminal-Bench—better suited to multi-step coding and tool collaboration.

More intelligence at similar latency

Versus GPT-5.4, OpenAI emphasizes stronger intelligence at comparable serving latency, and finishing Codex-style work with fewer tokens.

Steadier long-context professional work

Million-token context keeps long specs, multi-source materials, and tool traces in one thread—suited to knowledge work and complex delivery.

Higher research and security eval ladder

On GeneBench and cybersecurity-related evals, OpenAI reports step-ups versus the prior generation—useful before deciding whether to move to GPT-5.6.

Official evaluations

Results OpenAI published with GPT-5.5, compared with GPT-5.4 and peer flagships—covering composite intelligence, software engineering, knowledge work, computer use, and security research.

Artificial Analysis Intelligence Index: composite intelligence comparison
Third-party composite intelligence index with output-token volume on the x-axis. GPT-5.5 sits above GPT-5.4 and Claude Opus 4.7 throughout, and reaches matched scores with less output—an upgrade that also finishes work more efficiently.
Terminal-Bench 2.0: command-line software engineering tasks
Terminal-Bench 2.0 measures real command-line software engineering. GPT-5.5 peaks near 83%, clearly above GPT-5.4, and clears 75%+ even at lower output settings—multi-step terminal jobs stay solid without maxing depth.
Expert-SWE: expert-level software engineering tasks
Expert-SWE is authored by senior engineers to measure expert software engineering. GPT-5.5 peaks near 73%, about 5 points above GPT-5.4 at matched output—lower rework on large coding jobs.
GDPval: knowledge-work win rate versus industry experts
GDPval blinds model outputs against industry-expert deliverables. GPT-5.5 reaches 84.9% win+tie, above GPT-5.4 (83.0%) and Claude Opus 4.7 (80.3%)—reports and table work already above expert baseline.
OSWorld-Verified: graphical desktop computer use
OSWorld-Verified measures finishing tasks in a real GUI. GPT-5.5 approaches 80% with fewer tool-call turns—computer-use agents run faster and leaner.
Tau2-bench Telecom: multi-turn tool-calling support flows
Tau2-bench Telecom simulates multi-turn business flows with tool calls. GPT-5.5 approaches 98% accuracy with about one-third of GPT-5.4’s output—direct evidence for plugging it into business automation.
GeneBench: biological science reasoning
GeneBench measures long-horizon biological science reasoning. GPT-5.5 reaches 25% versus GPT-5.4’s 19%, with less output needed—clear gains on scientific long reasoning.
Capture-the-Flags: security offense/defense challenge tasks
CTF challenges measure end-to-end security research. GPT-5.5 hits 88% with about 50K output tokens; GPT-5.4 needed nearly 140K to reach 84%—one reason OpenAI applies its strongest safety protections here.

Three common workflows

GPT-5.5 still fits validated complex professional work; new heaviest loads can compare GPT-5.6.

Agent coding & repo changes

For multi-step coding, test repair, and cross-file features. Teams whose prompts and evals already bind to GPT-5.5 can keep shipping on it.

Knowledge work & decision memos

For research synthesis, table work, and one-page decision memos. Quote hard constraints and give a checkable recommendation.

Tool and computer-use flows

For workflows that call tools repeatedly and finish end to end. Pin success checks and abort conditions up front.

How to choose vs GPT-5.6 Sol and GPT-5.6 Terra

All three are OpenAI workhorses. The differences are generation depth, everyday cost, and whether your workflow is already validated.

DimensionGPT-5.5GPT-5.6 SolGPT-5.6 Terra
Context1M tokens1M tokens1M tokens
ReasoningTunable depth tiersMax depth · ultra availableBalanced depth
Coding / agentsFlagship coding & agentsHeaviest current long-horizon codingDay-to-day coding & mid-weight agents
Long tool runsValidated workflows can continueFlagship depth + ultraMore balanced completion and throughput
Speed postureProfessional-work primaryCompletion quality and depth firstEveryday primary favors balance
Prefer whenWorkflows already validated on 5.5New heaviest reasoning and long-horizon jobsNear–5.5 quality on the GPT-5.6 balanced tier

Using GPT-5.5 on iMini

Free quota, model naming, key specs, and how it differs from GPT-5.6.

Use GPT-5.5 free on iMini Agent

Million-token context, tunable depth, suited to validated professional workflows.

Open GPT-5.5

No OpenAI API key required.