xAI · Grok 4.5 · Jul 2026

Grok 4.5 — a coding model with live search and code execution

  • ~500K context
  • Leaner reasoning tokens
  • Search & execution ready
  • Editor-close coding
Grok 4.5

What is Grok 4.5?

Grok 4.5 is xAI’s July 2026 release, focused on multi-step coding in editor settings and on wiring web search, X search, and code execution into one workflow. Versus Grok 4.3, it emphasizes this kind of tool collaboration and token efficiency; context is ~500K—when materials run especially long, compare Grok 4.3, which still offers million-token context. On iMini Agent, pick Grok 4.5 to use it.

Vendor
xAI
Released
Jul 2026
Context
500K tokens
Max output
Per platform limit
Deep reasoning
Leaner tokens · tools ready
Best for
Coding, retrieval & code execution

What’s new versus Grok 4.3

Agent ability, token efficiency, and editor settings—the changes to check before making it your primary pick.

Stronger multi-step coding

xAI positions 4.5 as a stronger coding and tool-collaboration model—suited to multi-step edits and tool-connected jobs.

Leaner reasoning tokens

Uses tokens more tightly at similar quality—suited to high-turnaround tool loops.

Full tool surface

Can wire function calling, web search, X search, and code execution—so retrieval and execution live in one workflow.

Editor-close coding

Co-trained with products like Cursor, closer to the habit of editing a repo and making multi-step changes in an IDE.

Official evaluations

Engineering scores xAI published with the Grok 4.5 launch, shown against Claude Fable 5, Claude Opus 4.8, GPT-5.5, and peers. Competitor scores are taken from each vendor’s public system card or leaderboard.

DeepSWE 1.0: Grok 4.5 vs Fable, GPT-5.5, and Opus
DeepSWE 1.0 (pass@1). Grok 4.5 at 62.0%—above Claude Opus 4.8 (55.75%), below Claude Fable 5 (66.1%) and GPT-5.5 (64.31%). Datacurve builds the eval; AA runs each vendor’s harness.
DeepSWE 1.1: Grok 4.5 comparison
DeepSWE 1.1. Grok 4.5 at 53.0%—below Claude Fable 5 (70%), GPT-5.5 (67%), and Claude Opus 4.8 (59%), above GLM 5.2 (44%). Read both DeepSWE versions in the same family; don’t cherry-pick the higher score.
Terminal Bench 2.1: Grok 4.5 comparison
Terminal Bench 2.1. Grok 4.5 at 83.3%—nearly tied with GPT-5.5 (83.4%), above Claude Opus 4.8 (78.9%), slightly below Claude Fable 5 (84.3%). A core comparison for terminal-agent settings.
SWE Bench Pro: Grok 4.5 resolve rate comparison
SWE Bench Pro resolve rate. Grok 4.5 at 64.7%—above GPT-5.5 (58.6%) and GLM 5.2 (62.1%), below Claude Fable 5 (80.4%) and Claude Opus 4.8 (69.2%).

Three common workflows

Grok 4.5 fits work that chains coding, retrieval, and execution together.

Multi-step coding in the IDE

For multi-step features, test fixes, and tool collaboration. A strong coding pick on the Grok line.

Research with search

For knowledge work that needs web or X retrieval. Keep sources and conclusions clearly separated.

Checks with code execution

For short-to-mid jobs that write, run, and correct conclusions from execution results.

How to choose vs Grok 4.3 and GPT-5.6 Sol

All three are modern primary picks. The gap is mainly context length, product stack, and agent posture.

CriterionGrok 4.5Grok 4.3GPT-5.6 Sol
Context500K tokens1M tokens1M tokens
ReasoningAgent-oriented · more efficientAdjustable depth levelsMaximum depth · ultra available
Coding & agentsxAI flagship coding & agentsLong docs & tool callsOpenAI’s heaviest long-horizon coding
Long-horizon toolsTools & search readyMore headroom in long threadsFlagship depth + ultra
Speed postureFavors tool-collaboration efficiencyCan trade depth for speedFavors finish quality
Prefer whenCoding & agents on the xAI stackGrok jobs that need million-token contextYou’re on OpenAI’s depth stack

Using Grok 4.5 on iMini

Free quota, model naming, key specs, and how it differs from 4.3 and Sol.

Use Grok 4.5 free on iMini Agent

~500K context—suited to multi-step coding with retrieval and code execution.

Open Grok 4.5

No xAI API key required.