Skip to content
YD.
Email me
← Writing

Aug 15, 2026 · 1 min read

  • Notes
  • AI
  • Agents
  • Tooling

Grok 4.6 is for the long job

xAI’s Grok 4.6 matches Sol on a composite index and trails it on DeepSWE. The useful claim is the long loop: a first pass of a working interface, not a quiz score.

Field note

xAI shipped Grok 4.6 on 12 August. The pitch is a model that stays with a job across many steps — research, a codebase, a first version of an app — and then checks its own work.

Their Grok 4.6 post puts High at 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol Max. Fable 5 Max is 62. CursorBench 3.2 is 69.9% for Grok against 67.2% for Sol. DeepSWE v1.1 goes the other way: Sol 73%, Grok 65.9%. Terminal-Bench 3.0 is 26% versus 34.6%.

A mineral-white studio workbench with a near-black laptop and a long coiled cable.
Grok 4.6 is sold as a long loop, not a quiz score.

What I actually care about

If you ship interfaces, the useful sentence is further down: 4.6 is stronger at turning a product idea into a working first pass, including the visual language of the thing. That is a different skill from answering a coding quiz.

It is in Cursor and Grok Build today, with double included usage for the first week.

I will try it the way I try any new model. Give it a real screen, a real constraint, and see whether the first pass is something I can refine or something I throw away.