2026-07-21
This week the top of the board stayed put, and the reason is momentum. GPT-5.6 finished its broad rollout on July 9 and is now the default model inside ChatGPT, so the assistant most people actually open every morning got quietly sharper at daily chat and knowledge work. Pair that with the ecosystem reach ChatGPT already owns, roughly 54 percent of worldwide traffic across the biggest chatbots, and its number one spot is the easy call. Claude holds second on the strength of code. Fable 5 returned as the coding leader at 80.3 percent on SWE-Bench Pro, and in my own repo work it lands multi-file changes with the fewest retries, which is why its coding score stays the highest in this table. Gemini keeps third by being the accuracy pick, with 3.1 Pro strong on hardest-mode reasoning and 3.5 Flash still the best price-performance at the frontier. The one change I made is Grok. xAI shipped Grok 4.5 on July 8 with a clear coding focus, and it now reasons at a level worth a small bump on that factor, so I nudged its coding score up a tenth. It stays fourth because value and ecosystem still trail the leaders. Perplexity, Copilot, and DeepSeek round out the middle, each winning a lane: live search, Microsoft integration, and raw price. My advice remains simple. Pick by the job in front of you, since the gap between the top four is now measured in preferences, not capability.
GPT-5.6 is the new default
Since July 9 the model most ChatGPT users get by default is GPT-5.6, which sharpened everyday chat and knowledge work. That plus the widest ecosystem keeps ChatGPT at number one.
Claude owns coding
Fable 5 posts 80.3 percent on SWE-Bench Pro and lands multi-file edits with fewer retries in real repo work, so Claude keeps the highest coding score and second overall.
Grok 4.5 earns a small bump
xAI's July 8 release leans hard into coding and Opus-class reasoning with native X grounding. I raised Grok's coding score by a tenth while keeping its rank at four.
Pick by task, not by hype
Gemini for accuracy, Perplexity for live search, Copilot for Microsoft workflows, DeepSeek for price. The top four are now separated by preference more than capability.