πŸ† TopRankLand
← All Rankings
Software

Best AI Chatbots 2026: ChatGPT vs Claude vs Gemini vs Grok Tested

I tested every major AI chatbot in May 2026 and ranked the top 10. Here is who actually wins for coding, writing, reasoning, and real-time research.

Last updated: 2026-09-18 Β· 10 entries tracked daily

Rank Trend β€” Top 10

Lower = better rank. Showing last 89 days.

Current Rankings

#1
ChatGPT OpenAI
Free / $20 Plus / $100 Pro 9.7/10

GPT-5.5 powered chatbot with Sora, Agent Mode, and the largest ecosystem of apps, plugins, and integrations.

Reasoning & Problem Solving 9.8
Coding Capability 9.2
Writing & Creativity 9.4
Real-Time Information 9.0
Value & Pricing 9.4
Ecosystem Integration 9.7
#2
Claude Anthropic
Free / $20 Pro / $100 Max 9.5/10

Sonnet 5 became the new default on June 30 and writes the most natural prose of any model, while Fable 5 returned July 1 and retook the coding crown at 80.3% on SWE-Bench Pro.

Reasoning & Problem Solving 9.6
Coding Capability 9.7
Writing & Creativity 9.7
Real-Time Information 7.6
Value & Pricing 9.4
Ecosystem Integration 9.5
#3
Gemini Google
Free / $19.99 AI Pro / $249.99 Ultra 9.4/10

Gemini 3.1 Pro hits 94.3% on GPQA Diamond and lives natively inside Gmail, Docs, and Sheets.

Reasoning & Problem Solving 9.7
Coding Capability 9.3
Writing & Creativity 8.9
Real-Time Information 9.5
Value & Pricing 9.5
Ecosystem Integration 9.7
#4
Grok xAI
Free / $10 Lite / $30 SuperGrok / $300 Heavy 9.0/10

Grok 4 with native access to X firehose, the only model with truly current social and news data.

Reasoning & Problem Solving 9.1
Coding Capability 9.1
Writing & Creativity 8.4
Real-Time Information 9.8
Value & Pricing 7.0
Ecosystem Integration 7.5
#5
Perplexity Perplexity AI
Free / $20 Pro / $200 Max 8.6/10

Search-first answer engine with sourced citations and the Comet browser, now free across all platforms.

Reasoning & Problem Solving 8.7
Coding Capability 7.6
Writing & Creativity 8.0
Real-Time Information 9.5
Value & Pricing 8.7
Ecosystem Integration 8.5
#6
Free / $20 Pro / $30 M365 8.4/10

GPT-5.5 wrapped inside Word, Excel, PowerPoint, and Outlook for users who live in Microsoft 365.

Reasoning & Problem Solving 8.5
Coding Capability 8.7
Writing & Creativity 8.4
Real-Time Information 8.0
Value & Pricing 8.0
Ecosystem Integration 9.6
#7
DeepSeek DeepSeek
Free 8.2/10

V4-Pro with 1M token context, unlimited free access, and full open-source weights under MIT license.

Reasoning & Problem Solving 8.8
Coding Capability 8.5
Writing & Creativity 7.8
Real-Time Information 7.0
Value & Pricing 10.0
Ecosystem Integration 7.0
#8
Free 7.6/10

Llama 4 inside WhatsApp, Instagram, and Messenger with zero setup and zero cost.

Reasoning & Problem Solving 7.4
Coding Capability 7.0
Writing & Creativity 7.5
Real-Time Information 8.0
Value & Pricing 10.0
Ecosystem Integration 8.5
#9
Qwen Chat Alibaba
Free 7.6/10

Alibaba's open-weight Qwen3 model with image generation and free access for non-commercial use.

Reasoning & Problem Solving 7.5
Coding Capability 7.6
Writing & Creativity 7.0
Real-Time Information 6.8
Value & Pricing 9.0
Ecosystem Integration 6.5
#10
Le Chat Mistral
Free / $14.99 Pro 7.4/10

European AI assistant with strong privacy stance and Magistral reasoning model under EU data residency.

Reasoning & Problem Solving 7.5
Coding Capability 7.3
Writing & Creativity 7.6
Real-Time Information 7.0
Value & Pricing 8.0
Ecosystem Integration 7.0

Today's Analysis Β· 2026-09-18

ChatGPT holds first at 9.7, and the ranking carries forward unchanged. This week's story concerns the youngest people in your household, because the European Commission has put chatbots on the same list as social media.

The proposed EU Kids Act covers social media, video platforms, online games, app stores, chatbots and AI companions. Under the draft, children from 3 to 12 use child-oriented services through accounts a parent fully controls. At 13 and 14 a parent can open an introductory account with limited contacts and strict time limits. At 15 a teenager signs up alone. Fines reach up to 6% of annual sales, and the Commission says its own age-checking app is ready for member states to deploy by year end.

It is a draft. EU countries and the European Parliament both have to agree, and the age thresholds and verification method can still change. Families should still plan around it, because the direction is clear.

My practical advice for parents: pick one assistant for the household and set it up properly. ChatGPT has offered linked parent and teen accounts with content controls and quiet hours since late 2025, which is a strong reason it stays first for families. Gemini connects to Google Family Link, which makes it the natural choice in homes already running Android and Chromebooks. Claude's consumer apps are built for adults, and I recommend it for the grown-ups in the house.

On the product side, The Verge reported on 16 September that Anthropic added native Docs and Slides tools to Claude and merged chat with its Cowork environment into one interface. I have started drafting weekly reports there, and producing an editable slide deck from a conversation saves me real time. It keeps Claude a close second at 9.5.

EU Kids Act puts chatbots in scope

The draft covers chatbots and AI companions for under 15s, with parent-controlled accounts at 13 and 14 and fines up to 6% of annual sales.

Still a proposal

Member states and the European Parliament must approve it, and age limits and verification methods may change.

Best family setup today

ChatGPT's linked parent and teen accounts and Gemini's Family Link support make them the two easiest assistants to manage for a household.

Claude adds Docs and Slides

Anthropic's new native document and slide tools, merged with Cowork in one interface, strengthen Claude's work case at 9.5.

References

Update History

2026-09-17

ChatGPT holds first at 9.7, and the most useful item for ChatGPT subscribers this week sits in OpenAI's own changelog. On 14 September OpenAI announced that GPT-5.5 retires from ChatGPT, ChatGPT Work and Codex on 14 October, across every plan from consumer to Enterprise and Edu. The OpenAI API keeps it. Inside the apps, the replacement is GPT-5.6 Sol.

For most people this is a non-event, because the default model simply changes. The work is in everything that pinned a model by name. OpenAI's notice lists the places to check: workspace defaults, saved model settings, custom agents, scheduled tasks and scripts. If you set up a scheduled morning briefing or a custom agent over the summer, open it now and switch the model yourself. A task you test in September is a task that still runs on 15 October.

I keep ChatGPT at 9.7 because this is exactly how a platform should retire a model: a month of notice, a named successor and a checklist. Its ecosystem score of 9.7 reflects that maturity as much as the feature list.

Claude at 9.5 stays my pick for long writing and document work, with writing at 9.7 and coding at 9.7, both the best scores in this table. Gemini at 9.4 keeps reasoning at 9.7 and realtime at 9.5, and it is the assistant I use when the answer depends on something that happened this morning.

My practical advice applies to every assistant here. Keep one short document listing every automation you built on an AI chatbot, the model it uses and the date you last tested it. It takes ten minutes, and on retirement day it turns a surprise into a two minute fix.

Perplexity at 8.6 and DeepSeek at 8.2 stay where they are, and the ranking carries forward unchanged.

GPT-5.5 leaves ChatGPT on 14 October

OpenAI's changelog names GPT-5.6 Sol as the replacement in ChatGPT, ChatGPT Work and Codex. The API keeps GPT-5.5.

Check anything that pins a model

Workspace defaults, custom agents, scheduled tasks and scripts that select GPT-5.5 need a manual switch before the date.

A clean retirement supports the 9.7

A month of notice, a named successor and a checklist is how a mature platform handles change.

Keep an automation list

Write down each automation, its model and its last test date. It turns retirement day into a quick fix.

2026-09-16

ChatGPT holds first at 9.7, and this week I want to separate two questions readers keep merging: which assistant you pay for, and which assistant answers when you hold the side button on your phone.

iOS 27 shipped on 14 September with Siri rebuilt as a standalone app, trained with Google Gemini models, and the deepest features gated to an A17 Pro chip or newer. After two days with it my read is that the new Siri works as a front door. It handles device control, on-screen context and quick lookups, and the long thinking still belongs to whichever assistant you chose. Your subscription decision survives the update intact.

That decision got better value in September. GPT-6 Astra arrived on 3 September and keeps ChatGPT ahead on reasoning at 9.8, which is why it stays my single recommendation for anyone who wants one tool covering work, study and everyday questions. Gemini holds 9.4, and its 3.8 Flash release on 2 September made the free tier quick enough that I now run it as a second opinion on every research task. Claude at 9.5 remains the one I trust with a 200 page contract, and its writing score of 9.7 is still the highest here.

DeepSeek moves to 8.8 on reasoning after V4.1-Flash landed on 10 September. At 8.2 overall with a perfect value score, it is the assistant I point people to when the monthly fee is the blocker. Grok at 9.0 keeps the freshest feed for live events, and Perplexity at 8.6 stays my pick when I need answers with citations I can open myself.

One practical note for iPhone owners. Install the app of the assistant you pay for, turn on its own voice mode, and put it on your home screen next to Siri. Two front doors, each good at something specific, is the setup I run myself.

Siri is a front door, your subscription is the room

iOS 27 landed on 14 September with Siri as its own app, trained with Google Gemini models. It handles device tasks and on-screen context well, and the assistant you pay for still does the heavy thinking.

GPT-6 Astra keeps ChatGPT at 9.8 on reasoning

The 3 September release holds ChatGPT's lead in the factor that decides most hard questions. For one subscription covering work, study and daily use, this is still where I send people.

Gemini 3.8 Flash makes the free tier a real second opinion

The 2 September release is fast enough that checking an answer against Gemini costs seconds. Gemini scores 9.5 on real-time information and 9.5 on value.

DeepSeek V4.1-Flash lifts reasoning to 8.8

The 10 September release moves DeepSeek's reasoning score up a notch while the value score stays at a perfect 10. It is my answer for anyone who wants strong output with no monthly fee.

2026-09-15

ChatGPT keeps first at 9.7, and this week I spent my testing time on a question readers keep asking me: which assistant can you count on when your main one goes dark?

The September 3 disruption is now better understood. The Register reported that ChatGPT, Claude, and Grok failed in the same window. OpenAI traced its part to a routing error that began at 7:43 a.m. Pacific and was fixed by about 8:17. Anthropic described a partial infrastructure outage across Claude.ai, Claude Code, and its developer API. SpaceX said Grok's downtime came from an outage at its Memphis compute centre. Google recorded no formal Gemini incident on its status page.

Three separate causes landing together tells me the fix belongs with users: keep two assistants ready. No scores move on reliability this week.

ChatGPT stays on top for breadth. GPT-6 Astra handles deep reasoning, GPT-Live-1 powers the upgraded voice mode, and memory, file libraries, and agents all sit inside the same $20 plan. It is the single subscription that covers the most jobs.

Claude holds 9.5 as my writing and long-document assistant. Claude Fable 5.1 summarizes a dense report while keeping the author's caveats intact, which is exactly what I need before a meeting.

Gemini keeps 9.4 and earns my backup slot. The free tier runs Gemini 3.8 Flash, answers quickly, and already sits inside Gmail and Drive, so switching over during an outage costs nothing.

Grok at 9.0 remains the fastest route to what people on X are saying right now. Perplexity at 8.6 is the tool I open when every claim needs a citation. DeepSeek at 8.2 stays the best zero-budget reasoning model.

My setup recommendation for this autumn: pay for the assistant that fits your main work, and keep Gemini's free tier signed in on your phone as the spare.

No ranking changes this week.

Three causes behind the September 3 outage

OpenAI cited a routing error fixed in about 34 minutes, Anthropic a partial infrastructure fault, and SpaceX a Memphis compute centre outage for Grok.

ChatGPT covers the most jobs for $20

GPT-6 Astra, GPT-Live-1 voice, memory, file libraries, and agents share one subscription, which keeps it at 9.7.

Gemini's free tier is the ideal spare

Gemini 3.8 Flash is fast, free, and already inside Gmail and Drive, so it makes the easiest backup to keep signed in.

Claude stays the long-document pick

Claude Fable 5.1 keeps nuance and caveats intact in summaries, holding Claude at 9.5 for writing-heavy work.

2026-09-14

ChatGPT stays first at 9.7, and the past ten days handed me more reasons to keep it there. OpenAI released GPT-6 Astra in early September, then on September 10 shipped GPT-Live-1 to the API alongside a ChatGPT Voice upgrade that calls on GPT-5.6 or GPT-6 Astra whenever a spoken question needs search or reasoning. Voice, memory, file libraries, and agents in one $20 subscription is exactly the breadth my realtime and ecosystem columns reward.

Claude holds second at 9.5. Anthropic shipped Claude Fable 5.1 and Claude Mythos 5.1 on September 1, and for long documents and careful writing it remains the assistant I open first. Hand it a dense contract and ask which clauses shift liability, and it quotes the right paragraphs and flags the ambiguous wording clearly.

Gemini keeps third at 9.4. Gemini 3.8 Flash became generally available on September 2, and the value story keeps improving: fast answers, tight hooks into Gmail and Drive, and a free tier generous enough for most students.

One event matters for anyone relying on a single tool. On September 3 an Azure East US failure took ChatGPT, Claude, and Grok offline for about 90 minutes while Gemini stayed up. I did not move scores over one outage. I did change my own routine: keep a second assistant signed in on your phone, and Gemini's free tier is the obvious pick for that backup role.

Grok stays fourth at 9.0 on real-time information pulled from X. Perplexity holds fifth at 8.6 as the research tool that cites every claim it makes. DeepSeek at 8.2 remains the strongest free option for reasoning on a zero budget.

Nothing moves this week. The releases were large, and each one reinforced the product that already led its lane.

GPT-6 Astra and GPT-Live-1 widen ChatGPT's reach

OpenAI followed GPT-6 Astra with GPT-Live-1 on September 10, and ChatGPT Voice can now hand hard spoken questions to a reasoning model. That keeps ChatGPT the most complete assistant at 9.7.

Claude Fable 5.1 keeps Claude the writing pick

Anthropic released Fable 5.1 and Mythos 5.1 on September 1. For long documents, contracts, and careful drafting, Claude at 9.5 is still the assistant I trust with the text that matters.

Gemini 3.8 Flash strengthens the free tier

General availability on September 2 gives Gemini fast, capable answers at no cost, with Workspace integration that students and small offices already use every day.

Keep a backup assistant signed in

The September 3 Azure failure took three major chatbots down for roughly 90 minutes while Gemini kept working. A second free account costs nothing and saves a deadline.

2026-09-11

Two days after Apple put an on-device assistant at the centre of its September keynote, I went back through this ranking to see whether anything should move. Nothing does, and the reason is worth stating plainly. ChatGPT holds first at 9.7 because it covers more shapes of a working day than anything else shipping: memory that carries between sessions, voice that survives a walk around the block, a canvas for long document edits, and connectors into the drives where people already store their files.

Claude stays second at 9.5 and keeps the work I care most about, which is long documents and reasoning that has to hold up when someone checks it. Gemini holds 9.4 and owns Google Workspace outright. If your team lives in Docs and Sheets, that single fact outweighs any leaderboard position.

Apple's announcement changes the question underneath the category. An assistant answering on the phone itself takes the quick lookups and short dictations that used to mean opening an app. I expect every chatbot here to feel that in usage volume while the ranking order stays exactly where it is, because these assistants earn their keep on work that runs longer than a sentence.

My advice today is the same shape as last week and sharper in detail. Pick the assistant whose ecosystem already holds your documents, then judge it on tasks that take more than thirty seconds. Short prompts are heading toward whatever ships on your phone, and that frees a paid subscription to prove itself on work that actually justifies one.

The phone is absorbing the thirty-second questions

Apple's on-device assistant now answers the lookups and dictations that used to send you into an app. Every chatbot on this list will see that in usage volume, and none of it changes which one you hand a forty-page contract.

ChatGPT holds first on surface area

Cross-session memory, hands-free voice, document editing and file connectors add up to an assistant that fits more kinds of work than anything else shipping. That breadth is what keeps it at 9.7.

Where your files live decides more than model quality

Gemini inside Workspace and ChatGPT with drive connectors both remove the export step that quietly costs minutes every day. Capability at the top has converged enough that access wins the argument.

Judge candidates on the long tasks

A subscription earns itself on long documents, careful reasoning and multi-step work. Test exactly those, and let the phone keep the quick stuff.

2026-09-09

This was the busiest seven days the chatbot market has had all year, and it settled a question I have been asked constantly since spring. GPT-6 Astra shipped on September 3 and it is a genuine generational release, so ChatGPT stays at number one and earns a bump to 9.7. Astra lists at $10 per million input tokens and $50 per million output, which is 2.5 times what GPT-5.6 Sol charged on its promotional rate. That is the real story here. OpenAI is confident enough in the reasoning gap to charge like a premium vendor again.

Claude holds second, and I want to be specific about why. Anthropic shipped Fable 5.1 on September 1 and cut cache-read pricing by 75 percent in the same release. For anyone running long-context assistants over a fixed document set, that single pricing change does more for monthly cost than any benchmark gain announced this week. I raised Claude's value score to reflect it.

Gemini stays third on 9.4. The 3.8 Flash release on September 2 held the previous introductory price, which keeps Google the cheapest credible option at the high-volume tier, and Flash remains the model I reach for when latency matters more than depth.

My advice for the next month is to resist switching on launch-day excitement. Astra is the strongest reasoner available today, and it costs five times more per output token than the Flash tier that handles most everyday questions perfectly well. Pick per workload, and keep two providers wired up.

GPT-6 Astra justifies ChatGPT holding the top spot

Astra is a full generational step rather than another point release inside the GPT-5 line, and it arrived on September 3 with pricing that signals real confidence. At $10 in and $50 out per million tokens, OpenAI is charging 2.5 times the GPT-5.6 Sol promotional rate. Vendors only price that way when they believe the capability gap is defensible.

The 75 percent cache-read cut is the most valuable pricing move of the week

Anthropic shipped Fable 5.1 on September 1 alongside a 75 percent reduction in cache-read pricing. If you run an assistant over a stable corpus, cached input dominates your bill, so this lands harder on real monthly spend than any headline benchmark. It is the reason Claude keeps second place comfortably.

Gemini 3.8 Flash is still the right default for high volume

Google released 3.8 Flash on September 2 at the same introductory price as 3.7 Flash. Holding price through a capability refresh keeps Gemini the cheapest credible option for chat volume that runs into millions of tokens a day, and Flash latency remains the best in the category.

Run two providers and route by workload

The spread between the frontier tier and the fast tier is now roughly five times per output token. Sending every question to the most expensive model wastes money on tasks Flash answers correctly. Keep two accounts live, route hard reasoning to Astra or Claude, and let the fast tier absorb the rest.

2026-09-07

OpenAI has started rolling GPT-6 Astra out to a selected set of customers, and I want to be precise about what that means for a buying decision today. A staged rollout to vetted accounts is a signal about where the ceiling is heading, and it is not yet something a person paying twenty dollars a month can sit down and use. I am keeping ChatGPT at the top because its everyday surface remains the most complete package on the market: voice, memory, file handling, image generation and a mobile app that behaves the same way on every device. The Astra rollout raises my reasoning score by a tenth because the safeguards work OpenAI published alongside it tells me the model line is being hardened before it scales.

Claude stays a very close second on the strength of long-document work. I spend most of my week feeding it forty and fifty page PDFs, and it holds detail across the whole thing in a way that shows up immediately when you ask it to cite back to a specific paragraph. Gemini remains the pick for anyone who already lives inside Google Workspace, because the value there comes from the connection to Gmail, Drive and Calendar.

Grok moves up a tenth this week. xAI shipped an enterprise agent product with real access controls, network restrictions and audit logging, and that changes who can actually deploy it. Audit trails are the feature that gets a tool past a security review, and Grok now has them.

My advice for this month is to hold your subscription steady. There is a large model transition happening across all three leaders, and the sensible move is to wait for general availability before you reshuffle your workflow around any one of them.

ChatGPT keeps the crown because the product is deeper than the model

GPT-6 Astra is in a gated rollout, so it is not the reason I keep ChatGPT at number one. The reason is that voice mode, persistent memory, file analysis, image generation and the Canvas editor all sit in one app that works identically on desktop, web and phone. When I hand this to someone who has never used an AI tool, they get value in the first five minutes. That consistency is worth more day to day than a benchmark lead that lasts six weeks.

Claude is the one I reach for when the input is long

I ran the same 48 page contract through the top three this week and asked each to pull every clause touching termination rights. Claude found all of them and cited page numbers I could check. That reliability across long inputs is why it holds second place, and it is the specific reason I recommend it to lawyers, researchers and anyone whose real work arrives as a large document.

Grok earned its score bump with audit logs, not with benchmarks

xAI released an enterprise agent tier this week with granular access controls, network restrictions and audit logging. I raise Grok to 9.0 for that reason alone. Every company I have watched try to deploy an AI assistant has stalled at the security review, and the questions asked there are about who accessed what and when. Grok can now answer those questions, which moves it from an interesting model to a deployable one.

Wait for general availability before you switch subscriptions

All three leaders are mid transition right now. OpenAI is gating Astra, Anthropic shipped a 1M token context tier, and Google is rotating Gemini's model backing inside Workspace. Any ranking built on this week's gated previews will look wrong in a month. I am telling readers to keep whatever they are paying for through September and reassess when the new flagships are generally available and priced.

2026-09-05

The tie broke this week. OpenAI shipped GPT-6 Astra, staged to general availability at $10 in and $50 out per million tokens with a 1.05M context window, and it is the first model the company has flagged as reaching its critical-cyber safeguard threshold. Self-reported numbers put it at 57.7 on Terminal-Bench 4.0, 74.1 on DeepSWE and 96.0 on GPQA. That is a genuine capability step, and I am moving ChatGPT to 9.6 and first place on the strength of it. Claude drops to second at 9.5, which is a ranking change and not a criticism: Anthropic's Fable 5.1 landed September 1 with cache reads down to $0.25, and for anyone running long agentic sessions that price cut is still the most consequential release of the week. Claude keeps my recommendation for long-document reasoning and code that has to pass review. Gemini holds 9.4 after 3.8 Flash reached stable general availability on September 2, and the cheap tier is now genuinely production-ready with grounding good enough for real applications. Grok at 8.9 keeps its lead on real-time social data. Perplexity at 8.6 is still what I open when I want cited answers. My practical read for September: the safeguard flag on Astra is the story worth watching, because a frontier lab publicly gating a model on cyber capability changes what enterprise procurement asks for. Subscribe to one flagship, use free tiers elsewhere for a second opinion.

GPT-6 Astra takes first place at 9.6

Staged general availability at $10/$50 per million tokens with a 1.05M context window and self-reported 57.7 on Terminal-Bench 4.0 and 96.0 on GPQA. That combination of context, price and measured capability is what moves a ranking.

The critical-cyber safeguard flag is the real news

Astra is the first model OpenAI has publicly identified as reaching that threshold. When a frontier lab gates a release on cyber capability, enterprise procurement questionnaires change, and every competitor now has to answer the same question.

Fable 5.1 keeps Claude at 9.5 on economics

Cache reads dropping to $0.25 changes which long agentic workflows are affordable to run, since repeated context is where API bills grow. For multi-hour tool-using sessions Claude remains my recommendation.

Gemini 3.8 Flash made the cheap tier production-ready

Stable general availability on September 2 matters more than another benchmark point. A fast, inexpensive, generally available model with solid grounding is what most real deployments actually run on.

2026-09-04

The top of this list moved this week and I want to be precise about why. Anthropic shipped Claude Fable 5.1 on September 1, positioned as its strongest model for coding and knowledge work, tuned specifically for agentic tool use and paired with much cheaper cache reads. Google followed on September 2 with Gemini 3.8 Flash reaching stable general availability. Claude and ChatGPT stay tied at 9.5 and I am comfortable defending that tie, because they now win on different axes and most readers should pick based on what they do all day. Claude takes long-document reasoning, code that has to survive review, and anything where the model needs to hold a large codebase in working memory. The cheaper cache reads in the 5.1 release matter enormously for anyone running long agentic sessions, since that is where API bills quietly balloon. ChatGPT holds its 9.5 on breadth: the widest tool ecosystem, the strongest consumer app polish, and voice mode that remains the most natural of the three. Gemini at 9.4 has the best real-time grounding and the tightest integration with Workspace, and the Flash 3.8 general availability makes the cheap tier genuinely usable for production tasks. Grok at 8.9 stays where it is on real-time social data. Perplexity at 8.6 is still the one I open when I want an answer with its sources attached. My practical verdict for September: subscribe to one flagship and use the free tiers of the others for cross-checking. Paying for two is a waste at these capability levels.

Fable 5.1 makes long agentic sessions affordable

Anthropic paired the September 1 release with substantially cheaper cache reads. Anyone running multi-hour tool-using sessions knows that repeated context is where the bill grows, and cutting that cost changes which workflows are economically sensible to automate.

The Claude and ChatGPT tie is real, and the split is by task

Claude takes long-context reasoning and code that has to pass review. ChatGPT takes ecosystem breadth, app polish and voice. At 9.5 each, the right pick depends entirely on which of those you touch daily.

Gemini 3.8 Flash makes the cheap tier production-ready

Reaching stable general availability on September 2 matters more than a benchmark point. A fast, cheap, generally available model with solid grounding is what most real applications actually deploy against.

Paying for two flagship subscriptions is a waste right now

The capability gap between the top three is small enough that a second subscription buys very little. Pick the one that fits your daily work, then use free tiers elsewhere when you want a second opinion.

2026-09-01

August closed with sixteen new models from eight providers, and the shape of that flood tells you exactly where the chatbot race now sits. The frontier tier has stopped moving on raw intelligence and started moving on speed and price. Gemini 3.7 Flash landed on August 13 as the fastest frontier-class model available, Grok 4.6 arrived the day before, and Alibaba closed the month with Qwen3.8 Flash on August 26. Every one of those is a latency and cost play. I keep Claude and ChatGPT tied at the top because the work I actually give a chatbot, long reasoning chains, careful writing, code that has to run, still separates on quality, and tokens per second barely enters into it. Claude Opus 5 has held that position since late July while costing roughly half what the previous flagship tier did, and that price move matters more to a daily user than another benchmark point. Gemini stays a very close third on the strength of its real-time grounding and the Flash tier being genuinely quick enough to change how you use it. The interesting movement this month is at the bottom. Qwen Chat moves up to ninth on the back of Qwen3.8 Flash, which gives the free tier a materially faster default model than Le Chat currently offers. If you are choosing today, pick on the work you do most: Claude for sustained reasoning and writing, ChatGPT for breadth of tooling, Gemini for anything that needs live web grounding.

The frontier race has moved from intelligence to latency

Gemini 3.7 Flash on August 13 and Qwen3.8 Flash on August 26 are both speed-tier releases, and Grok 4.6 landed the day before Gemini. Three of the month's most visible launches optimize response time and cost. That tells me the top labs believe the quality gap at the frontier is now small enough that the winning move is making a very good model feel instant.

Claude Opus 5 at half the old flagship price is the value story of the quarter

Anthropic shipped Opus 5 on July 24 at roughly half the cost of the prior flagship tier for comparable intelligence. For anyone paying per token, that single change did more for practical usability than any benchmark movement in August. It is why Claude holds a 9.5 here alongside ChatGPT, with cost now counting in its favour.

Qwen Chat earns ninth place because its free tier just got faster

Qwen3.8 Flash arrived on August 26 and became the default for Alibaba's consumer chat surface. A free user now gets noticeably quicker responses there than on Le Chat, which is the whole reason I move Qwen Chat ahead this month. At the bottom of this list, response speed on the free tier is the thing people actually feel.

Pick by workload, and the top three separate cleanly

Claude wins sustained reasoning and long-form writing. ChatGPT wins breadth: images, voice, data analysis, and the widest third-party tooling. Gemini wins anything that needs live grounding in current web results. Those are three different jobs and the 9.5, 9.5, 9.4 spread reflects how close they are once you match the tool to the task.