Start with a contradiction. On the AA-Omniscience Hallucination Rate benchmark, OpenAI's newest flagship model, GPT 5.6 Sol (Max), hallucinates 89% of the time—way up from the older GPT-4o's 38%. On the same benchmark, GPT 5.6 Sol's knowledge accuracy actually improved over its predecessor. In other words, the model "knows more," but when it hits a question it doesn't know, it's also more likely to make something up rather than admit ignorance.

For comparison, Anthropic's Claude Fable 5 (Max) posts a 55% hallucination rate and 61% accuracy—both better than GPT 5.6 Sol's 59% accuracy. The gap is even more pronounced at the mid-tier: Claude Sonnet 5 hits 38% accuracy with a 37% hallucination rate, while ChatGPT 5.6 Terra leads on accuracy at 46% but hallucinates a staggering 85% of the time. In short, the two companies are roughly tied on "getting more answers right," but the moment a question stumps them, ChatGPT's models are clearly far more prone to just winging it.

The feature split is just as clear. Claude's Cowork, launched in January, can handle knowledge-based tasks on your behalf, like organizing files. Users can build instruction sets called skills, callable anytime mid-conversation with a "/" — and these skills work across Claude chat, Cowork, and Claude Code. Claude Artifacts can render code snippets, single-page HTML sites, React components, and charts on the fly, and whatever you build can be published and shared directly, while Model Context Protocol connectors let it pull in real-time data. ChatGPT has skills too, but they're limited to Codex and personal accounts in the API—not available in the regular chat window. Its answer to Artifacts is called Sites, aimed at enterprise internal use, also confined to Codex, and it can't pull live data either.

On the flip side, ChatGPT's voice experience is noticeably more mature—natural-sounding, and you can interrupt it mid-sentence without it losing track of the conversation. Paired with Advanced Voice Mode's real-time video, you can just point your camera at whatever's in front of you and ask about it directly. Image generation is another ChatGPT strength, capable of near-photorealistic output; Claude, by contrast, can only draw charts and icons using HTML and SVG.

Two User Bases, Two Ways of Working

Anthropic's Economic Index report, released in March, found that 42% of Claude conversations are personal, 45% work-related, and the rest schoolwork. OpenAI's own report, meanwhile, says 70% of ChatGPT usage has nothing to do with work. That split largely explains the diverging product strategies—one company leaning into enterprise collaboration and coding tools, the other doubling down on everyday interactions like voice and image.

An event in February accelerated some users' switch to the other side: contract negotiations between Anthropic and the US Department of War collapsed after the company refused to let its models be used for domestic mass surveillance and fully autonomous weapons systems, and Anthropic was subsequently flagged as a supply-chain risk. Hours later, OpenAI announced it had signed a similar contract with the US government, with some restrictions attached. The episode led many to view Anthropic as the "more principled" player, and Claude's download numbers saw a noticeable bump afterward.

The free-tier experience is diverging too. OpenAI has started running ads on its free tier and the ChatGPT Go plan; Anthropic, meanwhile, added more features to the free version of Claude—with no ads.

Pricing structures are roughly symmetrical. Claude's free tier is limited to the mid-tier Sonnet 5 model. Paid tiers run $20/month (or $17 billed annually) for Pro, $100 for Max 5x (five times Pro's usage), and $200 for Max 20x (twenty times the usage). Pro users pay per token to use Fable 5, while Max plans can allocate half their weekly usage to Fable 5, with the option to buy more once that's used up. ChatGPT similarly offers $20, $100, and $200 tiers, plus an additional $8 plan with higher usage limits but still carrying ads. Codex is free to use, but its usage cap is so low it's practically symbolic.

新模型分數更高,胡謅機率卻沖上89%:Claude與ChatGPT怎麼分道揚鑣
新模型分數更高,胡謅機率卻沖上89%:Claude與ChatGPT怎麼分道揚鑣