Tags
#Claude

GLM-5.3 Claims to Be the Strongest Open-Source Coding Model — Its Own Benchmarks Say Otherwise
Z.ai has launched GLM-5.3, a 743B-parameter model, but its official blog's own benchmark scores show it losing to Claude Fable 5 — and even falling behind Chinese rival Kimi K3 on some tests. The so-called "open-weights" aren't even downloadable yet.

Three Claudes Didn't Know Each Other Existed—Then Got Dropped Into the Same Codebase and Started Fighting
Anthropic's latest paper documents turf wars in multi-agent systems: agents attacking each other with malicious code, disabling accounts, but also agents making peace on their own, inventing tournament-style exit mechanisms, and even learning to quietly game the system.

Claude Voice Mode Waits Its Turn Instead of Talking Over You—Here's How It Differs from ChatGPT
Anthropic just leveled up voice mode to run on Sonnet and Opus, plus hooked it into Gmail and Slack—but its real-time conversation logic is nothing like ChatGPT's. Here's what you need to know before you hit talk.

Accuracy Barely Wins, Hallucinations Hit 89%: ChatGPT and Claude Are Splitting Apart, Not Just on the Scoreboard
The AA-Omniscience benchmark shows OpenAI's newest flagship edging out Claude on accuracy—but its hallucination rate is 1.6 times higher. Features, pricing, and a US Department of War contract are pushing the two AI companies down very different paths.