Aug 15, 2026 · 2 min read
GLM-5.3 and the scoreboard
GLM-5.3 tops CyberGym and jumps on Terminal-Bench 3.0. ExploitBench and a missing vision stack are the reminder that a leaderboard is not a hiring manager.

Field note
Z.ai released GLM-5.3 on 14 August. Same base as GLM-5.2. The first line of the announcement is “Scaling post-training is all we did.” Weights follow in two weeks.
On CyberGym it posts 84.5%, ahead of Fable 5 at 83.8% and GPT-5.6 Sol at 83.6%. Terminal-Bench 3.0 jumped from 4.6 to 28.3. On their private Code Bench, High effort, they report 31.4% at about 50K output tokens, against Claude Opus 4.8 at 29.5% with 120K. Fable 5 is still ahead at Max effort.

A table is not a hiring manager
ExploitBench tells the other half. GLM-5.3 is 54.4% there. Fable 5 is 78.0%. Sol is 76.5%. Discovery is the headline. Exploitation is still behind. The model is also text-only.
vidIQ still shows “Claude vs ChatGPT” as the search people type. The benches this week are a four-way argument between Sol, Fable, Grok, and a ~750B Chinese lab model with no vision.
I do not pick a model off a table. I pick the one that survives a production UI: tool calls that do not wander, copy that matches the product, and a loop I can afford. GLM-5.3 looks like a coding agent. It is not a design partner. That is fine, if you know which job you hired it for.