The Information Machine
Updated today·following since 20 Aug 2026·Day 7·13 sources

Z.ai GLM-5.3 versus frontier coding models

The gist

Zhipu's 320B GLM-5.3-Flash runs 17x cheaper than full GLM-5.3 on coding tasks

Multiple independent benchmarks show a 320B open-weight model running on Chinese chips matching or beating closed frontier models on coding at a fraction of the cost, with detailed cost-quality tradeoff data now available across both the full model and its Flash variant. Bloomberg Intelligence has flagged Z.ai's unit economics as unsustainable given rising inference costs.

The full picture

Together AI's 900-rollout DeepSWE test finds GLM-5.3-Flash costs 17x less than full GLM-5.3, with Flash giving up 5.6 pass@1 points that narrow to 2.6 at pass@4. A Blender architectural task via MCP showed comparable output from Flash for 16.7x less ($0.0526 versus $0.8807). Z.ai disclosed on August 26, 2026 that GLM-5.3-Flash is the model it had been previewing as Ox Alpha: a 320B-parameter MoE with 18 billion active parameters, MIT-licensed, natively multimodal, with a 1M-token context window, running on domestic Chinese accelerators. The full GLM-5.3, released August 14, 2026, scored 91.25% on KingBench 3, outperforming Opus 5 and Kimi K3, tied Fable 5 on pass@1 in DeepSWE testing at $3.99 per rollout versus $21.63, and trailed GPT-5.6 Sol by 3.7 pass@1 points while winning on pass@4 at roughly half the cost. Z.ai's own benchmarks show GLM-5.3 outperformed models from Anthropic and OpenAI on a cyber performance test. Bloomberg Intelligence called Z.ai's commercial footing unsustainable; GLM-5.3's release coincided with a 9% drop in Z.ai shares.

How it developed
29 August 2026

Three independent comparisons published August 28 put Z.ai's GLM-5.3-Flash at roughly 17x less cost than the full GLM-5.3, with the pass@1 gap narrowing from 5.6 to 2.6 points at pass@4 in a 900-rollout DeepSWE test and 16.7x less on a Blender MCP task.

A three-model price test placed Flash at $0.03 versus $0.29 for Claude Opus 4.6 on the same prompt. Rohan Paul found that wrapping full GLM-5.3 in the Atomic Agent harness roughly doubled token usage while adding only $0.77 to cost.

28 August 2026

Blender MCP task: Flash produces comparable architectural output for 16.7x less than full GLM-5.3 with better prompt following

GLM-5.3-Flash became available on Databricks on August 27, adding a second major cloud platform alongside CoreWeave for Z.ai's 320B-parameter mixture-of-experts model. Databricks reports it delivers 10% higher quality than GLM-5.2 at one-tenth the cost on OfficeQA Pro v2 and adds multimodal support for reading figures and verifying web pages. A Hugging Face analysis the same day found Chinese labs are releasing open-weight models several times larger than any open models US labs have shipped publicly in 2025-2026.

27 August 2026

GLM-5.3-Flash launches on Databricks, delivering 10% higher quality than GLM-5.2 at one-tenth the cost on OfficeQA Pro v2

Together AI published 904-rollout DeepSWE comparisons August 27 showing Z.ai's GLM-5.3, released August 14, tied Fable 5 on pass@1 and beat it on pass@4 at 5.4x lower cost; KingBench 3 placed it at 91.25%, ahead of Opus 5 and Kimi K3. Bloomberg Intelligence analyst Robert Lea said Z.ai 'remains on a completely unsustainable commercial footing,' with shares down 9% on the launch day, and founder Tang Jie argued post-training is now the highest-value scaling dimension.

26 August 2026

CoreWeave announces GLM-5.3-Flash coming to its Serverless Inference platform

21 August 2026

Together AI: GLM-5.3 trails GPT-5.6 Sol by 3.7 pass@1 points but wins pass@4 at half the cost; GLM-first cascade hits 85.9%

20 August 2026

Tang Jie argues parameter count is incomplete as a capability metric and frames GLM-5.3 as a controlled post-training experiment

17 August 2026

Hugging Face analysis finds Chinese labs releasing open-weight models several times larger than any US labs have shipped publicly in 2025-2026

Sources
8 more sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free