
OpenAI’s first custom inference chip reportedly beats Nvidia Blackwell and Rubin on throughput and efficiency metrics.
Z.ai’s GLM-5.3-Flash is a 320-billion-parameter open-source model that nearly matches GLM-5.3 on Artificial Analysis’s Intelligence Index while costing far less and running inference on Chinese AI chips.

Z.ai has released GLM-5.3-Flash, an open-source model with 320 billion total parameters and 18 billion active parameters. The model is described as the first natively multimodal model in the GLM-5 series, ships under an MIT license, offers a one-million-token context window, and has weights available on Hugging Face.
For readers tracking model availability, the key takeaway is that GLM-5.3-Flash is positioned as a high-capability, lower-cost alternative rather than only a smaller technical demo.

Artificial Analysis measurements put GLM-5.3-Flash at 57 points on its Intelligence Index at maximum reasoning effort. That is three points behind the larger GLM-5.3, which scores 60, and level with GPT-5.6 Terra and Muse Spark 1.2.
The cost difference is the headline: GLM-5.3-Flash is listed at 0.09 dollars per task versus 0.68 dollars for GLM-5.3, making it roughly 7.5 times cheaper. On Z.ai’s API, pricing is 0.15 dollars per million input tokens and 0.50 dollars per million output tokens.
Z.ai says all inference traffic for the model’s pre-launch anonymous testing as “ox-alpha” on OpenCode and OpenRouter ran on Chinese AI chips. The model became the most popular model of the week during that test, according to the article.
This matters because much of the AI software ecosystem has long been optimized around Nvidia’s CUDA stack. Z.ai built serving software on top of SGLang, split processing into independently scalable stages, and says the work tripled throughput over its first attempt on the same hardware.
GLM-5.3-Flash appears strongest as a cost-performance story: it comes close to the larger GLM-5.3 on benchmark scoring while lowering task cost and using non-Nvidia inference hardware. It also keeps pace with its larger sibling on agentic tasks, according to the article’s summary of GDPval-AA v2 results.
The article also flags a weakness: Artificial Analysis found the model was less token-efficient, with roughly 90 percent of output tokens used for reasoning. For developers and buyers, the practical takeaway is to compare total workload cost, not just headline token pricing or benchmark rank.

OpenAI’s first custom inference chip reportedly beats Nvidia Blackwell and Rubin on throughput and efficiency metrics.

GPT-5.6-Cyber is aimed at giving defenders earlier access to AI-assisted vulnerability research.

AMD is adding Taalas’ model-in-silicon inference technology to its AI accelerator roadmap.

Nova models stay online for existing customers, but active development reportedly shifts to frontier model research.