SOURCE-LINKED INTELLIGENCE
huggingface/transformers: Release v5.16.1
# Release v5.16.1 This is a special release as we include GLM! (and a few small fixes) # GLM-5.3-Flash GLM-5.3-Flash, the first **natively multimodal model** in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest **30T-token** multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less comp
Read original source ↗ Open in workspace
- recordType
- software-release
- evidenceStatus
- publisher-reported
- region
- Global
Evidence & attribution
- GitHub AI software releases · 2026-08-26T14:50:01.000Z
First collected: 2026-09-21T01:21:57.065Z. This is not the publication date.