AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

ggml-org/llama.cpp: v0.5.0

GitHub AI software releases · article · Sep 23, 2026 · UTC

## Overview This release focuses on backend performance and correctness, broader model coverage, and more robust server/router operation. It adds HRM-Text (DFM Mimir 1B) support, MiMo-V2.6 and HunyuanOCR conversion support, ggml 0.25.0 backend improvements, multi-address HTTP binding, image outputs from function calls, and several chat parser/UI fixes. ### Highlights - Accelerate CUDA `conv2d` with implicit GEMM ([#29135](https://github.com/ggml-org/llama.cpp/pull/29135)) - Add Metal MoE and SSM_CONV fusion optimizations ([#28948](https://github.com/ggml-org/llama.cpp/pull/28948)) - Allow the server to bind to multiple addresses ([#28690](https://github.com/ggml-org/llama.cpp/pull/28690)) ### API changes - Add `llama_adapter_lora_init_from_file_ptr()` for loading LoRA from an open FILE ([#28993](https://github.com/ggml-org/llama.cpp/pull/28993)) - Document `llama_model_load_from_file_ptr()` as reading from the current position and requiring aligned mmap ([#28993](https://github.com/ggm

Read original source ↗ Open in workspace

recordType
software-release
evidenceStatus
publisher-reported
region
Global

Evidence & attribution

First collected: 2026-09-23T21:32:19.581Z. This is not the publication date.