gh-llama-cpp Releases v0.5.0 with Performance Improvements and New Features
gh-llama-cpp has released version 0.5.0, focusing on backend performance, correctness, and broader model coverage.
gh-llama-cpp has released version 0.5.0, which includes several performance and feature enhancements. The release focuses on backend performance and correctness, broader model coverage, and more robust server/router operation. Key highlights include the addition of HRM-Text (DFM Mimir 1B) support, MiMo-V2.6 and HunyuanOCR conversion support, and ggml 0.25.0 backend improvements. Other notable features are multi-address HTTP binding and image outputs from function calls. Additionally, the release addresses several chat parser/UI fixes to enhance user experience. The updated ggml to version 0.25.0 expands support for hyper-connection, flash-attention, and fused MoE/SSM across backends, with improvements in robustness, quantization, data-layout, and RPC/meta functionalities.
Source: gh-llama-cpp

