Skip to content
Intelligence RadarNews & Launches1 MIN READ

Hugging Face Announces LFM2.5-VL-DSpark for Vision-Language Models

Hugging Face has released an experimental DSpark draft model for their vision-language model LFM2.5-VL-3B, offering significant speedup in inference without compromising output quality.

Fathom Intelligence
Fathom IntelligenceFathom Layer Expert

Hugging Face today announced the release of an experimental DSpark draft model for their vision-language model (VLM) LFM2.5-VL-3B. This new model includes a speculative decoding path that enhances inference speed without affecting output quality, trading a minimal increase in memory footprint for substantial performance gains.

The DSpark draft model for LFM2.5-VL-3B is designed to work seamlessly with existing frameworks such as llama.cpp, MLX-VLM, and SGLang. Performance benchmarks conducted on six diverse vision-based tasks, following the MMSpec benchmark, reveal significant improvements:

  • On-device inference: With MLX on an M5 Max, decoding runs 2.30x to 3.13x faster, with end-to-end latency improvements of 1.56x to 2.62x. On an M3 Ultra with llama.cpp, decoding speeds up by 1.57x to 2.14x, and end-to-end performance by 1.30x to 1.77x.

  • GPU inference: On an H100 GPU, the model delivers decoding speeds 20.4x to 2.66x faster, with end-to-end improvements of 1.64x to 2.27x.

These enhancements aim to accelerate the deployment and use of vision-language models in various applications, making them more accessible and efficient for developers and researchers alike.


Source: huggingface

END OF SIGNAL

Don't Miss the Next Signal

Get high-impact hardware and AI launches distilled into your inbox. No noise, just the changes that matter.