LM Studio 0.4.14 Adds Stable MTP Speculative Decoding for Local LLMs

LM Studio 0.4.14 introduces stable MTP Speculative Decoding to accelerate local large language models while maintaining quality, supporting Qwen3.6 and GGUF formats.

LM Studio 0.4.14 Adds Stable MTP Speculative Decoding for Local LLMs

Element Labs has released LM Studio 0.4.14 (Build 4), introducing a stable version of MTP Speculative Decoding for local large language model acceleration.

Element Labs releases patch with stable multi-token prediction support.

The update enables the software to predict future tokens using a lightweight model and verify them on the target model, which speeds up output while maintaining quality.

Users can currently run this feature with Qwen3.6-35B-A3B-MTP-GGUF or Qwen3.6-27B-MTP-GGUF models, alongside standard GGUF and llama.cpp formats.

The release also addresses several bugs, including non-MTP speculative decoding errors that occurred when MTP was enabled, and a display issue with the 'lms get gemma4' command.

Additionally, the 'lms chat' command now identifies which LM Link device hosts each remote model during use.

Element Labs has not provided pricing or availability details for this software update.

Discussion

0 comments

Log in to join the thread with a thoughtful take, question, or correction.

Add to the discussion