Moore Threads MUSA Backend Merged into SGLang Mainline at Beijing Meetup

Moore Threads MUSA backend merges into SGLang mainline. 41 PRs merged, supports DeepSeek, Qwen, GLM. Meetup in Beijing highlights ecosystem growth.

Moore Threads MUSA Backend Merged into SGLang Mainline at Beijing Meetup

Moore Threads has merged its MUSA backend into the mainline of SGLang, an open-source inference engine. The company held a meetup in Beijing on May 10, 2025, to discuss the integration and the broader GPU software ecosystem.

MUSA backend now native in SGLang

Moore Threads submitted 47 pull requests to SGLang, with 41 already merged. The MUSA backend now supports models including DeepSeek, Qwen, GLM, Wan, and LTX. The MATE operator library provides high-performance Attention and GEMM operators compatible with FlashAttention, FlashMLA, and DeepGEMM. The torchada adapter lets CUDA code run on Moore Threads GPUs with a single import. TileLang integration is also underway, with a maintainer reporting a 20x speedup on Attention Sinks kernels using about 50 lines of code.

Moore Threads MUSA backend integration with SGLang meetup
Attendees at the SGLang × MUSA Meetup in Beijing on May 10, 2025.

The Mooncake project, which includes Moore Threads as a core maintainer, reduced P2P transfer time from 53 seconds to 7.2 seconds, according to a contributor. BAAI researchers reported that FlagOS achieved a 4x speedup on Fused MoE and FP8 GEMM through fusion and quantization. For DeepSeek-V4 Day0 adaptation, they noted a 56.7% reduction in TTFT and a 65.7% increase in throughput.

Moore Threads CTO Zhang Yubo said the company aims to integrate into existing ecosystems with zero learning cost rather than creating a closed one. The MUSA interface reuses familiar GPU programming habits. The company has not confirmed pricing or availability for the MTT S5000 or related hardware.

Moore Threads CTO Zhang Yubo speaking at meetup
Zhang Yubo discusses MUSA's zero learning cost integration strategy.

Discussion

0 comments

Log in to join the thread with a thoughtful take, question, or correction.

Add to the discussion