AMD Launches vLLM-ATOM Plugin to Accelerate AI Inference for DeepSeek, Kimi

AMD launches vLLM-ATOM plugin to accelerate AI inference on Instinct GPUs. Supports DeepSeek, Kimi, Qwen3, and gpt-oss with zero learning cost for migration.

AMD Launches vLLM-ATOM Plugin to Accelerate AI Inference for DeepSeek, Kimi

released the vLLM-ATOM plugin on May 11, 2024, to optimize artificial intelligence inference on its Instinct GPU hardware. The software tool targets the growing demand for efficient large language model processing on AMD accelerators. It aims to bridge the gap between popular open-source frameworks and AMD-specific hardware capabilities.

Architecture uses vLLM for scheduling and ATOM for routing

The plugin architecture consists of three distinct layers. The top layer utilizes vLLM for scheduling and key-value cache management. The middle layer features the ATOM plugin for routing and plugin integration. The bottom layer relies on AITER kernels for low-level computation optimization.

AMD vLLM-ATOM plugin architecture showing vLLM, ATOM, and AITER layers
The tool utilizes a three-layer design to optimize inference.

vLLM-ATOM supports a range of prominent large language models. Verified model versions include Qwen3-235B-A22B-Instruct-2507-FP8 and DeepSeek-R1-0528. The software also handles openai/gpt-oss-120b and amd/Kimi-K2.5-MXFP4. It is designed to work with MoE, hybrid MoE, dense models, and vision-language model scenarios.

According to Wccftech, the plugin improves inference performance for models like DeepSeek-R1, Kimi-K2, and gpt-oss-120B. AMD states that the solution offers zero learning cost for deployment migration. Users can theoretically move existing vLLM-based service flows to AMD backends without changing commands, APIs, or workflows.

Discussion

0 comments

Log in to join the thread with a thoughtful take, question, or correction.

Add to the discussion