Skip to content
Tech News
← Back to articles

DeepSeek V4 Flash on a Single AMD MI300X

read original more articles
Why This Matters

This article highlights the successful deployment of DeepSeek V4 Flash on a single AMD MI300X GPU, demonstrating its capability to handle large language models efficiently with optimized configurations. The achievement underscores the potential of AMD's high-memory, high-bandwidth hardware for AI workloads, offering a cost-effective alternative to NVIDIA solutions for large-scale AI inference. It also provides valuable insights into necessary fixes and tuning for AMD hardware to run advanced AI models reliably.

Key Takeaways

DeepSeek V4 Flash on a single AMD MI300X

This repository contains the configuration and patches I use to run deepseek-ai/DeepSeek-V4-Flash-0731 on one AMD MI300X in production. It includes the Docker Compose stack, SHA-256-pinned file overlays, reference diffs against upstream, and tuning tables. The checkpoint runs as shipped, without additional weight quantization or offload.

Results from the pinned stack (vLLM ROCm nightly 0.26.1rc1.dev229+g124154a88.rocm723 , AITER 0.1.19 ):

Metric Result Single-stream decode (median per-stream, DSpark-7) 168.6 tok/s Prefill with tuned kernels ≈ 7.9–8.5K tok/s (6,988–7,019 tok/s on fresh prompts in the shipping profile) 8 concurrent streams 542 tok/s aggregate, 90.3 tok/s median per stream 64-stream burst 830 tok/s aggregate, no OOM, no engine errors Context 256K validated (the architecture supports 1M) Weights in HBM 156.67 GiB — no additional quantization or weight offload

The official vLLM recipe targets NVIDIA and newer AMD hardware. Running the model reliably on MI300X required fixes for its FP8 format, MoE routing at high concurrency, causal speculative verification, CPU-KV synchronization, and several untuned kernel shapes. This repository collects those fixes and pins the versions used in production.

Why MI300X

The MI300X has 192 GB of HBM3 and 5.3 TB/s of memory bandwidth, with 2.4× the HBM capacity of an H100 SXM5 (AMD). Doubleword's write-up estimates that it costs roughly half as much at list price. For this 304B-parameter checkpoint, the memory capacity allows a simple single-GPU deployment:

The entire model fits in HBM without PCIe weight streaming or layer offload.

There is room for a 20 GB GPU KV pool and a 96 GiB CPU tier for evicted prefix-cache entries.

One card handles 2–8 typical concurrent streams and bursts of up to 64 streams.

... continue reading