One line: zai-org/GLM-5.3-Flash (FP8, 320B/18B-active MoE) served by SGLang on 8× Instinct MI325X in Docker, tensor-parallel 8, EAGLE speculative decoding on the NextN layer, at ~5–6× the upstream AMD recipe (~110–123 tok/s single-stream vs 18.9) after five local patch families — all reproduced below.
| GPUs | 8× AMD Instinct MI325X (gfx942), 256 GB HBM each, 304 CUs, 64 KB LDS/workgroup |
| Host | AMD EPYC, 256 logical cores, 3 TiB RAM |