Triton
- Running 125B MoE on Three 3090s - vLLM Optimization Edition
Qwen3.8-Flash-Next ran at 80 tok/s last time. Once it went into real use, first-request TTFT was 100 seconds, prefix caching cut the context by 30%, and decode in the production configuration was in the 60s, not 80. This covers the 11 patches added to reach 95-99 tok/s single-stream, 243 tok/s with four streams, image input and MTP, and the approaches that were tested and not adopted.