<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Triton on minamism</title><link>https://minamism.com/tags/triton/</link><description>Recent content in Triton on minamism</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Fri, 25 Sep 2026 05:00:00 +0900</lastBuildDate><atom:link href="https://minamism.com/tags/triton/index.xml" rel="self" type="application/rss+xml"/><item><title>Running 125B MoE on Three 3090s - vLLM Optimization Edition</title><link>https://minamism.com/posts/flash-next-vllm-optimization/</link><pubDate>Fri, 25 Sep 2026 05:00:00 +0900</pubDate><guid>https://minamism.com/posts/flash-next-vllm-optimization/</guid><description>Qwen3.8-Flash-Next ran at 80 tok/s last time. Once it went into real use, first-request TTFT was 100 seconds, prefix caching cut the context by 30%, and decode in the production configuration was in the 60s, not 80. This covers the 11 patches added to reach 95-99 tok/s single-stream, 243 tok/s with four streams, image input and MTP, and the approaches that were tested and not adopted.</description></item></channel></rss>