<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Posts on minamism</title><link>https://minamism.com/posts/</link><description>Recent content in Posts on minamism</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Thu, 10 Sep 2026 06:00:00 +0900</lastBuildDate><atom:link href="https://minamism.com/posts/index.xml" rel="self" type="application/rss+xml"/><item><title>3.14M Context on 3x 3090: The Power of an Architecture That Puts KV Cache in Host RAM</title><link>https://minamism.com/posts/qsa-kv-offload/</link><pubDate>Thu, 10 Sep 2026 06:00:00 +0900</pubDate><guid>https://minamism.com/posts/qsa-kv-offload/</guid><description>Qwen Sparse Attention reads the same number of tokens per step no matter how long the context is. Thanks to this, the KV cache can be placed in host RAM. With the same three GPUs, the upper limit went from 234k to 3.14M tokens.</description></item></channel></rss>