Posts
- 3.14M Context on 3x 3090: The Power of an Architecture That Puts KV Cache in Host RAM
Qwen Sparse Attention reads the same number of tokens per step no matter how long the context is. Thanks to this, the KV cache can be placed in host RAM. With the same three GPUs, the upper limit went from 234k to 3.14M tokens.