<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Vllm on minamism</title><link>https://minamism.com/tags/vllm/</link><description>Recent content in Vllm on minamism</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Fri, 11 Sep 2026 01:00:00 +0900</lastBuildDate><atom:link href="https://minamism.com/tags/vllm/index.xml" rel="self" type="application/rss+xml"/><item><title>vLLM Can Run INT5–7</title><link>https://minamism.com/posts/vllm-int5-7-quantization/</link><pubDate>Fri, 11 Sep 2026 01:00:00 +0900</pubDate><guid>https://minamism.com/posts/vllm-int5-7-quantization/</guid><description>vLLM can natively execute INT5/6/7 weights. Here is an overview of how well they perform, along with results showing that quantizing embed_tokens and lm_head causes minimal loss.</description></item></channel></rss>