Quantization
- vLLM Can Run INT5–7
vLLM can natively execute INT5/6/7 weights. Here is an overview of how well they perform, along with results showing that quantizing embed_tokens and lm_head causes minimal loss.
vLLM can natively execute INT5/6/7 weights. Here is an overview of how well they perform, along with results showing that quantizing embed_tokens and lm_head causes minimal loss.