Cuda
- Your 3090 Isn't Missing Anything for FP8/NVFP4 Weight-Only Inference
FP8/FP4 tensor cores are used only when the activations are quantized to FP8/FP4 as well. With weight-only (A16) quantization, both an RTX 3090 and Blackwell compute on BF16 tensor cores. This post shows the evidence in the vLLM source, in the SASS built for the 5090, and in measurements on a 3090 (kernel time, locked clocks, Nsight Compute). It states what the 3090 actually lacks (the prefill compute ceiling of W8A8/W4A4) and answers the objections I expect.