<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Cuda on minamism</title><link>https://minamism.com/ja/tags/cuda/</link><description>Recent content in Cuda on minamism</description><generator>Hugo</generator><language>ja-JP</language><lastBuildDate>Wed, 30 Sep 2026 01:00:00 +0900</lastBuildDate><atom:link href="https://minamism.com/ja/tags/cuda/index.xml" rel="self" type="application/rss+xml"/><item><title>FP8 / NVFP4 の weight-only 推論で RTX 3090 が失っているものは無い</title><link>https://minamism.com/ja/posts/weight-only-fp8-nvfp4-3090/</link><pubDate>Wed, 30 Sep 2026 01:00:00 +0900</pubDate><guid>https://minamism.com/ja/posts/weight-only-fp8-nvfp4-3090/</guid><description>FP8 / FP4 のテンソルコアが使われるのは、アクティベーションも FP8 / FP4 に量子化したときだけである。weight-only (A16) では、RTX 3090 も Blackwell も BF16 のテンソルコアで計算する。vLLM のコード、5090 用にビルドされた SASS、3090 での実測 (カーネル時間・クロック固定・Nsight Compute) で根拠を示し、3090 が実際に失っているもの (W8A8 / W4A4 の prefill の演算上限) と、想定される反論への回答をまとめる。</description></item></channel></rss>