← All reports

Report · 2026-10-06

Qwen3-30B-A3B FP8 on 1× H100: BF16 vs FP8 KV cache

1× NVIDIA H100 80GB HBM3 Qwen3-30B-A3B-FP8 FP8 weights; KV cache varies by configuration vLLM 0.31.0 TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms

Validation run on a Lambda Cloud 1× H100 SXM5 instance with vLLM 0.31.0. Quick mode: 120 s windows, 2 repeats, 50 GSM8K items with one seed. FP8 KV cache raised SLO goodput and tokens per joule but failed the accuracy gate on this small suite.

ConclusionUnder SLO (TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms), serving the same billable throughput as 64 cards on vllm-default takes: vllm-fp8kv failed the accuracy gate (-8.0 pts) and is excluded from capacity math, despite 5,057 tok/s goodput.

Results

Each value is the median of 2 runs.

vllm-defaultvllm-fp8kv

SLO goodput

output tok/s per node · higher is better
vllm-default4,572vllm-default: 4,572 output tok/s per node
vllm-fp8kv5,057vllm-fp8kv: 5,057 output tok/s per node

Tokens per joule

tok/J, GPU board power · higher is better
vllm-default7.18vllm-default: 7.18 tok/J, GPU board power
vllm-fp8kv8.21vllm-fp8kv: 8.21 tok/J, GPU board power

Decode bandwidth utilization

% of rated HBM bandwidth
vllm-default57%vllm-default: 57 % of rated HBM bandwidth
vllm-fp8kv45%vllm-fp8kv: 45 % of rated HBM bandwidth

Axis runs to 100% of rated bandwidth.

GPU board power

kW, steady state
vllm-default0.64vllm-default: 0.64 kW, steady state
vllm-fp8kv0.62vllm-fp8kv: 0.62 kW, steady state

GPU board power only — excludes CPUs, memory, NICs, fans and PSU losses.

All numbers

Cards for equal goodput = 64 × reference goodput ÷ configuration goodput, rounded up.

Metricvllm-defaultvllm-fp8kv
Engine buildvLLM 0.31.0 · KV BF16vLLM 0.31.0 · KV FP8
Decode bandwidth utilization (MBU)57%45%
GPU board power0.64 kW0.62 kW
SLO goodput4,572 tok/s5,057 tok/s
Tokens per joule7.188.21
TTFT p99340 ms352 ms
TPOT p9944.6 ms43.7 ms
SLO metyesyes
Accuracy (gsm8k-5shot)96.088.0
Accuracy Δ vs. referencereference-8.0 pts
Cards for equal goodput64—

Setup

Any change to these fields makes it a different report.

Accelerator
1× NVIDIA H100 80GB HBM3
Rated HBM BW
3.35 TB/s per GPU
Driver / runtime
580.105.08
Host
Lambda Cloud instance
Engine
vLLM 0.31.0
Kernel libraries
vllm/vllm-openai:latest
Model
Qwen3-30B-A3B-FP8 (MoE)
Quantization
FP8 weights; KV cache varies by configuration
Parallelism
1 replica × TP1
Workload trace
chat-mix-v1
Input tokens
median 1,024 · p90 4,096
Output tokens
median 256 · p90 1,024
Arrival
Poisson, rate swept to SLO boundary
Power source
GPU board power sum (NVIDIA) · 1 Hz · 120 s steady state
Runs per config
2 (variance band ±1.0%)
Accuracy gate
gsm8k-5shot · max drop 1.0 pts

Not covered

Do not extrapolate this report to the following.

  • Whole-node power (cloud VM: no BMC access)
  • Context lengths beyond the chat-mix-v1 trace
  • Quick mode: short windows, few repeats — validation only, not for publishing
  • Whole-node power: energy figures are GPU board power only (no CPUs, memory, NICs, fans or PSU losses)

Reproduce

Run on the same hardware and software versions. Results should fall within ±1.0%.

shell
tokenwatt run --spec lambda-qwen3-30b-a3b-fp8-tp1-dp1.spec.yaml --label 'vllm-default' --role baseline --out runs/vllm-default.json
tokenwatt run --spec lambda-qwen3-30b-a3b-fp8-tp1-dp1.spec.yaml --label 'vllm-fp8kv' --role tuned --out runs/vllm-fp8kv.json

tokenwatt report runs/vllm-default.json runs/vllm-fp8kv.json --title 'Qwen3-30B-A3B FP8 on 1× H100: BF16 vs FP8 KV cache'

Download raw JSON