Mid-size MoE decode on an 8-GPU node
Format example. Three configurations of the same node, model and SLO: the production image as deployed, the vendor's latest recommended image, and a tuned decode path.
ConclusionUnder SLO (TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms), serving the same billable throughput as 64 cards on Production takes: with Vendor latest, 64 → 47 cards and 21% less node energy per token; with Tuned, 64 → 35 cards and 39% less node energy per token.
Results
Each value is the median of 5 runs.
ProductionVendor latestTuned
SLO goodput
output tok/s per node · higher is better
Tokens per joule
tok/J at the node · higher is better
Decode bandwidth utilization
% of rated HBM bandwidth
Axis runs to 100% of rated bandwidth.
Node power
kW, steady state
Higher power is expected when the node does more work; tokens per joule is the efficiency measure.
All numbers
Cards for equal goodput = 64 × reference goodput ÷ configuration goodput, rounded up.
| Metric | Production | Vendor latest | Tuned |
|---|---|---|---|
| Engine build | — | — | — |
| Decode bandwidth utilization (MBU) | 31% | 44% | 63% |
| Node power | 6.10 kW | 6.60 kW | 7.00 kW |
| SLO goodput | 2,400 tok/s | 3,300 tok/s | 4,500 tok/s |
| Tokens per joule | 0.39 | 0.50 | 0.64 |
| TTFT p99 | 1,710 ms | 1,650 ms | 1,590 ms |
| TPOT p99 | 47.0 ms | 46.0 ms | 48.0 ms |
| SLO met | yes | yes | yes |
| Accuracy (gsm8k-500) | 88.4 | 88.3 | 88.2 |
| Accuracy Δ vs. reference | reference | -0.1 pts | -0.2 pts |
| Cards for equal goodput | 64 | 47 | 35 |
Setup
Any change to these fields makes it a different report.
- Accelerator
- 8× Accelerator A (example)
- Rated HBM BW
- 5.0 TB/s per GPU
- Driver / runtime
- example-driver 1.2.3
- Host
- 2-socket server, 8 accelerators
- Engine
- vLLM 0.x (example)
- Kernel libraries
- vendor kernel library (example)
- Model
- Example-MoE-30B-A3B (MoE)
- Quantization
- FP8 weights, FP8 KV cache
- Parallelism
- 8 replicas × TP1
- Workload trace
- chat-mix-v1
- Input tokens
- median 1,024 · p90 4,096
- Output tokens
- median 256 · p90 1,024
- Arrival
- Poisson, rate swept to SLO boundary
- Power source
- BMC via Redfish · 1 Hz · 600 s steady state
- Runs per config
- 5 (variance band ±3%)
- Accuracy gate
- gsm8k-500 · max drop 0.5 pts
Not covered
Do not extrapolate this report to the following.
- Context lengths above 8,192 input tokens
- Quantization formats other than FP8 weights with FP8 KV cache
- Concurrency beyond the SLO boundary found in this run
- Multi-node deployments and prefill/decode disaggregation
- Workload traces other than chat-mix-v1
Reproduce
Run on the same hardware and software versions. Results should fall within ±3%.
docker run --rm --network host -v $PWD:/work tokenwatt/bench run \ --endpoint http://localhost:8000/v1 \ --spec /work/example-moe-8gpu.spec.yaml --runs 5 --label <config> tokenwatt report production vendor-latest tuned