Verify what you got
Drivers, firmware, form factor, memory bandwidth and power baselines — checked before a node carries traffic.
Seen: same instance type, two drivers — one served vLLM, one couldn't start it.Between a powered GPU and a reliable API sit drivers, operators, inference engines, kernels and gateways — and any release can break the stack. We validate, tune and measure it on AMD Instinct and NVIDIA, for public APIs and private ones. Measured, not claimed.
New capacity waits on interconnection, not on GPUs. Every energized watt that doesn't turn into a billable token is the most expensive kind of idle.
Peak tokens per second at unlimited latency tells you nothing about how many cards you need to serve your users within your TTFT and tail-latency targets.
A new driver, operator, engine image or quantization can swing throughput by double digits — or stop the stack from starting. Without a fixed ruler, nobody can tell what changed.
Each stage is a place we've seen stacks fail on real hardware — and each one ends with a measurement that shows it's fixed.
Drivers, firmware, form factor, memory bandwidth and power baselines — checked before a node carries traffic.
Seen: same instance type, two drivers — one served vLLM, one couldn't start it.Kubernetes or Slurm, GPU operators, vLLM or SGLang, an API gateway — integrated from open source and validated step by step.
Seen: four version and naming mismatches between correct components; two left no logs.SLO goodput, tokens per joule and an accuracy gate decide what goes to production — not peak throughput.
Seen: FP8 KV cache — +10.6% throughput, −4 points accuracy.Health baselines, regression checks on every upgrade, utilization and cost per million tokens.
Seen: an internal assistant for 1,000 people — two GPUs (one spare), 0.4% average utilization.For public APIs and internal ones. We work with a small number of design partners across the whole path — talk to us.
Every TokenWatt report pins the hardware, driver, engine, model, quantization, workload trace and SLO — then states exactly what it does and doesn't cover.
How much of the rated HBM bandwidth the decode loop actually turns into useful reads. The clearest signal of software headroom.
From the BMC or a metered PDU, sampled through the steady-state window. GPU board power is reported separately and labeled.
Output tokens per second at the highest load where p99 TTFT and p99 time-per-output-token both stay inside your targets.
A fixed evaluation set with a pass threshold agreed before the run. Speed that costs quality doesn't count.
Bench is a client. It talks to any OpenAI-compatible server — vLLM, SGLang, TensorRT-LLM — and reads telemetry from the node. It never touches your weights.
Declare TTFT and tail-latency targets, the request-length distribution and the eval set up front. The spec file is part of the report.
Production image, vendor's latest recommended image, and any candidate change — measured the same way, on the same node.
Machine-readable JSON plus a one-page summary with the capacity math: how many cards and kilowatts the same billable throughput needs.
# 1. describe what "good" means for you $ cat slo.yaml slo: ttft_p99_ms: 2000 tpot_p99_ms: 50 workload: trace: chat-mix-v1 # input/output length distribution accuracy: suite: gsm8k-500 max_drop_pts: 0.5 power: source: redfish # or ipmi-dcmi, pdu # 2. measure each configuration $ docker run --rm --network host \ -v $PWD:/work tokenwatt/bench run \ --endpoint http://localhost:8000/v1 \ --spec /work/slo.yaml --label production # 3. compare and write the report $ tokenwatt report production vendor-latest tuned ✓ report written: tw-report-2026-10-06.json ✓ summary written: tw-report-2026-10-06.html
We also build tuned inference images for AMD Instinct. That's exactly why the rules below are strict and public.
AMD Instinct and NVIDIA are measured with the same definitions, the same workload traces and the same power sources.
Every comparison includes the vendor's own latest recommended configuration. We never benchmark against a configuration we picked to lose.
Definitions, workload traces and the measurement core are published. Anyone can rerun a report and get the same answer within its stated band.
Bench is read-only. No weights, prompts or logs leave your environment unless you choose to publish a report.
Node power comes from the BMC or a metered PDU. Nameplate TDP is never used to estimate energy.
Every report lists the context lengths, quantizations and concurrency levels it did not test. No extrapolation, no peak-only numbers.
One telemetry core, three products, all priced per GPU and installed by your own engineers.
Vendor-neutral inference measurement for your models, SLOs and nodes.
Open core. Paid tier adds power capture, history and reports.
Find the GPU that isn't broken — just slow.
Read-only agent. Installs in minutes.
Decode-path tuned inference images for AMD Instinct, verified with Bench.
Per-GPU subscription. Your gains, verified by you.
Every public report ships with its spec file, raw JSON and the command that reproduces it.
Every run we do on real hardware gets written up: setup, numbers, surprises and the commands to reproduce it.
Using measured SLO throughput for Qwen3-30B-A3B on one H100 and live OpenRouter prices, we work out the break-even utilization for rented, owned and idle GPUs. Rented GPUs need 64–77% utilization to match the cheapest price; idle owned GPUs break even at 2%. Includes a calculator.
Sizing private LLM inference from a measured H100 benchmark. One H100 can carry the peak chat traffic of roughly 12,000 employees, but at 1,000 employees it sits 99.6% idle and a million tokens costs $67.70 instead of $0.30. Utilization, not GPU count, decides the economics. Includes a calculator.
Full-mode rerun of Qwen3-30B-A3B-FP8 on one H100 with vLLM 0.31.0. With 10-minute windows the SLO boundary fell from 11.3 to 9.5 req/s and goodput from 4,572 to 3,996 tok/s; the FP8 KV-cache accuracy drop shrank from 8 to 4 points on a 600-sample suite.
We're looking for a small number of US inference teams and GPU clouds to work with closely. You get Bench on your fleet and a say in the roadmap; we get honest feedback.