Lab notes #3: Nine providers, one Llama 3.3 70B — same accuracy, 18× different speed
We audited every OpenRouter provider of Llama 3.3 70B Instruct with the same GSM8K suite and latency probe. Accuracy was statistically indistinguishable (94.0–98.0%), decode speed ranged from 19 to 347 tokens/s, and output prices varied 7×. A naive first pass wrongly scored two providers at 0% and 44%.
When the same open model is served by many providers, buyers usually compare price and maybe speed. They rarely check quality. A provider can quantize harder, change kernels or cap generation, and the model name stays the same.
So we pinned every request to one provider at a time and measured them all the same way. The model was Llama 3.3 70B Instruct, chosen because it has many providers on OpenRouter and they declare different precisions (FP8, BF16, FP16, or nothing).
Short version:
- Accuracy: no meaningful difference. All nine measurable providers scored between 94.0% and 98.0% on 200 GSM8K questions. That spread is within sampling noise for a 200-item test. Providers declaring FP8 were not worse than those declaring FP16 or BF16.
- Speed: 18× apart. Median decode speed ranged from 19 to 347 tokens per second. Median time to first token ranged from 330 to 766 ms.
- Price: 7× apart, and not tied to speed. Output prices ran from $0.32 to $2.25 per million tokens. The cheapest provider was also the most accurate, and the slowest.
- A naive audit gets this badly wrong. Our first pass scored Cloudflare at 0% and Novita at 44%. The cause was API incompatibilities and rate limits, not model quality.
The whole audit cost $1.13 in API usage. The raw data is at /data/audits/llama-3.3-70b-2026-10-07.json.
Method
| Model | meta-llama/llama-3.3-70b-instruct via OpenRouter, 2026-10-06/07 |
| Routing | each request pinned with provider.only = [<endpoint>], allow_fallbacks = false; a routing check confirmed the serving provider before each audit |
| Accuracy | GSM8K, the same 5-shot prompts, first 200 test questions, temperature 0, stop at the next "Question:", max 384 tokens, scored on the final number |
| Latency | 20 streamed requests per provider, sent one at a time (about 1,000 input words, up to 256 output tokens) — we don't load-test other people's APIs |
| Failures | rate limits and server errors retried up to 4 times with backoff; remaining failures counted separately |
Results
| Provider | Declared precision | GSM8K (200) | ±95% CI | Failed | TTFT p50 | Decode p50 | $ in / out per 1M |
|---|---|---|---|---|---|---|---|
| DeepInfra (turbo) | FP8 | 98.0% | ±1.9 | 0 | 514 ms | 19 tok/s | $0.10 / $0.32 |
| AkashML | FP8 | 96.5% | ±2.5 | 0 | 732 ms | 26 tok/s | $0.20 / $0.52 |
| SambaNova (turbo) | not declared | 96.5% | ±2.5 | 0 | 428 ms | 287 tok/s | $0.45 / $0.90 |
| Parasail | FP8 | 95.5% | ±2.9 | 0 | 624 ms | 67 tok/s | $0.22 / $0.50 |
| Groq | not declared | 95.5% | ±2.9 | 0 | 330 ms | 347 tok/s | $0.59 / $0.79 |
| Together | not declared | 95.5% | ±2.9 | 0 | 643 ms | 68 tok/s | $1.04 / $1.04 |
| CoreWeave | FP16 | 95.0% | ±3.0 | 0 | 382 ms | 82 tok/s | $0.71 / $0.71 |
| Novita | BF16 | 94.5% | ±3.2 | 1 | 766 ms | 52 tok/s | $0.14 / $0.40 |
| Cloudflare | FP8 | 94.0% | ±3.3 | 3 | 743 ms | 64 tok/s | $0.29 / $2.25 |
Google Vertex's two endpoints were listed but couldn't be routed to: every pinned request returned "No endpoints found". They aren't in the table.
Is a 4-point spread real? The best and worst providers differ by 4.0 points; on 200 questions each, that gives z ≈ 2.0 before correcting for comparing nine providers at once. After correction it isn't significant. On this test, every provider serves Llama 3.3 70B at essentially the same quality. A harder suite or more samples could still separate them.
Speed and price are separate decisions
- If you need interactive speed, Groq (347 tok/s) and SambaNova (287 tok/s) are in a different class. The rest decode at 19–82 tok/s from a single request.
- If you need cheap bulk tokens, DeepInfra's turbo endpoint costs $0.32 per million output tokens and scored highest. It's also the slowest, which is fine for batch jobs and frustrating for chat.
- Price doesn't buy speed. Cloudflare's output price is the highest on the list at $2.25, with mid-pack speed. Together charges $1.04 for 68 tok/s.
These latencies come from one client location in the US, one request at a time, through OpenRouter, on one evening. Under load, and from other regions, they will differ. Treat them as a snapshot, not a ranking.
The audit that almost went wrong
Our first pass sent the same requests to every provider with a fixed seed and eight requests in parallel. It produced this:
| Provider | First pass | After the fix | What went wrong |
|---|---|---|---|
| Cloudflare | 0.0% | 94.0% | rejects the seed parameter; all 200 requests returned HTTP 400 |
| Novita | 44.0% | 94.5% | rate limits (HTTP 429) on parallel requests |
| Groq | 85.0% | 95.5% | intermittent failures |
| AkashML | 87.5% | 96.5% | intermittent failures |
Had we counted every failed request as a wrong answer and published, we'd have accused two providers of serving a broken model. Bench now retries rate limits and server errors with backoff, drops seed when a provider rejects it (at temperature 0 it doesn't change results), and reports failures separately from wrong answers.
If you compare providers yourself: check failure rates before you trust a score.
Not covered
- Harder or longer tasks: reasoning, long context, tool use, code. GSM8K is a quick sanity check, not a full quality evaluation.
- Throughput under load, and latency from other regions or times of day.
- Variation over time: providers change hardware, kernels and quantization without notice. That's why audits like this need repeating.
Reproduce
cd bench
python scripts/make_gsm8k_suite.py --out suites/gsm8k-200.jsonl
python scripts/audit_openrouter.py --model meta-llama/llama-3.3-70b-instruct --concurrency 4
Next: the same audit for Qwen and DeepSeek models, where providers declare FP4 and FP8 side by side.