Serving cost and latency are decided by the stack as much as by
the model. We benchmark your model on your target hardware across
candidate serving stacks, tune the winner, and hand back the
configuration together with scripts that reproduce every number
in the report.
01.1 Case study: Qwen3-8B on Blackwell, three stacks plus FP8
measured 2026-08-25
What the $349 benchmark actually produces, run on our own
hardware: Qwen3-8B on one RTX PRO 6000 Blackwell (96 GB),
identical prompts and settings across vLLM 0.27.1, SGLang 0.5.9
and llama.cpp, then an FP8 pass on the winner to isolate
quantization. Greedy decoding, token counts matched across
engines before timing was compared.
Findings: stack choice barely matters for one user (83 to 96
tok/s everywhere) but is a 4x decision under concurrent load;
time to first token differs 4x between engine families; FP8 gave
a clean 1.5x with zero factual regressions on a fixed 20-prompt
check. Getting FP8 to run on this workstation-class Blackwell
chip (sm_120) required routing around a kernel assertion in the
default FP8 path, which is exactly the class of work the kernel
tier exists for.
Every number ships with raw CSVs, an environment manifest and a
one-command reproduction script. The script was re-run end to
end after the report was written: all figures reproduced within
6 percent. Read the full
sample report or download the raw
evidence bundle. To see what numbers like these do to your
serving bill, use the self-hosted
LLM cost calculator.
SKU E Single-cell serving verification
$59 fixed
The narrowest possible engagement, for one exact question: does
this specific combination serve correctly on workstation-class
Blackwell, and at what cost. Fixed scope is one public model and
revision, one runtime and version, one quantization format, one
tensor or expert parallel topology, and one input, output and
concurrency point. Run on our own RTX PRO 6000 Blackwell (sm_120,
96 GB).
- a Startup pass or fail, and a basic correctness check
against a reference.
- b Throughput (tokens per second, or TTFT and ITL) and peak
VRAM for that cell.
- c The exact command, an environment manifest, and the raw
JSON and logs.
Delivered in writing within 2 days. Public model weights only. No
root cause diagnosis, no code changes, and no access to your
systems: those are the benchmark and kernel tiers below. If the
cell turns out to already be known good, the report proving it is
the deliverable.
Order a single-cell serving
verification, $59
Identify your cell first with the open-source probe
uvx --from git+https://github.com/jahnclawdmonet/blackwell-doctor blackwell-doctor, and check whether your error is
already reproduced in the
Blackwell
serving error index (for example the
FlashInfer
sm_120 workspace OOM and
FP8
Xid 13 and the cuDNN runner on SM120, plus
greedy
nondeterminism at temperature=0).
SKU C Inference stack benchmark and tuning
$349 fixed
Fixed scope: one model, one target GPU, up to three candidate
serving stacks chosen from vLLM, SGLang, TensorRT-LLM, ONNX
Runtime and llama.cpp, plus one quantization pass (FP8, INT8,
INT4, AWQ or GPTQ, whichever fits the model and hardware).
- a Benchmark report: tokens per second, p50 and p99
latency, VRAM use, and cost per 1M tokens, before and
after.
- b The tuned configuration files for the recommended
stack.
- c Reproduction scripts. Every number in the report can be
re-run in your environment.
Delivered within 5 days after access is in place. If the
benchmark shows your current setup is already at the practical
limit, the report proving that is the deliverable.
SKU D Custom kernel engineering
from $1,000 per kernel
Quoted only after C, and only when the C report has isolated a
bottleneck that kernel work can move. Scope: CUDA or Triton
kernels written for your hardware generation, fused ops, and
custom quantization kernels.
- a The contract states the target speedup before work
starts.
- b If the target is missed, the final payment is
waived.
- c A time cap is written into the contract, so the work
cannot run open ended.
Deliverables: kernel source, an integration patch for your
stack, before and after benchmarks, and tests.