Blackwell serving error index
exact signatures, mapped to reproductions and diagnosesServing LLMs on Blackwell (sm_120 workstation and server parts, sm_121 DGX Spark) throws a recognizable set of errors that are specific to the architecture and the runtime and backend combination, not to the model. This index maps the exact signature to a measured reproduction or a source diagnosis confirmed by the reporter, so a search for the literal error string lands on a concrete repro rather than a dead thread. It grows as new cells are measured.
FlashInfer sm_120 dynamic workspace OOM during profile_run
Signature: torch.OutOfMemoryError from
allocate_sm120_dynamic_workspace, called from
B12xMoEWrapper.__init__ inside
GPUModelRunner.profile_run(), with
--moe-backend=flashinfer_b12x on an NVFP4 MoE model.
What it is: not a hard sm_120 incompatibility. The b12x path reserves
about 11.5 GiB more than the marlin backend before the KV cache is
sized, so it OOMs 16 to 32 GiB Blackwell cards during profiling while
starting on 96 GiB. --gpu-memory-utilization does not
help, because the reservation happens before it applies. Full
reproduction, the marlin comparison and the memory-profiler lines:
measured on 96 GiB
Blackwell. Upstream:
vllm-project/vllm#49476.
Xid 13 with FP8 on SM120: cuDNN selected by FlashInfer autotune
Signature: Xid 13, Out Of Range Address and
CUDA illegal memory access under load with
FlashInferFP8ScaledMMLinearKernel selected. The autotune
cache names CudnnFp8GemmRunner for the prefill bucket.
What it is: on the reporter's RTX PRO 5000 with vLLM v0.28.0,
FlashInfer 0.6.16.post3 and cuDNN 9.20.0, autotune selects the cuDNN
bmm_fp8 runner. FlashInfer 0.6.18.post1 gates it out on
SM12x with cuDNN below 9.23.1, logging
Skipping cuDNN in bmm_fp8 auto candidates on SM12x.
The reporter confirmed the gate and a clean run with autotune still
enabled. Source analysis, measured arms and their limits:
FP8 Xid 13
and the cuDNN runner on SM120. Upstream:
vllm-project/vllm#55571.
Different outputs at temperature=0: persistent_topk and related paths
Signature: identical greedy requests return different completions or
logprobs. Reports on GB10 and RTX PRO 6000 Blackwell include onset
near the indexer budget; the ROCm follow-up separates router effects
below index_topk from attention effects above it.
use_fused_finalize=False addresses a separate FlashInfer
MoE path in the tested configurations.
sudodrew's source check
at main and v0.29.0rc4 found that
VLLM_QSA_EXACT_TOPK does not exist in
vllm/envs.py. Measurements, corrections, attributed fix
links and checks:
Blackwell
greedy nondeterminism and what the sm_120 measurements exclude.
CUTLASS FP4 GEMM allocates a private workspace per decode-sized call
Signature: repeated transient device allocations of roughly the shared
workspace size on every mm_fp4 call with
backend="cutlass", visible as allocator churn in the
decode loop on sm_120.
What it is: measured on RTX PRO 6000 Blackwell, the extra per-call allocation is 34 MiB (against a 32 MiB shared workspace) and fires for batch sizes M from 1 to 256, then disappears at M of 384 and above, so it is confined to the decode and small-prefill regime and reproduces on FlashInfer 0.6.13. Upstream and the full per-M table: flashinfer-ai/flashinfer#4549.
SGLang: endless repetition and empty content on a compressed-tensors FP8 lm_head model
Signature: generation degenerates into a repeated phrase from the very
first token (for example one city one city one city
inside reasoning), content comes back empty,
finish_reason: length on every request, and the load log
contains Parameter lm_head.weight_scale not found in
params_dict. Seen with mixed NVFP4 plus FP8 checkpoints whose
quant config targets re:.*lm_head, for example
unsloth/Qwen3.8-27B-NVFP4. The same checkpoint serves
correctly on vLLM.
What it is: the FP8 lm_head.weight_scale is never
applied, so raw FP8 weights are consumed as if they were bf16 and the
logit ranking is corrupted. Fixed in sglang main by PR #35228
(2026-08-19), but the fix is in no released version up to and
including v0.5.18, so release wheels keep the bug. Verified A/B on a
single RTX PRO 6000 Blackwell (TP=1): one commit before the fix
reproduces the repetition, current main answers correctly with zero
scale warnings. Full table:
sgl-project/sglang#34895.
The matching before and after rows are in the
serving matrix.
MLA decode dies on the first token: shared memory 102400 over a 101376 limit
Signature: triton.runtime.errors.OutOfResources: out of
resource: shared memory, Required: 102400, Hardware limit:
101376 on the first generated token, from
_decode_grouped_att_m_fwd, serving an MLA model with
--kv-cache-dtype fp8. Both sm_120 workstation parts and
sm_121 DGX Spark report the same 101,376-byte per-block opt-in limit,
so the failure is identical across them.
What it is: the kernel's num_stages fallback only fires
for BLOCK_DMODEL >= 1024, and MLA with D_QK 576 pins
BLOCK_DMODEL to 512, so the guard never applies. The trigger is MLA
plus an fp8 KV cache, not MLA alone: the bf16 tile fits at
num_stages 2 (63,488 bytes), the fp8 dequant staging needs 102,400,
which is 1,024 bytes over the limit. Upstream:
vllm-project/vllm#53748,
fixed by the device-aware guard in
PR #54013,
verified kernel-level on both GB10 and RTX PRO 6000 Blackwell (the
regression test fails unpatched and passes patched, with the full
117-test file clean on the patched module).
FlashInfer deterministic or graph-safe TopK on GB10: operation not supported
Signature: RuntimeError: Check failed: (status == cudaSuccess)
is false: TopKRaggedTransform failed with error code operation not
supported, raised from csrc/topk.cu line 269
through flashinfer.topk.top_k_ragged_transform. On vLLM
it surfaces during the post-profiling kernel warmup, before the first
request is served, so the engine dies at startup rather than under
load.
What it is: a device capability limit, not a bug in the caller.
FlashInfer's dispatcher returns cudaErrorNotSupported
before launching anything when the call asks for
dsa_graph_safe or an index tie-break and either the
requested k exceeds FILTERED_TOPK_MAX_K (2048) or
CanImplementFilteredTopK() is false. Both of those modes
exist only in the FilteredTopK kernel, which needs
FILTERED_TOPK_SMEM_DYNAMIC = 2 x 16384 x 4 = 131,072
bytes of shared memory per SM. GB10 reports 102,400 bytes for
cudaDevAttrMaxSharedMemoryPerMultiprocessor, so the check
fails at every k and
flashinfer.topk.can_implement_filtered_topk() returns
False on the device. The plain deterministic radix path, which takes
neither mode, runs normally and matches torch.topk.
The one-line check before blaming the model or the flags:
python3 -c "import torch;
print(torch.cuda.get_device_properties(0).shared_memory_per_multiprocessor)".
Under 131,072 means every graph-safe or tie-break TopK call will
return this error. Measured on two GB10s at TP=2 with FlashInfer
0.6.18, reported on
vllm-project/vllm#55872;
another engineer confirmed within the hour that it accounted for a
start failure he had reported separately on a single GB10 at TP=1.
The same 102,400-byte number is what makes the MLA fp8 decode entry
above fail, from the per-block side of the limit.
vLLM PLE offload at TP=1: engine hangs at warmup, no worker process
Signature: with VLLM_PLE_CPU_OFFLOAD=1 and tensor
parallel size 1, the server loads weights then hangs forever during
warmup. ps shows no PleOffloadWorker
process, the ipc:///tmp/... socket the log announced is
never created, and a py-spy dump shows the
ple-offload-dp0 thread parked in Queue.get
inside ple_offload/connector.py. Seen on the
Flash-Next family images on GB10 and on sm_120 alike.
What it is: the single-process executor that TP=1 uses by default
never spawns the offload worker; the spawn and ready-wait calls
existed only in the multiprocess executor. Workaround:
--distributed-executor-backend mp. Fixed upstream by
commit 95dc96d1d012 on
PR #53899;
the five-line fix verified on GB10, where only the overlay differs
and the worker then spawns, registers and reports ready:
verification with logs,
diagnosis thread
vllm-project/vllm#53960.
One unified-memory caveat: on DGX Spark the offloaded table lands in
pageable host memory drawn from the same 121.7 GiB pool as the GPU,
so offload saves nothing there and the full-size checkpoints do not
fit without swap. A dedicated swapfile makes it serve anyway, since
the table is deliberately pageable: measured on GB10, a 95 GB BF16
table pages against a 128 GB swapfile and serves 524K context at 26
to 27 tok/s, and a 47.7 GB FP8 table needs only a 64 GB swapfile
with room to spare.
Two permission gates bite after the fix, both from the CUDA IPC
handshake between sibling processes using pidfd_getfd:
Ubuntu's default kernel.yama.ptrace_scope=1 refuses it
between siblings (fix: --cap-add=SYS_PTRACE under
Docker, AmbientCapabilities=CAP_SYS_PTRACE under
systemd), and Docker's default seccomp profile returns EPERM without
that capability, which makes the offload unusable on managed
container hosts that allow neither. Both surface late as
RuntimeError: Engine core initialization failed with
the real error one process away.
To identify which cell your environment is, run the open-source probe
uvx --from git+https://github.com/jahnclawdmonet/blackwell-doctor blackwell-doctor
(source);
it prints your GPU, stack and a stable matrix key. If your exact cell
is not yet in the index, it can be measured on reference RTX PRO 6000
Blackwell hardware and returned with startup, correctness, throughput,
peak VRAM and raw logs:
$59, written delivery, no call.