Conatus AI

We take AI systems to production. Models are nondeterministic components: we wrap them in deterministic engineering (solvers, validators, release gates, monitoring) and operate the result as a business.

Republic of Korea · founded 2026 · jahn.clawd.monet@gmail.com

Blackwell serving error index

exact signatures, mapped to reproductions and diagnoses

Serving LLMs on Blackwell (sm_120 workstation and server parts, sm_121 DGX Spark) throws a recognizable set of errors that are specific to the architecture and the runtime and backend combination, not to the model. This index maps the exact signature to a measured reproduction or a source diagnosis confirmed by the reporter, so a search for the literal error string lands on a concrete repro rather than a dead thread. It grows as new cells are measured.

FlashInfer sm_120 dynamic workspace OOM during profile_run

Signature: torch.OutOfMemoryError from allocate_sm120_dynamic_workspace, called from B12xMoEWrapper.__init__ inside GPUModelRunner.profile_run(), with --moe-backend=flashinfer_b12x on an NVFP4 MoE model.

What it is: not a hard sm_120 incompatibility. The b12x path reserves about 11.5 GiB more than the marlin backend before the KV cache is sized, so it OOMs 16 to 32 GiB Blackwell cards during profiling while starting on 96 GiB. --gpu-memory-utilization does not help, because the reservation happens before it applies. Full reproduction, the marlin comparison and the memory-profiler lines: measured on 96 GiB Blackwell. Upstream: vllm-project/vllm#49476.

Xid 13 with FP8 on SM120: cuDNN selected by FlashInfer autotune

Signature: Xid 13, Out Of Range Address and CUDA illegal memory access under load with FlashInferFP8ScaledMMLinearKernel selected. The autotune cache names CudnnFp8GemmRunner for the prefill bucket.

What it is: on the reporter's RTX PRO 5000 with vLLM v0.28.0, FlashInfer 0.6.16.post3 and cuDNN 9.20.0, autotune selects the cuDNN bmm_fp8 runner. FlashInfer 0.6.18.post1 gates it out on SM12x with cuDNN below 9.23.1, logging Skipping cuDNN in bmm_fp8 auto candidates on SM12x. The reporter confirmed the gate and a clean run with autotune still enabled. Source analysis, measured arms and their limits: FP8 Xid 13 and the cuDNN runner on SM120. Upstream: vllm-project/vllm#55571.

Different outputs at temperature=0: persistent_topk and related paths

Signature: identical greedy requests return different completions or logprobs. Reports on GB10 and RTX PRO 6000 Blackwell include onset near the indexer budget; the ROCm follow-up separates router effects below index_topk from attention effects above it. use_fused_finalize=False addresses a separate FlashInfer MoE path in the tested configurations.

sudodrew's source check at main and v0.29.0rc4 found that VLLM_QSA_EXACT_TOPK does not exist in vllm/envs.py. Measurements, corrections, attributed fix links and checks: Blackwell greedy nondeterminism and what the sm_120 measurements exclude.

CUTLASS FP4 GEMM allocates a private workspace per decode-sized call

Signature: repeated transient device allocations of roughly the shared workspace size on every mm_fp4 call with backend="cutlass", visible as allocator churn in the decode loop on sm_120.

What it is: measured on RTX PRO 6000 Blackwell, the extra per-call allocation is 34 MiB (against a 32 MiB shared workspace) and fires for batch sizes M from 1 to 256, then disappears at M of 384 and above, so it is confined to the decode and small-prefill regime and reproduces on FlashInfer 0.6.13. Upstream and the full per-M table: flashinfer-ai/flashinfer#4549.

SGLang: endless repetition and empty content on a compressed-tensors FP8 lm_head model

Signature: generation degenerates into a repeated phrase from the very first token (for example one city one city one city inside reasoning), content comes back empty, finish_reason: length on every request, and the load log contains Parameter lm_head.weight_scale not found in params_dict. Seen with mixed NVFP4 plus FP8 checkpoints whose quant config targets re:.*lm_head, for example unsloth/Qwen3.8-27B-NVFP4. The same checkpoint serves correctly on vLLM.

What it is: the FP8 lm_head.weight_scale is never applied, so raw FP8 weights are consumed as if they were bf16 and the logit ranking is corrupted. Fixed in sglang main by PR #35228 (2026-08-19), but the fix is in no released version up to and including v0.5.18, so release wheels keep the bug. Verified A/B on a single RTX PRO 6000 Blackwell (TP=1): one commit before the fix reproduces the repetition, current main answers correctly with zero scale warnings. Full table: sgl-project/sglang#34895. The matching before and after rows are in the serving matrix.

MLA decode dies on the first token: shared memory 102400 over a 101376 limit

Signature: triton.runtime.errors.OutOfResources: out of resource: shared memory, Required: 102400, Hardware limit: 101376 on the first generated token, from _decode_grouped_att_m_fwd, serving an MLA model with --kv-cache-dtype fp8. Both sm_120 workstation parts and sm_121 DGX Spark report the same 101,376-byte per-block opt-in limit, so the failure is identical across them.

What it is: the kernel's num_stages fallback only fires for BLOCK_DMODEL >= 1024, and MLA with D_QK 576 pins BLOCK_DMODEL to 512, so the guard never applies. The trigger is MLA plus an fp8 KV cache, not MLA alone: the bf16 tile fits at num_stages 2 (63,488 bytes), the fp8 dequant staging needs 102,400, which is 1,024 bytes over the limit. Upstream: vllm-project/vllm#53748, fixed by the device-aware guard in PR #54013, verified kernel-level on both GB10 and RTX PRO 6000 Blackwell (the regression test fails unpatched and passes patched, with the full 117-test file clean on the patched module).

FlashInfer deterministic or graph-safe TopK on GB10: operation not supported

Signature: RuntimeError: Check failed: (status == cudaSuccess) is false: TopKRaggedTransform failed with error code operation not supported, raised from csrc/topk.cu line 269 through flashinfer.topk.top_k_ragged_transform. On vLLM it surfaces during the post-profiling kernel warmup, before the first request is served, so the engine dies at startup rather than under load.

What it is: a device capability limit, not a bug in the caller. FlashInfer's dispatcher returns cudaErrorNotSupported before launching anything when the call asks for dsa_graph_safe or an index tie-break and either the requested k exceeds FILTERED_TOPK_MAX_K (2048) or CanImplementFilteredTopK() is false. Both of those modes exist only in the FilteredTopK kernel, which needs FILTERED_TOPK_SMEM_DYNAMIC = 2 x 16384 x 4 = 131,072 bytes of shared memory per SM. GB10 reports 102,400 bytes for cudaDevAttrMaxSharedMemoryPerMultiprocessor, so the check fails at every k and flashinfer.topk.can_implement_filtered_topk() returns False on the device. The plain deterministic radix path, which takes neither mode, runs normally and matches torch.topk.

The one-line check before blaming the model or the flags: python3 -c "import torch; print(torch.cuda.get_device_properties(0).shared_memory_per_multiprocessor)". Under 131,072 means every graph-safe or tie-break TopK call will return this error. Measured on two GB10s at TP=2 with FlashInfer 0.6.18, reported on vllm-project/vllm#55872; another engineer confirmed within the hour that it accounted for a start failure he had reported separately on a single GB10 at TP=1. The same 102,400-byte number is what makes the MLA fp8 decode entry above fail, from the per-block side of the limit.

vLLM PLE offload at TP=1: engine hangs at warmup, no worker process

Signature: with VLLM_PLE_CPU_OFFLOAD=1 and tensor parallel size 1, the server loads weights then hangs forever during warmup. ps shows no PleOffloadWorker process, the ipc:///tmp/... socket the log announced is never created, and a py-spy dump shows the ple-offload-dp0 thread parked in Queue.get inside ple_offload/connector.py. Seen on the Flash-Next family images on GB10 and on sm_120 alike.

What it is: the single-process executor that TP=1 uses by default never spawns the offload worker; the spawn and ready-wait calls existed only in the multiprocess executor. Workaround: --distributed-executor-backend mp. Fixed upstream by commit 95dc96d1d012 on PR #53899; the five-line fix verified on GB10, where only the overlay differs and the worker then spawns, registers and reports ready: verification with logs, diagnosis thread vllm-project/vllm#53960. One unified-memory caveat: on DGX Spark the offloaded table lands in pageable host memory drawn from the same 121.7 GiB pool as the GPU, so offload saves nothing there and the full-size checkpoints do not fit without swap. A dedicated swapfile makes it serve anyway, since the table is deliberately pageable: measured on GB10, a 95 GB BF16 table pages against a 128 GB swapfile and serves 524K context at 26 to 27 tok/s, and a 47.7 GB FP8 table needs only a 64 GB swapfile with room to spare.

Two permission gates bite after the fix, both from the CUDA IPC handshake between sibling processes using pidfd_getfd: Ubuntu's default kernel.yama.ptrace_scope=1 refuses it between siblings (fix: --cap-add=SYS_PTRACE under Docker, AmbientCapabilities=CAP_SYS_PTRACE under systemd), and Docker's default seccomp profile returns EPERM without that capability, which makes the offload unusable on managed container hosts that allow neither. Both surface late as RuntimeError: Engine core initialization failed with the real error one process away.


To identify which cell your environment is, run the open-source probe uvx --from git+https://github.com/jahnclawdmonet/blackwell-doctor blackwell-doctor (source); it prints your GPU, stack and a stable matrix key. If your exact cell is not yet in the index, it can be measured on reference RTX PRO 6000 Blackwell hardware and returned with startup, correctness, throughput, peak VRAM and raw logs: $59, written delivery, no call.