You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This document describes the environment variables used in the ATOM project.
Metadata H2D
Variable
Type
Default
Description
ATOM_H2D_BACKEND
str
packed
packed combines forward metadata into one H2D and GPU scatter per consumer group. direct copies each member separately. Both preserve source reuse gates and full cudagraph padding. Set before starting the runner. See metadata publication.
Data parallelism
Variable
Type
Default
Description
ATOM_DP_RANK
int
0
The rank ID for the current process in data parallelism.
ATOM_DP_RANK_LOCAL
int
0
The local rank ID for the current process (used in SPMD mode).
ATOM_DP_SIZE
int
1
Total number of data parallel ranks.
ATOM_DP_MASTER_IP
str
127.0.0.1
Master IP address for DP ranks coordination.
ATOM_DP_MASTER_PORT
int
29500
Master port for DP ranks coordination.
ATOM_DP_LB_REQ_EQUIV
int
512
Token-equivalent decode pressure assigned to each in-flight request by least_tokens routing.
ATOM_DP_SESSION_AFFINITY
bool
false
Load-place each new session, then keep later turns on the same prefix-cache owner. Reads X-Dynamo-Session-ID, falling back to X-Correlation-ID.
Prefill delayer (TP/DCP and DP attention)
Coalesces waiting prefills while decode continues. DP attention enables it by
default through ATOM_ENABLE_PREFILL_DELAYER. For a single scheduler (DP=1,
PP=1), including TP/DCP, it is opt-in: set ATOM_PREFILL_DECODE_INTERVAL above
zero and keep the master switch enabled. Interval 0 leaves TP scheduling
unchanged; setting only the master switch does not enable TP coalescing.
On TP, the interval and coalescer are enabled together.
This applies to the standard scheduler, including connector-based P/D roles.
RapidServe's dedicated PrefillScheduler/DecodeScheduler do not use the delayer.
After each executed prefill, the decode interval runs before all coalescing
bounds. Once it expires, fill, queue age, KV pressure, partial-prefill and stall
bounds decide when to release. MAX_QUEUE_MS stops extra coalescing after that
interval; it does not guarantee end-to-end TTFT. DP decisions reduce local
signals across ranks to keep their phases aligned.
The local fill signal discounts HBM cache hits and uses the admission path's
chunk limits. To bound CPU work, it stops probing fresh requests after their
total prompt length reaches one batch budget. It reports only work found so
far; it does not treat unseen requests as a full batch. A deep, cache-heavy
queue can therefore release through the stall or hold bounds before reaching
the fill target. If parked transfers exhaust the unreserved slots, a fresh
request may signal possible work with zero estimated tokens until admission
resolves its connector match. This is an estimate, not a reservation: checkpoint dependencies,
connector results and resource changes during admission can still reduce a batch.
Repeated probes reuse immutable prompt hashes while rechecking pool contents
and resource fit. During decode protection, only the existence of queued or
partial work is checked; partial-prefill hold bounds start after the interval.
Local hybrid models with state checkpointing can wait for an in-flight
producer's reusable prompt-end checkpoint. The expected prefix must exceed both
the HBM hit and any offload match by at least one prefill budget. P/D transfers
and already-started offload loads keep their own progress paths. Deferred
requests retain their relative order, while at most 16 later queue entries are
examined for independent work each pass. Each request's wait expires after
TTFT_MAX_TICKS scheduler passes from its first dependency wait; bypassing it
does not restart that deadline. MAX_QUEUE_MS bounds coalescing, not this
dependency wait: time spent queued before a producer becomes runnable does not
make duplicate prefill useful. Pure-attention models do not use checkpoint waits.
Variable
Type
Default
Description
ATOM_ENABLE_PREFILL_DELAYER
bool
true
Master switch for the prefill coalescer.
ATOM_PREFILL_DELAYER_TARGET_FILL
float
0.9
Release once accumulated pending tokens reach target_fill × max_num_batched_tokens (averaged across prefillable ranks). In (0, 1]; higher = fewer, larger prefills at some TTFT cost. Clamped to (0, 1].
ATOM_PREFILL_DELAYER_TTFT_MAX_TICKS
int
200
Max consecutive scheduler ticks a held prefill waits before force-release. Values < 1 clamped to 1.
ATOM_PREFILL_DELAYER_PARTIAL_MAX_TICKS
int
100
Tighter bound for a held mid-chunked-prefill (it holds allocated KV). Values < 1 clamped to 1.
ATOM_PREFILL_DELAYER_STALL_TICKS
int
10
After this many consecutive non-growing ticks, release (burst ended, more won't come). Values < 1 clamped to 1.
ATOM_PREFILL_DELAYER_KV_HIGH_WATERMARK
float
0.9
At/above this KV usage a prefillable rank force-releases (can't accumulate a bigger batch anyway).
ATOM_PREFILL_DELAYER_TOKEN_USAGE_LOW_WATERMARK
float|""
"" (None)
If set, a prefillable rank below this KV usage force-releases (GPU starving).
ATOM_PREFILL_DELAYER_MAX_QUEUE_MS
float|""
"" (None)
After decode protection, release coalescing when the oldest schedulable waiting prefill reaches this age since arrival. Empty disables the age guard. Checkpoint dependency waits use TTFT_MAX_TICKS. This is not a hard TTFT limit.
ATOM_PREFILL_DECODE_INTERVAL
int
0
Protect this many scheduler passes after an executed prefill. On DP=1, PP=1, a positive value also enables local coalescing when the master switch is on; 0 leaves TP scheduling unchanged. On DP>1, 0 disables only the interval.
ATOM_PREFILL_DELAYER_DEBUG
bool
false
Per-tick FIRE/HOLD debug logging.
ATOM_PREFILL_DELAYER_LOG_EVERY
int
1000
Emit aggregate stats (per-exit fire counts + hold rate) every N decisions (0 disables).
Model loading
Variable
Type
Default
Description
ATOM_DISABLE_MMAP
bool
false
If set to true, disable memory-mapped file loading for model weights. Useful in containerized environments where mmap may cause issues.
ATOM_LOADER_NUM_THREADS
int
16
Worker threads for weight loading. >1 (default 16) enables the batched parallel loader (routed expert weights staged in a CPU buffer, flushed with a single H2D copy when every routed expert of that parameter has arrived) with that many threads; set to 1 to fall back to the original sequential per-expert path. Raise on high-core hosts if loading is CPU-bound.
ATOM_LOADER_STRICT_COVERAGE
bool
true
Fail loading when a fused MoE parameter does not receive every routed expert from the checkpoint. Set to false to downgrade to a warning and load anyway, leaving those expert slots at their init values — useful when bringing up a checkpoint known to be partial, misleading otherwise (the symptom is an accuracy drop much later).
ATOM_LOADER_PREFETCH
bool
true
Warm the page cache by reading this rank's share of the checkpoint sequentially on a background thread, instead of leaving it to demand faults through the mmap. The fault pattern sustains ~3.2 GB/s on a local NVMe that a single sequential reader drives at 6.06 GB/s, so this is an access-pattern fix, not a queue-depth one. Measured on DeepSeek-R1 MXFP4 (350 GiB, TP=4): cold load 154s → 69s. Set to false to restore demand faulting. Also applies with ATOM_DISABLE_MMAP=true, whose whole-file reads go through the same page cache (DeepSeek-V4.1-Flash, TP=4, cold: 282–315s without it, 169–207s with it).
ATOM_LOADER_PREFETCH_THREADS
int
4
Concurrent sequential readers per rank used by the prefetcher. The device saturates at ~2 streams, so raising this mostly adds contention with the loader; 0 is clamped to 1 (use ATOM_LOADER_PREFETCH=false to switch prefetching off).
ATOM_LOADER_PREFETCH_BLOCK_MB
int
16
Read block size for the prefetcher, in MiB.
ATOM_LOADER_FADVISE
bool
false
Issue posix_fadvise(SEQUENTIAL|WILLNEED) per shard before reading it. Off by default and ignored while ATOM_LOADER_PREFETCH is on: WILLNEED is a hint the kernel drops for most of a 350 GiB checkpoint, and running both makes the kernel read ahead over random-ish ranges while the prefetcher streams the same files, so the two compete for the device. Only useful with prefetching disabled.
ATOM_ONLINE_QUANT_STREAMING
bool
false
Opt in to quantizing eligible online-quant modules as soon as their checkpoint weights are complete, then release source storage to reduce load-time peak memory. Only active with a valid online quantization config. See the streaming online quantization guide.
ATOM_ONLINE_QUANT_STREAMING_HOST_STAGING
bool
true
Assemble streamed module weights in CPU storage before one H2D transfer. Keeps the checkpoint walk parallel; disabling it buffers loader calls and forces the checkpoint walk to one thread.
ATOM_ONLINE_QUANT_STREAMING_THREADS
int
4
Tail workers for H2D, per-module quantization, and source release. More workers increase overlap and in-flight memory; 0 runs finalization inline.
Plugin mode
Variable
Type
Default
Description
ATOM_DISABLE_VLLM_PLUGIN
bool
0 (false)
If set to 1, disable the vLLM plugin registration entirely.
Kernel / backend selection
Variable
Type
Default
Description
ATOM_USE_TRITON_GEMM
bool
0 (false)
If set to 1, use AITER Triton FP4 weight preshuffled GEMM. Otherwise use AITER ASM FP4 weight preshuffled GEMM.
ATOM_FP8_BLOCKSCALE_USE_E8M0_SCALE
bool
unset (per checkpoint)
E8M0 rather than FP32 128x128 FP8 block scales, for the weight and the activation quantized for it; on gfx950 E8M0 operands take the AITER FlyDSL GEMM. Unset: E8M0 on every arch but gfx942 when the checkpoint declares scale_fmt: ue8m0 (DeepSeek-V4), whose power-of-two scales E8M0 restates exactly; FP32 otherwise. 1/0 force it on/off; forcing it on for a checkpoint with non-power-of-two FP32 scales rounds them.
ATOM_GROUP32_WEIGHT_PRESHUFFLE
bool
1 (true)
On gfx950, (16, 16)-shuffle native FP8 32x32 group32 weights (DeepSeek-V4.1) at load and after a weight sync, so AITER's preshuffled group32 GEMM (FlyDSL) reads them; that GEMM emits BF16 only. Set to 0 to keep checkpoint bytes and the row-major group32 GEMM. V4.1's grouped wo_a is shuffled either way.
ATOM_USE_FP4_NON_SHUFFLE_TRITON_GEMM
bool
0 (false)
If set to 1, use AITER Triton FP4 GEMM with non-shuffled weights. Takes precedence over the FP4 preshuffled GEMM path selected by ATOM_USE_TRITON_GEMM.
ATOM_MHC_USE_BF16
bool
1 (true)
Use AITER BF16 hi/lo mHC computation for attention, FFN and head. After loading, replace FP32 fn storage with mhc_shuffle_fn output; no FP32 copy is retained. Set to 0 for FP32 mHC. Takes effect at model load; restart to change modes. On gfx1250, AITER enables shuffled residuals only while its runtime mhc_fused_post_pre policy remains fused (M < 1024); larger M uses ordinary residual layout and the standalone post/pre fallback.
ATOM_USE_TRITON_MXFP4_BMM
bool
0 (false)
If set to 1, use FP4 BMM in MLA attention module.
ATOM_USE_FLYDSL_GATHER_KV_B_PROJ
bool
1 (true)
Use the FlyDSL fused gather + kv_b_proj GEMM for MLA's cached-prefix path. Covers page_size-1 fp8 (e4m3) KV with an fp8 weight on gfx950 — i.e. Kimi-K3 / DeepSeek MLA under --kv-cache-dtype fp8. Unavailable imports or unsupported tensor configurations use Triton; kernel execution errors propagate. Set 0 to force Triton.
ATOM_USE_FLYDSL_FP8_PREFILL_ATTN
bool
0 (false)
Enable FlyDSL FP8 MLA prefill on gfx950, including fused K concatenation with QKV quantization and direct FP8 output from FlyDSL gather where supported. AITER capability checks are cached per layer. Missing FP8 attention support or unsupported device/model configurations select BF16 attention before quantization and emit a warning once per process. FP16 activations and synthetic RoPE-padding configurations also use the normal attention path. Cached K/V reuse descales max(new_token_descale, 1e-6) * 2; larger outliers saturate. Unsupported FP8 gather configurations use BF16 gather followed by fused K/V dynamic quantization when FP8 attention is supported. Kernel execution errors propagate without retry. Added 2026-09-10.
ATOM_PA_FLYDSL
bool
0 (false)
Route the MHA paged decode to aiter's FlyDSL kernel (aiter #4332) instead of the gluon one, where FlyDSL's domain covers the call. Off by default: the kernel is not fully tested on every shape ATOM ships. The env is the first of two gates; the second mirrors aiter's own validation so an unsupported shape falls back to gluon rather than raising inside aiter, and the and short-circuits so nothing is evaluated when this is off. Global, not per call site: the vLLM bridge is rerouted too. Needs aiter 94dca7bc6 or later.
ATOM_PA_FLYDSL_PLAN
bool
1 (true)
Use aiter #5546's GPU work planner on the dense decode, built once per forward in the metadata builder. Requires ATOM_PA_FLYDSL=1 — with FlyDSL off no plan is built at all. Sizes each request's partition count from its real context length instead of splitting the batch uniformly, which is worth p90 interactivity +20.5% at concurrency 20 on the MiniMax-M3 agentic trace and nothing at concurrency 1, where there is nothing to rebalance. Kept off the two MiniMax-M3 sparse call sites, whose contexts are a fixed topk window and where it measures a net loss. Plans are only minted during cudagraph capture — aiter's planner takes the batch as a tl.constexpr, so a new value is a cold kernel specialization and a plan no captured graph can ever release — which leaves the planner inert under --enforce-eager.
GLM-5.3
Variable
Type
Default
Description
ATOM_GLM5_KPOOL
bool
1 (true)
Enable the pooled sparse indexer. Setting 0 is an exact token-granular A/B only at or below index_topk; longer requests are refused.
ATOM_GLM5_FORCE_DENSE_MLA
bool
0 (false)
Disable sparse MLA for short-context bring-up comparisons.
ATOM_GLM5_DISABLE_FUSED_MHC
bool
0 (false)
Force the PyTorch mHC reference path instead of AITER's fused kernels.
MiniMax-M3
Variable
Type
Default
Description
ATOM_MONO_ENABLE
bool
1 (true)
Fused per-layer decode, on by default (0 disables it); MiniMax-M3 is the only model with a mono path so far. Route decode steps of up to 16 tokens (requests × speculative query tokens; a verify's tokens run as one row each) to the fused layer kernels in atom/models/minimax_m3/mono/, Eagle3 aux hidden states included. Only a supported configuration is routed (TP4, ptpc_fp8 attention linears, fp8 KV and index cache, max_model_len up to 1M, no index-cache reuse, no TBO / DP / PP / plugin mode, FULL cudagraph or eager); everything else keeps the original model. The same entry class serves both paths.
ATOM_MONO_CHECK
bool
0 (false)
Debug aid for ATOM_MONO_ENABLE: every sparse layer also runs the original decoder layer on the same input (after the mono layer, so mono reads only the cache entries it inserted) and rank 0 logs the per-layer, per-token difference of the residual and of the reduced output. Run it with --enforce-eager.
ATOM_MONO_TRACE
path
unset
Debug aid for ATOM_MONO_ENABLE: rank 0 appends every mono step's input tokens, positions, top-2 logits and greedy pick to this file (one JSON line per step), with a fingerprint of every stage (dense layers, each sparse layer's output, the cache history it reads). Two runs of the same prompt then align token by token. Run it with --enforce-eager.
ATOM_MONO_TIMELINE
path prefix
unset
Debug aid for ATOM_MONO_ENABLE: the mono layer kernels are built with their per-phase s_memrealtime stamps, each sparse layer into its own buffer, and after 20 warm-up steps of every decode token count S the next 5 steps are saved per rank as <prefix>_r<rank>_s<S>_<i>.pt (int64 [layer][CTA][stamp], 100 MHz ticks, 0 = not reached). Run it with --enforce-eager: a graph replay runs no Python.
MoE all2all (MoRI) wire format
Both are opt-in and default to off; they only apply with DP attention + expert
parallelism. They are not symmetric — FP4 dispatch only moves a quantization
the MoE GEMM was going to perform anyway (it consumes FP4 activations either
way, and per_1x32 is per-row, so it does not matter which rank runs it), while
FP8 combine adds a quantization that would not otherwise happen, since the
expert output is bf16. Treat the dispatch knob as format matching and the
combine knob as a quality/throughput tradeoff.
Variable
Type
Default
Description
ATOM_MORI_FP4_DISPATCH
bool
0 (false)
If set to 1, quantize activations to packed FP4 (E2M1, per_1x32) before the MoE all2all instead of sending bf16 — a quarter of the bytes on the dispatch wire — which selects EpDispatchIntraNodeKernel_fp4. MoRI picks its dispatch kernel from the dtype of the tensor handed to dispatch() but sizes its staging buffers from the config built at init, so this also switches scale_dim to hidden_dim/32 and the scale type to e8m0. All three are resolved together by mori_prepare_finalize.resolve_mori_dispatch(); never set one without the others, as a mismatch strides the staging scale buffer wrong and faults on the first real batch instead of erroring cleanly.
ATOM_MEGA_COMBINE_WIRE
str
bf16
MegaMoE (ATOM_MORI_V2_FUSED=1) combine wire: bf16, fp8 (mxfp8) or fp4 (mxfp4). Prefill-only: decode steps always combine in bf16, and the choice is DP-agreed so every rank reduces in the same format.
ATOM_MORI_COMBINE_QUANT
str
none
Combine-side codec passed into the MoRI config. none returns bf16; fp8_blockwise selects EpCombineIntraNodeKernel_*_fp8bwq_*; MoRI also accepts fp8_direct_cast.
Fusion passes
RMSNorm
Variable
Type
Default
Description
ATOM_USE_MODEL_SENSITIVE_RMSNORM
bool
0 (false)
If set to 1, use AITER's model-sensitive RMSNorm rounding mode. This can change numerical results and prefill performance, so it is opt-in.
TP AllReduce fusion
Variable
Type
Default
Description
ATOM_ENABLE_ALLREDUCE_RMSNORM_FUSION
bool
1 (true)
If set to 1, fuse allreduce with RMSNorm in tensor parallel mode.
DeepSeek-style
Variable
Type
Default
Description
ATOM_ENABLE_DS_INPUT_RMSNORM_QUANT_FUSION
bool
1 (true)
If set to 1, fuse RMSNorm with quantization.
ATOM_ENABLE_DS_QKNORM_FUSION
bool
1 (true)
If set to 1, use the fused Q/K RMSNorm path (fused_qk_rmsnorm) in the DeepSeek MLA attention module when Q-LoRA is enabled and QK norm+quant fusion is not used. If set to 0, apply separate RMSNorm for the Q and KV branches instead.
ATOM_ENABLE_DS_QKNORM_QUANT_FUSION
bool
1 (true)
If set to 1, fuse QK norm with quantization in MLA attention module.
ATOM_DUAL_STREAM_MOE_TOKEN_THRESHOLD
int
1024
Upper bound on MoE token count (num_tokens in the MoE forward) for using the dual-stream path: shared experts on a secondary CUDA stream while routed experts run on the default stream. If num_tokens exceeds this value, that forward uses single-stream MoE instead. Set to 0 to disable dual-stream setup entirely (no alt stream, no maybe_dual_stream_forward registration).
ATOM_DUAL_STREAM_PIECEWISE
bool
0
Opt-in: allow a PIECEWISE-captured graph piece to hold the MoE dual-stream fork/join (shared experts on alt_stream overlapping routed experts). Capture is not the obstacle — set_forward_context runs inside graph_capture(), so the main stream the fork waits on is the stream capture runs on — and vLLM and SGLang both keep this overlap on inside piecewise graphs (SGLang runs dual-stream only inside a graph). Measured on V4-Pro-DSpark under AF_PIECEWISE: the fork survives capture, hides 77.5% of shared-expert time, and leaves GSM8K and MTP acceptance unmoved. Off by default only because no throughput win has been demonstrated, and because each replayed piece then carries its own driver-allocated stream (368 vs 2 distinct streams on a tp8 rank trace). The dispatcher is shared, so this moves V2/V3.2/K3 as well. Eager (NONE) and whole-model FULL are unaffected.
DSpark block sampling
DSpark drafts a num_speculative_tokens-wide block in one backbone pass, then
samples it left-to-right with a low-rank first-order Markov head
(logits_k = base_logits_k + W1[x_{k-1}] @ W2ᵀ, x_k = argmax(logits_k)). The
unfused loop casts the whole [V, r]W2 table to fp32 on every iteration and
materializes two [B, V] fp32 tensors that only an argmax reads. See
atom/model_ops/dspark_markov_sample.py.
Variable
Type
Default
Description
ATOM_DSPARK_FUSED_MARKOV_SAMPLE
bool
1 (true)
Sample the DSpark block with a fused Triton kernel that computes the rank-r bias GEMV, adds the base logits in the GEMM epilogue and reduces to token ids in registers — so W2 stays bf16 and is read exactly once per block position, and no [B, V] intermediate exists. Covers both native DSpark block samplers, Kimi-K3 (r=256) and DeepSeek-V4 (r=512); the op is shape-generic and hands anything it cannot index back to the reference, but only K3 has been run on hardware. Tie-breaking matches torch.argmax (lowest index). The bias moves from an fp32 matmul to bf16 MFMA with an fp32 accumulator: every product is exact in fp32 either way, so the result is equal to the reference up to accumulation order. Measured on Kimi-K3 (MI355X, TP8, fp8 KV, full GSM8K 5-shot at 64 concurrency): acceptance 87.08% against 87.06% unfused with the accept-length distribution equal to within 0.1pp, and flexible-extract inside the run-to-run band. Saves 145 µs per drafting step at B=1 and ~235 µs at B=64. Set to 0 to force the reference spelling if an acceptance-rate regression is suspected — the two paths are not bit-identical by construction, so this is the fastest way to rule the kernel in or out. Read at Markov-head construction, so set it before the server starts.
Qwen3 style
Variable
Type
Default
Description
ATOM_ENABLE_QK_NORM_ROPE_CACHE_QUANT_FUSION
bool
0 (false)
If set to 1, fuse QK norm, RoPE, and cache quantization into one kernel for Qwen3 dense and MoE models.
If set to 1, use Triton kernel to fuse SiLU and mul with quantization in MLP module.
Draft CUDAGraphs (all drafter flavors)
A drafter declares its forward passes as DraftGraphs (atom/spec_decode/drafter.py).
At the end of CUDAGraph capture the runner runs each one once per captured batch
size, so the per-shape JIT — aiter's flydsl builds an hgemm per tile config,
in-process — is paid at startup instead of stalling a serving step. At serve
time a pass runs at the batch the target just ran, which ForwardMode.decide
picks out of those same capture_sizes — that is what makes a warmed shape and a
reachable shape one set rather than two lists that drift. A pass must support capture; the switch below then decides whether its warmup also records. V4.1 DSpark currently warms and drafts eagerly because its request windows and expert dispatch are outside the whole-block capture contract. Its target supports PIECEWISE tensor-stage graphs.
Variable
Type
Default
Description
ATOM_DRAFT_CUDAGRAPH
bool
1 (true)
Capture each declared draft pass into a per-capture_sizes CUDAGraph as it is warmed, so a draft pass replays instead of relaunching every kernel. 0 keeps the warmup (and therefore the JIT saving) but drafts eagerly. Only declared passes with capture support are recorded; every drafter ATOM ships declares at least one pass, including the separate-draft Kimi-K3 path, whose block pass builds its paged metadata at warmup so nothing host-side is left inside the recording. EPLB no longer declines the padding: the target pads on every cudagraph decode step and its rows reach the same expert-load recorder, so declining on the draft protected nothing. A DP-sync dummy DOES replay, in lockstep with the ranks holding work — is_dummy_run is per-rank, so gating on it splits one DP group across two collectives. Measured on V4-Flash-DSpark tp1: GSM8K 0.9527 / acceptance 65.25% captured against 0.9497 / 65.21% eager, i.e. indistinguishable; on tp4 with the LM head inside the capture, draft kernel launches went 30 → 0 per pass and draft wall time 915.8 → 118.9 µs. Read per pass at warmup time, so set it before the server starts. Grep a trace for a trailing graph in a propose_* label to confirm which passes replayed.
DSpark drafting
The Kimi-K3 DSpark draft writes the target's context rows into its own paged MLA
cache once per draft layer per drafting step; the switch below shortens that
path. The first write of each process logs which path it took, and logs again if
that ever changes, so a fusion left inert by an unrecognised layout says so.
Variable
Type
Default
Description
ATOM_DSPARK_FUSED_CTX_KV
bool
1 (true)
Write the context rows with one Triton kernel (RMSNorm + RoPE + concat + paged store) instead of four launches plus a throwaway empty_like for the RoPE's query side. Falls back per call when the cache layout or the RoPE is not the plain one the kernel understands (seg / shuffled-KV layouts keep their own write kernels), and until the RoPE's cos/sin cache has reached the device. Measured on Kimi-K3 (MI355X, TP8, fp8 KV): one 4.65 µs kernel replaces a 14 µs three-kernel chain, saving ~39 µs per drafting step at B=1 and ~36 µs at B=64. Set to 0 to force the per-op chain; that chain is the fallback above rather than debug code, so it stays reachable either way (it runs the first write of every layer).
ATOM_DSPARK_DISABLE_COMPILE
bool
0 (false)
Run the DSpark draft eager while the target stays compiled. Prefer it over --level 0, which drops compilation for both models; --enforce-eager does not reach it, because support_torch_compile keys off compilation_config.level alone. Flips the decorator's own bypass rather than handing the draft a cloned config, so the shared static_forward_context registry stays one object.
Speculative acceptance
Variable
Type
Default
Description
ATOM_ENABLE_RELAXED_MTP
bool
0 (false)
Accept a draft token when it lands in the target's top 10 within 0.6 of the top logit, instead of requiring the argmax. Intended for quantized MTP heads, whose drafts are right about the region and wrong about the exact winner often enough that strict acceptance throws away usable tokens. Read once at rejection_sampler import, so it must be set before the server starts.
Engram (DeepSeek-V4.1)
The n-gram tables are per-layer and large enough that where they live, and
whether they are rebuilt, both show up at startup. Both switches below are
all-or-nothing on purpose: a half-registered set would keep the host path for
some layers and the device path for others, which is the confusing state.
Variable
Type
Default
Description
ATOM_ENGRAM_UVA
bool
1 (true)
Page-lock this rank's shard of the hash tables in place and let a device kernel read the rows it needs across the bus, dequantizing there. No copy and no HBM for the table. 0 falls back to gathering the rows on the host, which returns the same rows but costs ~50 ms of CPU per decode step with the GPU idle behind it. Anything that would make the device path unsafe — no CUDA, more TP ranks than hash heads, a registration that will not fit — falls back on its own, so the switch is for taking the host path deliberately. The fallback is the whole TP group's: the lookup ends in an all-gather, so one rank that cannot register turns every rank around rather than leaving the others in a collective it never enters.
ATOM_ENGRAM_CACHE_DIR
path
~/.cache/atom/engram
Where the compressed-vocab table is cached between runs. The table is reproducible from the tokenizer, so this only trades startup time for disk; point it at shared storage to let several servers build it once. A truncated or stale cache is rebuilt rather than raised.
Attention side streams (DeepSeek-V4.1)
A layer's compressor reads the hidden row and its own arena state, and its
indexer reads the normed query latent; neither reads what the query projection
and the fused rope/window launch produce, so the three are branches of one
dependency graph that a single stream serializes.
Variable
Type
Default
Description
ATOM_DSV41_SIDE_STREAMS
int (0/1/2)
0
How many of a layer's branches leave the main stream. 0 — none. 1 — the compressor, on the MoE's alt_stream, waited at the scorer rather than at attention, because visible counts the index row a boundary crossed in this same forward writes, so the scorer is its first reader a whole top-k chain earlier. It borrows that stream because the MoE joins it inside its own forward, a sublayer after the compressor was joined, so the two are never live at once. 2 — the same, plus the indexer on a stream of its own; it is the only branch live beside both others, so it is the only one worth a queue. A value outside 0–2 raises rather than rounding.
Level 0 is the default because forking measured slower, not faster. Median
steady-state layer period, MI355X TP4 bf16 KV DSpark-5, 1024/1024 at
concurrency 64, ~10k sampled layers per trace:
level
layer period
vs 0
0
636.40 µs (repeat: 636.72)
—
1
645.76 µs
+1.47%
2
640.32 µs
+0.62%
The anchor is the gap between consecutive topk_gating launches, so a branch
that moves to another stream cannot drop out of the sample. The two level-0
traces were taken either side of the other two, which puts the floor — session
drift included — at 0.05%, making those deltas 12× and 29× the noise; the
unfiltered medians (604.0/605.9 against 618.8 and 607.6) order the levels the
same way. End-to-end throughput cannot resolve this: one wall-clock number per
run carries 2–6% spread, which is why an earlier A/B called the same
arrangement a wash. Whoever revisits it should start from these numbers and
from GPU_MAX_HW_QUEUES, not from where the forks sit.
V4 attention backend (Migration)
Selects between the legacy per-seq Python dispatch path in atom/models/deepseek_v4.py
and the new batched V4AttentionBackend (atom/model_ops/v4_attention_backend.py).
The new backend removes ~256 GPU→CPU .item() syncs per forward and is required
to enable CUDAGraph capture for V4. Legacy stays available during PR-A migration
for byte-equal A/B verification via dump-bisect; it is removed once all phases
land. See atom/model_ops/v4_backend_gate.py for the selector.
Variable
Type
Default
Description
ATOM_V4_BACKEND
str
legacy
legacy keeps the per-seq dispatch loop. new routes through V4AttentionBackend. Layer-restricted by ATOM_V4_BACKEND_LAYERS if set.
ATOM_V4_BACKEND_LAYERS
csv int
"" (= all)
Comma-separated layer ids that use the new backend (others stay legacy). Empty means: apply ATOM_V4_BACKEND uniformly. Used for layer-by-layer bisect during migration (e.g. 0,3,15,30).
State checkpoints
For models carrying per-request recurrent state (GDN: Qwen3-Next / Qwen3.5;
Kimi-K3's KDA; DeepSeek-V4's compressor ring), a checkpoint lets a later prefix
hit resume mid-prompt instead of recomputing from zero. Where they are placed
is a policy, set by --state-checkpoint-interval-tokens (three regimes carried
by the sign — see the configuration guide) and the
flag below. Details in the state-checkpoint section of the
scheduling & KV cache guide.
Variable
Type
Default
Description
ATOM_STATE_CHECKPOINT_DEMAND
bool
1 (true)
Set to 0 to stop a prefix hit that was refused for want of a checkpoint from placing a rung of its own, leaving the prompt-end anchor as the only placement. Overrides --state-checkpoint-demand, so the policy can be A/B'd without editing a launch script. The rung is most of the checkpoint write traffic and little of the read-back, and every write evicts something — StateSlotPool.mark_speculative carries the measurement.
LMCache offload tier
ATOM's own offload knobs are defined in atom/utils/envs.py; their behavior is
documented in atom/kv_transfer/offload/README.md. Where a
kv_connector_extra_config key also exists, it takes precedence over the env
var. LMCACHE_EC_PIN_TIMEOUT_SEC belongs to LMCache and is read only to
derive a bound.
Variable
Type
Default
Description
ATOM_KV_OFFLOAD
str
"" (off)
Enables LMCache KV offload without --kv-transfer-config, so a launcher that owns that flag for P/D can still add offload. lmcache selects the in-process lmcache_offload connector, lmcache_mp the standalone-server lmcache_mp connector (start lmcache server first). With a P/D connector in --kv-transfer-config, both run behind a multi connector. Setting it while --kv-transfer-config already names an offload connector is an error.
ATOM_KV_OFFLOAD_EXTRA_CONFIG
JSON object
""
The offload connector's kv_connector_extra_config: lmcache.<field> overrides (e.g. {"lmcache.chunk_size": 256}) and, for lmcache_mp, lmcache.mp.* options (e.g. {"lmcache.mp.port": 5556}). Requires ATOM_KV_OFFLOAD.
OFFLOAD_COPY_WORKERS
int
1
Save executor threads per offload worker. Also scales the default OFFLOAD_MAX_PENDING_SAVES.
OFFLOAD_LOAD_WORKERS
int
1
Load executor threads per offload worker. DSV4's in-process path ignores it (its SLOT load path needs a serial load executor).
OFFLOAD_MAX_PENDING_SAVES
int
unset: max(2, 2 × OFFLOAD_COPY_WORKERS) for connectors, 2 for the scheduler's state tier
Bound on running-plus-queued saves. KV and state saves share it because both pin the same pool. A non-integer raises on the connector path and warns (using 2) on the state-tier path. Overridden by max_pending_saves in kv_connector_extra_config.
OFFLOAD_MIN_LOAD_TOKENS
int
8192
Smallest external-tier hit worth loading; shorter hits are recomputed. Negative values clamp to 0.
OFFLOAD_MIN_SAVE_TOKENS
int
8192
Shortest prefix native lmcache_mp stores, as an absolute boundary for normal and late saves. With the default equal to OFFLOAD_MIN_LOAD_TOKENS, a shorter prefix could never be loaded back. Other connectors ignore it.
OFFLOAD_LOOKUP_MEMO_STEPS
int
32
Scheduler steps a memoised tier-lookup answer is replayed before the tier is asked again.
OFFLOAD_LOOKUP_RETRY_STEPS
int
32
Scheduler steps a failed tier lookup suppresses the next attempt.
OFFLOAD_PROFILE
bool
0
Emit [OFFLOAD-SAVE-PROF] / [OFFLOAD-LOAD-PROF] per-transfer records. An empty value reads as off.
OFFLOAD_SINGLE_STREAM
bool
0
Experimental: run the staging pack and copy legs on one stream.
OFFLOAD_GPU_STAGING_CHUNKS
int
derived from KV geometry
GPU staging buffer size in LMCache chunks (≥ 1).
OFFLOAD_GPU_STAGING_MAX_BYTES
int
unset
Upper bound on the GPU staging buffer in bytes; must hold at least one chunk.
OFFLOAD_RELEASE_GPU_STAGING_AFTER_TRANSFER
bool
0
Free the GPU staging buffer after each transfer instead of keeping it.
OFFLOAD_SLOT_STAGING_SLOTS
int
1
DSV4 in-process SLOT sidecar staging rows. Overridden by slot_sidecar_staging_slots.
OFFLOAD_COMMITTED_SIDECAR_CAPACITY
int
65536
DSV4 in-process committed SLOT sidecar index capacity. Overridden by committed_sidecar_index_capacity.
OFFLOAD_PUBLICATION_TIMEOUT_S
float
5.0
DSV4 in-process wait for a saved SLOT sidecar to become visible (finite, ≥ 0).
OFFLOAD_PUBLICATION_POLL_INTERVAL_S
float
0.01
Poll period of that wait (finite, > 0).
LMCACHE_MP_TRANSFER_MODE
str
auto
lmcache_mp transfer mode: auto or lmcache_driven (engine_driven is rejected). Overridden by lmcache.mp.mp_transfer_mode.
LMCACHE_EC_PIN_TIMEOUT_SEC
float
LMCache's own (300)
LMCache's source-pin timeout. ATOM reads it only to derive the engine's save-abandon window (pin + 30s), so the two stay ordered: a lost store report is reclaimed only after LMCache would already have force-unpinned its source. Non-positive disables ATOM's reclamation. ATOM sets no default of its own and, when unset, assumes LMCache's.
KV cache events
The scheduler can publish prefix-cache changes (BlockStored, BlockRemoved,
AllBlocksCleared, and BlockStored(medium=REMOTE) for KV received from a
PD producer) over ZMQ so external routers and cache managers can mirror what
each engine holds. Every batch carries a monotonic 8-byte sequence number;
a consumer that sees a gap can ask for the missed batches over the optional
replay socket. Events are advisory and never stall inference: the in-process
queue drops the oldest batch when full. These variables are the only way to
configure the feature today: they build KVEventsConfig (see atom/config.py)
and there is no CLI flag.
Endpoint rules: each publisher binds its PUB and replay endpoints offset by its
data-parallel rank (see ATOM_KV_EVENTS_ENDPOINT), so with tcp:// the two
configured ports must be at least data_parallel_size apart or rank N's PUB
lands on rank N-1's replay port. Under pipeline parallelism only the head stage
publishes; downstream stages bind nothing. Prefill/decode disaggregation runs
separate engine processes, and only decode publishes; if other engines share
the host, give each its own endpoints.
Variable
Type
Default
Description
ATOM_KV_EVENTS_ENABLE
bool
0 (false)
Set to 1 to publish KV cache events.
ATOM_KV_EVENTS_PUBLISHER
str
zmq
zmq or null (accepts events and discards them).
ATOM_KV_EVENTS_ENDPOINT
str
tcp://127.0.0.1:5557
ZMQ PUB bind address. Under data parallelism every rank binds its own socket: tcp:// endpoints get the DP rank added to the port, ipc:///inproc:// endpoints get a _dp<rank> suffix. Rank 0 uses the configured value unchanged.
ATOM_KV_EVENTS_TOPIC
str
""
Subscription topic prefix sent as the first frame of every message.
ATOM_KV_EVENTS_HWM
int
0
ZMQ high-water mark on the PUB socket (0 = unlimited).
ATOM_KV_EVENTS_BUFFER_STEPS
int
10000
Depth of the in-process queue between the scheduler and the sender thread. When full, the oldest batch is dropped and counted in the publisher's dropped stat; the dropped batch still consumes a sequence number, so subscribers see the loss as a gap.
ATOM_KV_EVENTS_REPLAY_ENDPOINT
str
""
ZMQ ROUTER bind address for replay requests. Empty disables replay (PUB-only). A consumer sends an 8-byte big-endian start sequence and receives every retained batch with seq >= start, followed by a REPLAY_DONE terminal frame carrying the retained [oldest, latest] window. Offset per DP rank the same way as the PUB endpoint. Replay is serviced on the sender thread with non-blocking sends: a client that stops reading has its replay abandoned (counted in replay_aborted) rather than stalling live publication.
ATOM_KV_EVENTS_REPLAY_BUFFER_STEPS
int
10000
Number of most recently sent batches retained for replay. Independent of ATOM_KV_EVENTS_BUFFER_STEPS; each entry holds an encoded payload including token ids, so size it against the event rate and memory budget. Must be >= 1 when replay is enabled.
Profiling & debugging
Variable
Type
Default
Description
ATOM_METRICS_UPDATE_INTERVAL_S
float
1.0
Shared interval in seconds for ordinary/DP/PP engine metrics pushes and API snapshot refresh. Must be finite and positive; read when each loop starts, so set it before starting every service process. Prometheus scraping is configured independently. Does not cache rendered /metrics responses or change when histogram observations are recorded.
ATOM_ENABLE_METRICS_DEVICE_TIMER
bool
0 (false)
Set to 1 before starting the service to collect GPU forward duration and cumulative request prefill GPU time. Uses CUDA/HIP events, a reusable pool capped at 256 pending pairs, and FIFO polling that stops at the first incomplete event. Adds event recording and query overhead; disabled services emit no GPU timing samples. Agentic dashboard CI explicitly enables it.
ATOM_TORCH_PROFILER_DIR
str
—
When set, enables PyTorch profiler and writes traces to this directory. Create subdirectories per rank (e.g., rank_0, dp0_tp0).
ATOM_PROFILER_MORE
bool
0 (false)
When ATOM_TORCH_PROFILER_DIR is set and this is 1, enables detailed profiling: record_shapes, with_stack, and profile_memory. Applies to both the run-phase profiler and the CUDA-graph capture profiler.
ATOM_ENABLE_DETAILED_ANNOTATION
bool
0 (false)
When profiling is active, appends detailed attention aggregates to the prefill[]/decode[] trace labels: sqsq (Σ N_Q²), sqsk (Σ N_Q·N_KV), and sk (Σ N_KV), where N_Q is the scheduled query tokens and N_KV the KV length per request. Used to estimate attention FLOPs for downstream roofline analysis.
ATOM_LOG_MORE
bool
0 (false)
If set to 1, use verbose logging format (includes process name, PID, path, line number, function name).
Garbage collection
CPython's generation-2 pass is stop-the-world and walks every tracked
container, so its cost tracks the live heap — which in a serving process is
almost entirely startup state (model, compiled graph, tokenizer, KV block
pool) that is never garbage. Measured on DeepSeek-V4-Flash-DSpark tp1: 242.8 ms
in the EngineCore, up to 596 ms in a ModelRunner worker, while reclaiming zero
objects once startup was done. See atom/utils/gc_utils.py.
Variable
Type
Default
Description
ATOM_GC_FREEZE
bool
1 (true)
Move the startup heap into CPython's permanent generation once warmup is done, so collections stop scanning it. Applied in every process that outlives startup — the API server, the atomesh frontend, every EngineCore and every ModelRunner worker; undone on engine shutdown so an in-process teardown does not leak. Set 0 to keep the pre-freeze behaviour.
ATOM_GC_DEBUG
bool
0 (false)
Log every collection: generation, duration, objects reclaimed, objects tracked. Costly — counting the tracked set on every pass added ~90s of startup on a V4-Flash tp1 — but the only way to see these pauses, since a stall in the EngineCore idles the workers with no event in their torch trace.
ATOM_GC_THRESHOLD
csv int
"" (= CPython default 700,10,10)
t0,t1,t2 for gc.set_threshold(). Thresholds are per-interpreter, so each process reads it independently; anything that is not three integers is logged and ignored, applying nothing. Raising these does not make a pass cheaper, it makes passes rarer — the same total scan lands in fewer, longer stop-the-world pauses, which is a trade against tail latency and not measured here. It is also not uniform across processes: at concurrency 4096 the API server's collector ran 13,956 times in twenty minutes over a set that grew to 688,646 objects and reclaimed zero, while each ModelRunner worker reclaimed thousands per pass, where spacing collections out defers real work. Read atom:gc_collected_total for the process you mean to tune before setting this — and note that only the API server exports it, so a worker has to be read with ATOM_GC_DEBUG=1.
Raising a process's thresholds is free only while its collector keeps finding
nothing to free, which is a property of that one process, so it is exported
rather than assumed. All three are wired into the API server only; the engine
and worker processes serve no /metrics, so ATOM_GC_DEBUG=1 is what reads
them there.
atom:gc_collected_total (/metrics, per generation — prometheus_client
appends the _total) is the invariant as a series. Flat after startup is the
expected shape; a rising line means the process has started building reference
cycles, and spacing its collections out would defer real work into a growing
heap. atom:gc_collections_total, atom:gc_uncollectable_total and
atom:gc_threshold sit beside it for context. All are O(1) reads taken at
scrape time, which is a bound and not a preference: /metrics renders on the
loop that delivers every stream. The frozen count is deliberately not
here — gc.get_freeze_count() walks the permanent generation (11.9 ms at
430k frozen, the cost freezing exists to remove) for a number that changes
twice in a process's life. The startup log has it, and so does the census.
reclaim_watch logs one warning — once, not per check — if that line
ever rises, because a counter nobody looks at is not a safeguard. It sees
only what a collection reclaimed, so cyclic garbage that reaches gen-2 before
it dies is invisible to it where gen-2 passes are rare; the gen-2 size in
/debug/gc_census is what shows that.
GET /debug/gc_census breaks the scanned set down by type and by owning
library. Unlike the metrics it walks every tracked object (~1s at a million),
so it is asked for, never scraped, and it runs in a worker thread rather than
on the loop that delivers the streams. It reports counts only: naming what a
container holds would mean serialising the keys of parsed request bodies
into an unauthenticated response. top and types_per_owner bound the two
breakdowns.
Incremental detokenizer
Streaming decodes each delta from two tokenizer.decode calls that share a
window start, so subtracting one from the other isolates the new text without
emitting a half-formed UTF-8 character. These two settings govern the shortcut
that removes one of those calls. Nothing here is related to the KV prefix
cache, which is what "cache" means everywhere else in this repository.
Variable
Type
Default
Description
ATOM_DETOKENIZER_DELTA_REUSE
auto | on | off
auto
Whether the incremental detokenizer may reuse the delta it last emitted in place of one of its two tokenizer.decode calls per update. The two decodes share a window start so that subtracting one from the other isolates the new text; the first one only ever yields a length, and its span is what the previous call already emitted. Measured on DeepSeek-V4-Pro: 1.74x at one token per update, 1.40x at sixty-four. Holds only where decoding a token span does not depend on where the window started — true for the byte-level BPE tokenizers measured (DeepSeek, Qwen3.5, GLM-5.2), not for a SentencePiece-style decoder that adds or strips a leading space by position. auto therefore verifies it at startup (~15 ms) instead of assuming it from a class name, and leaves it off on any failure. Case-insensitive; an unrecognised value is logged and leaves reuse off, since off is what someone spelling this setting is reaching for. Unrelated to the KV prefix cache.
ATOM_DETOKENIZER_AUDIT_EVERY
int
1000
How often a reused delta is checked against a real decode once reuse is on. A mismatch answers that call from the decode, then turns reuse off for the process and logs once. Only audited calls are checked, so at the default up to 999 deltas can ship between a tokenizer starting to disagree and the audit that notices it; 1 checks every update, which is what the startup probe runs at. An unusable value — 0, negative, or not a number — is logged and the default used; this is not how reuse is turned off, ATOM_DETOKENIZER_DELTA_REUSE=off is.
Debug dump (atom.utils.debug_helper)
Env-gated dump / compare / monkey-patch primitives for forward bisect &
batch invariance investigation. All entries are no-op when their
controlling *_DIR is unset, so they are safe to leave wired into
production paths. See .claude/skills/dump-bisect-debug.md for the
methodology and atom/utils/debug_helper/ for the implementation.
Variable
Type
Default
Description
ATOM_FWD_DUMP_DIR
str
—
Enables install_block_forward_hooks. Per-Block hidden state is saved to {DIR}/layer{LL}_{Cls}_rank{R}[_call{NNN}].pt.
ATOM_FWD_DUMP_LAYERS
csv int
"" (= all)
Comma-separated layer ids to dump (e.g. 0,5,15,30). Empty string means dump every layer.
ATOM_FWD_DUMP_BLOCK_CLASS
csv str
Block
Module class names to hook. Multiple values supported (e.g. Block,DeepseekV4Attention,MoE,Compressor,Indexer) for sub-stage bisect. Override per model.
ATOM_FWD_DUMP_LAYER_ATTR
str
layer_id
Attribute name on the block carrying its index. Some non-DeepSeek models use layer_idx.
ATOM_FWD_DUMP_ONE_SHOT
bool
1 (true)
When 1, only the first call per layer is dumped (typical: warmup). Set to 0 to enumerate every call (_call000.pt, _call001.pt, …) — required when bisecting per-seq dispatch loops.
ATOM_WEIGHT_DUMP_DIR
str
—
Enables maybe_dump_weights_and_exit. Per-rank params + buffers for selected layers dumped to {DIR}/weight_rank{R}_layer{L}.pt. Skips .experts.* (FP4 packed).
ATOM_WEIGHT_DUMP_LAYERS
csv int
0
Comma-separated layer ids to dump weights for.
ATOM_WEIGHT_DUMP_EXIT
bool
1 (true)
When 1 (default), call sys.exit(0) after dumping. Set to 0 to continue inference after dump.
ATOM_DEBUG_TOPK
int
0
Set to K > 0 to log top-K logits per row from Sampler.forward via maybe_log_topk(). Only rank 0 writes.
ATOM_DEBUG_TOPK_PATH
str
—
Optional output file for top-K logs. Writes to stderr if unset.
CLI for comparing dumps:
python -m atom.utils.debug_helper.compare slot-invariance --dir DIR --n-slots 4
python -m atom.utils.debug_helper.compare ref-vs-target --dir DIR
python -m atom.utils.debug_helper.compare layer-bisect --dir DIR --threshold 0.99
python -m atom.utils.debug_helper.compare schema --a A.pt --b B.pt
Benchmarks (optional)
Variable
Type
Default
Description
OPENAI_API_KEY
str
—
API key for OpenAI-compatible benchmark requests.
VLLM_USE_MODELSCOPE
bool
false
If set to true, use ModelScope for model downloads in benchmarks.
SAVE_TO_PYTORCH_BENCHMARK_FORMAT
bool
false
If set, save benchmark results in PyTorch benchmark format.
Internal / Set by ATOM
The following variables are set internally by ATOM; users typically do not need to configure them:
Variable
Description
AITER_QUICK_REDUCE_QUANTIZATION
Set to INT4 for Llama models with bf16/fp16.
TORCHINDUCTOR_CACHE_DIR
Set by compiler interface for inductor cache.
TRITON_CACHE_DIR
Set by compiler interface for Triton cache.
Reference
Environment variables are defined and accessed via atom.utils.envs:
fromatom.utilsimportenvs# Example: check data parallel sizedp_size=envs.ATOM_DP_SIZE
See atom/utils/envs.py for the full list of lazy-evaluated environment variables.
Mooncake PD matched rails
Variable
Type
Default
Description
ATOM_MOONCAKE_MATCHED_RAILS
str
""
Use auto to discover ACTIVE local HCAs in the primary HCA's numbered name family, or set a comma-separated allowlist. Enables a lazy single-HCA engine pool on P, selected by D's advertised HCA name. Requires RDMA, a single primary HCA, and corresponding same-name rails across hosts. Unset preserves existing behavior.
See Mooncake matched rails for independent P/D
rank configuration, deployment requirements, and registration lifetime.