Skip to content

Latest commit

 

History

History
499 lines (403 loc) · 52.8 KB

File metadata and controls

499 lines (403 loc) · 52.8 KB

ATOM Environment Variables

This document describes the environment variables used in the ATOM project.

Metadata H2D

Variable Type Default Description
ATOM_H2D_BACKEND str packed packed combines forward metadata into one H2D and GPU scatter per consumer group. direct copies each member separately. Both preserve source reuse gates and full cudagraph padding. Set before starting the runner. See metadata publication.

Data parallelism

Variable Type Default Description
ATOM_DP_RANK int 0 The rank ID for the current process in data parallelism.
ATOM_DP_RANK_LOCAL int 0 The local rank ID for the current process (used in SPMD mode).
ATOM_DP_SIZE int 1 Total number of data parallel ranks.
ATOM_DP_MASTER_IP str 127.0.0.1 Master IP address for DP ranks coordination.
ATOM_DP_MASTER_PORT int 29500 Master port for DP ranks coordination.
ATOM_DP_LB_REQ_EQUIV int 512 Token-equivalent decode pressure assigned to each in-flight request by least_tokens routing.
ATOM_DP_SESSION_AFFINITY bool false Load-place each new session, then keep later turns on the same prefix-cache owner. Reads X-Dynamo-Session-ID, falling back to X-Correlation-ID.

Prefill delayer (TP/DCP and DP attention)

Coalesces waiting prefills while decode continues. DP attention enables it by default through ATOM_ENABLE_PREFILL_DELAYER. For a single scheduler (DP=1, PP=1), including TP/DCP, it is opt-in: set ATOM_PREFILL_DECODE_INTERVAL above zero and keep the master switch enabled. Interval 0 leaves TP scheduling unchanged; setting only the master switch does not enable TP coalescing. On TP, the interval and coalescer are enabled together. This applies to the standard scheduler, including connector-based P/D roles. RapidServe's dedicated PrefillScheduler/DecodeScheduler do not use the delayer.

After each executed prefill, the decode interval runs before all coalescing bounds. Once it expires, fill, queue age, KV pressure, partial-prefill and stall bounds decide when to release. MAX_QUEUE_MS stops extra coalescing after that interval; it does not guarantee end-to-end TTFT. DP decisions reduce local signals across ranks to keep their phases aligned.

The local fill signal discounts HBM cache hits and uses the admission path's chunk limits. To bound CPU work, it stops probing fresh requests after their total prompt length reaches one batch budget. It reports only work found so far; it does not treat unseen requests as a full batch. A deep, cache-heavy queue can therefore release through the stall or hold bounds before reaching the fill target. If parked transfers exhaust the unreserved slots, a fresh request may signal possible work with zero estimated tokens until admission resolves its connector match. This is an estimate, not a reservation: checkpoint dependencies, connector results and resource changes during admission can still reduce a batch. Repeated probes reuse immutable prompt hashes while rechecking pool contents and resource fit. During decode protection, only the existence of queued or partial work is checked; partial-prefill hold bounds start after the interval.

Local hybrid models with state checkpointing can wait for an in-flight producer's reusable prompt-end checkpoint. The expected prefix must exceed both the HBM hit and any offload match by at least one prefill budget. P/D transfers and already-started offload loads keep their own progress paths. Deferred requests retain their relative order, while at most 16 later queue entries are examined for independent work each pass. Each request's wait expires after TTFT_MAX_TICKS scheduler passes from its first dependency wait; bypassing it does not restart that deadline. MAX_QUEUE_MS bounds coalescing, not this dependency wait: time spent queued before a producer becomes runnable does not make duplicate prefill useful. Pure-attention models do not use checkpoint waits.

Variable Type Default Description
ATOM_ENABLE_PREFILL_DELAYER bool true Master switch for the prefill coalescer.
ATOM_PREFILL_DELAYER_TARGET_FILL float 0.9 Release once accumulated pending tokens reach target_fill × max_num_batched_tokens (averaged across prefillable ranks). In (0, 1]; higher = fewer, larger prefills at some TTFT cost. Clamped to (0, 1].
ATOM_PREFILL_DELAYER_TTFT_MAX_TICKS int 200 Max consecutive scheduler ticks a held prefill waits before force-release. Values < 1 clamped to 1.
ATOM_PREFILL_DELAYER_PARTIAL_MAX_TICKS int 100 Tighter bound for a held mid-chunked-prefill (it holds allocated KV). Values < 1 clamped to 1.
ATOM_PREFILL_DELAYER_STALL_TICKS int 10 After this many consecutive non-growing ticks, release (burst ended, more won't come). Values < 1 clamped to 1.
ATOM_PREFILL_DELAYER_KV_HIGH_WATERMARK float 0.9 At/above this KV usage a prefillable rank force-releases (can't accumulate a bigger batch anyway).
ATOM_PREFILL_DELAYER_TOKEN_USAGE_LOW_WATERMARK float|"" "" (None) If set, a prefillable rank below this KV usage force-releases (GPU starving).
ATOM_PREFILL_DELAYER_MAX_QUEUE_MS float|"" "" (None) After decode protection, release coalescing when the oldest schedulable waiting prefill reaches this age since arrival. Empty disables the age guard. Checkpoint dependency waits use TTFT_MAX_TICKS. This is not a hard TTFT limit.
ATOM_PREFILL_DECODE_INTERVAL int 0 Protect this many scheduler passes after an executed prefill. On DP=1, PP=1, a positive value also enables local coalescing when the master switch is on; 0 leaves TP scheduling unchanged. On DP>1, 0 disables only the interval.
ATOM_PREFILL_DELAYER_DEBUG bool false Per-tick FIRE/HOLD debug logging.
ATOM_PREFILL_DELAYER_LOG_EVERY int 1000 Emit aggregate stats (per-exit fire counts + hold rate) every N decisions (0 disables).

Model loading

Variable Type Default Description
ATOM_DISABLE_MMAP bool false If set to true, disable memory-mapped file loading for model weights. Useful in containerized environments where mmap may cause issues.
ATOM_LOADER_NUM_THREADS int 16 Worker threads for weight loading. >1 (default 16) enables the batched parallel loader (routed expert weights staged in a CPU buffer, flushed with a single H2D copy when every routed expert of that parameter has arrived) with that many threads; set to 1 to fall back to the original sequential per-expert path. Raise on high-core hosts if loading is CPU-bound.
ATOM_LOADER_STRICT_COVERAGE bool true Fail loading when a fused MoE parameter does not receive every routed expert from the checkpoint. Set to false to downgrade to a warning and load anyway, leaving those expert slots at their init values — useful when bringing up a checkpoint known to be partial, misleading otherwise (the symptom is an accuracy drop much later).
ATOM_LOADER_PREFETCH bool true Warm the page cache by reading this rank's share of the checkpoint sequentially on a background thread, instead of leaving it to demand faults through the mmap. The fault pattern sustains ~3.2 GB/s on a local NVMe that a single sequential reader drives at 6.06 GB/s, so this is an access-pattern fix, not a queue-depth one. Measured on DeepSeek-R1 MXFP4 (350 GiB, TP=4): cold load 154s → 69s. Set to false to restore demand faulting. Also applies with ATOM_DISABLE_MMAP=true, whose whole-file reads go through the same page cache (DeepSeek-V4.1-Flash, TP=4, cold: 282–315s without it, 169–207s with it).
ATOM_LOADER_PREFETCH_THREADS int 4 Concurrent sequential readers per rank used by the prefetcher. The device saturates at ~2 streams, so raising this mostly adds contention with the loader; 0 is clamped to 1 (use ATOM_LOADER_PREFETCH=false to switch prefetching off).
ATOM_LOADER_PREFETCH_BLOCK_MB int 16 Read block size for the prefetcher, in MiB.
ATOM_LOADER_FADVISE bool false Issue posix_fadvise(SEQUENTIAL|WILLNEED) per shard before reading it. Off by default and ignored while ATOM_LOADER_PREFETCH is on: WILLNEED is a hint the kernel drops for most of a 350 GiB checkpoint, and running both makes the kernel read ahead over random-ish ranges while the prefetcher streams the same files, so the two compete for the device. Only useful with prefetching disabled.
ATOM_ONLINE_QUANT_STREAMING bool false Opt in to quantizing eligible online-quant modules as soon as their checkpoint weights are complete, then release source storage to reduce load-time peak memory. Only active with a valid online quantization config. See the streaming online quantization guide.
ATOM_ONLINE_QUANT_STREAMING_HOST_STAGING bool true Assemble streamed module weights in CPU storage before one H2D transfer. Keeps the checkpoint walk parallel; disabling it buffers loader calls and forces the checkpoint walk to one thread.
ATOM_ONLINE_QUANT_STREAMING_THREADS int 4 Tail workers for H2D, per-module quantization, and source release. More workers increase overlap and in-flight memory; 0 runs finalization inline.

Plugin mode

Variable Type Default Description
ATOM_DISABLE_VLLM_PLUGIN bool 0 (false) If set to 1, disable the vLLM plugin registration entirely.

Kernel / backend selection

Variable Type Default Description
ATOM_USE_TRITON_GEMM bool 0 (false) If set to 1, use AITER Triton FP4 weight preshuffled GEMM. Otherwise use AITER ASM FP4 weight preshuffled GEMM.
ATOM_FP8_BLOCKSCALE_USE_E8M0_SCALE bool unset (per checkpoint) E8M0 rather than FP32 128x128 FP8 block scales, for the weight and the activation quantized for it; on gfx950 E8M0 operands take the AITER FlyDSL GEMM. Unset: E8M0 on every arch but gfx942 when the checkpoint declares scale_fmt: ue8m0 (DeepSeek-V4), whose power-of-two scales E8M0 restates exactly; FP32 otherwise. 1/0 force it on/off; forcing it on for a checkpoint with non-power-of-two FP32 scales rounds them.
ATOM_GROUP32_WEIGHT_PRESHUFFLE bool 1 (true) On gfx950, (16, 16)-shuffle native FP8 32x32 group32 weights (DeepSeek-V4.1) at load and after a weight sync, so AITER's preshuffled group32 GEMM (FlyDSL) reads them; that GEMM emits BF16 only. Set to 0 to keep checkpoint bytes and the row-major group32 GEMM. V4.1's grouped wo_a is shuffled either way.
ATOM_USE_FP4_NON_SHUFFLE_TRITON_GEMM bool 0 (false) If set to 1, use AITER Triton FP4 GEMM with non-shuffled weights. Takes precedence over the FP4 preshuffled GEMM path selected by ATOM_USE_TRITON_GEMM.
ATOM_MHC_USE_BF16 bool 1 (true) Use AITER BF16 hi/lo mHC computation for attention, FFN and head. After loading, replace FP32 fn storage with mhc_shuffle_fn output; no FP32 copy is retained. Set to 0 for FP32 mHC. Takes effect at model load; restart to change modes. On gfx1250, AITER enables shuffled residuals only while its runtime mhc_fused_post_pre policy remains fused (M < 1024); larger M uses ordinary residual layout and the standalone post/pre fallback.
ATOM_USE_TRITON_MXFP4_BMM bool 0 (false) If set to 1, use FP4 BMM in MLA attention module.
ATOM_USE_FLYDSL_GATHER_KV_B_PROJ bool 1 (true) Use the FlyDSL fused gather + kv_b_proj GEMM for MLA's cached-prefix path. Covers page_size-1 fp8 (e4m3) KV with an fp8 weight on gfx950 — i.e. Kimi-K3 / DeepSeek MLA under --kv-cache-dtype fp8. Unavailable imports or unsupported tensor configurations use Triton; kernel execution errors propagate. Set 0 to force Triton.
ATOM_USE_FLYDSL_FP8_PREFILL_ATTN bool 0 (false) Enable FlyDSL FP8 MLA prefill on gfx950, including fused K concatenation with QKV quantization and direct FP8 output from FlyDSL gather where supported. AITER capability checks are cached per layer. Missing FP8 attention support or unsupported device/model configurations select BF16 attention before quantization and emit a warning once per process. FP16 activations and synthetic RoPE-padding configurations also use the normal attention path. Cached K/V reuse descales max(new_token_descale, 1e-6) * 2; larger outliers saturate. Unsupported FP8 gather configurations use BF16 gather followed by fused K/V dynamic quantization when FP8 attention is supported. Kernel execution errors propagate without retry. Added 2026-09-10.
ATOM_PA_FLYDSL bool 0 (false) Route the MHA paged decode to aiter's FlyDSL kernel (aiter #4332) instead of the gluon one, where FlyDSL's domain covers the call. Off by default: the kernel is not fully tested on every shape ATOM ships. The env is the first of two gates; the second mirrors aiter's own validation so an unsupported shape falls back to gluon rather than raising inside aiter, and the and short-circuits so nothing is evaluated when this is off. Global, not per call site: the vLLM bridge is rerouted too. Needs aiter 94dca7bc6 or later.
ATOM_PA_FLYDSL_PLAN bool 1 (true) Use aiter #5546's GPU work planner on the dense decode, built once per forward in the metadata builder. Requires ATOM_PA_FLYDSL=1 — with FlyDSL off no plan is built at all. Sizes each request's partition count from its real context length instead of splitting the batch uniformly, which is worth p90 interactivity +20.5% at concurrency 20 on the MiniMax-M3 agentic trace and nothing at concurrency 1, where there is nothing to rebalance. Kept off the two MiniMax-M3 sparse call sites, whose contexts are a fixed topk window and where it measures a net loss. Plans are only minted during cudagraph capture — aiter's planner takes the batch as a tl.constexpr, so a new value is a cold kernel specialization and a plan no captured graph can ever release — which leaves the planner inert under --enforce-eager.

GLM-5.3

Variable Type Default Description
ATOM_GLM5_KPOOL bool 1 (true) Enable the pooled sparse indexer. Setting 0 is an exact token-granular A/B only at or below index_topk; longer requests are refused.
ATOM_GLM5_FORCE_DENSE_MLA bool 0 (false) Disable sparse MLA for short-context bring-up comparisons.
ATOM_GLM5_DISABLE_FUSED_MHC bool 0 (false) Force the PyTorch mHC reference path instead of AITER's fused kernels.

MiniMax-M3

Variable Type Default Description
ATOM_MONO_ENABLE bool 1 (true) Fused per-layer decode, on by default (0 disables it); MiniMax-M3 is the only model with a mono path so far. Route decode steps of up to 16 tokens (requests × speculative query tokens; a verify's tokens run as one row each) to the fused layer kernels in atom/models/minimax_m3/mono/, Eagle3 aux hidden states included. Only a supported configuration is routed (TP4, ptpc_fp8 attention linears, fp8 KV and index cache, max_model_len up to 1M, no index-cache reuse, no TBO / DP / PP / plugin mode, FULL cudagraph or eager); everything else keeps the original model. The same entry class serves both paths.
ATOM_MONO_CHECK bool 0 (false) Debug aid for ATOM_MONO_ENABLE: every sparse layer also runs the original decoder layer on the same input (after the mono layer, so mono reads only the cache entries it inserted) and rank 0 logs the per-layer, per-token difference of the residual and of the reduced output. Run it with --enforce-eager.
ATOM_MONO_TRACE path unset Debug aid for ATOM_MONO_ENABLE: rank 0 appends every mono step's input tokens, positions, top-2 logits and greedy pick to this file (one JSON line per step), with a fingerprint of every stage (dense layers, each sparse layer's output, the cache history it reads). Two runs of the same prompt then align token by token. Run it with --enforce-eager.
ATOM_MONO_TIMELINE path prefix unset Debug aid for ATOM_MONO_ENABLE: the mono layer kernels are built with their per-phase s_memrealtime stamps, each sparse layer into its own buffer, and after 20 warm-up steps of every decode token count S the next 5 steps are saved per rank as <prefix>_r<rank>_s<S>_<i>.pt (int64 [layer][CTA][stamp], 100 MHz ticks, 0 = not reached). Run it with --enforce-eager: a graph replay runs no Python.

MoE all2all (MoRI) wire format

Both are opt-in and default to off; they only apply with DP attention + expert parallelism. They are not symmetric — FP4 dispatch only moves a quantization the MoE GEMM was going to perform anyway (it consumes FP4 activations either way, and per_1x32 is per-row, so it does not matter which rank runs it), while FP8 combine adds a quantization that would not otherwise happen, since the expert output is bf16. Treat the dispatch knob as format matching and the combine knob as a quality/throughput tradeoff.

Variable Type Default Description
ATOM_MORI_FP4_DISPATCH bool 0 (false) If set to 1, quantize activations to packed FP4 (E2M1, per_1x32) before the MoE all2all instead of sending bf16 — a quarter of the bytes on the dispatch wire — which selects EpDispatchIntraNodeKernel_fp4. MoRI picks its dispatch kernel from the dtype of the tensor handed to dispatch() but sizes its staging buffers from the config built at init, so this also switches scale_dim to hidden_dim/32 and the scale type to e8m0. All three are resolved together by mori_prepare_finalize.resolve_mori_dispatch(); never set one without the others, as a mismatch strides the staging scale buffer wrong and faults on the first real batch instead of erroring cleanly.
ATOM_MEGA_COMBINE_WIRE str bf16 MegaMoE (ATOM_MORI_V2_FUSED=1) combine wire: bf16, fp8 (mxfp8) or fp4 (mxfp4). Prefill-only: decode steps always combine in bf16, and the choice is DP-agreed so every rank reduces in the same format.
ATOM_MORI_COMBINE_QUANT str none Combine-side codec passed into the MoRI config. none returns bf16; fp8_blockwise selects EpCombineIntraNodeKernel_*_fp8bwq_*; MoRI also accepts fp8_direct_cast.

Fusion passes

RMSNorm

Variable Type Default Description
ATOM_USE_MODEL_SENSITIVE_RMSNORM bool 0 (false) If set to 1, use AITER's model-sensitive RMSNorm rounding mode. This can change numerical results and prefill performance, so it is opt-in.

TP AllReduce fusion

Variable Type Default Description
ATOM_ENABLE_ALLREDUCE_RMSNORM_FUSION bool 1 (true) If set to 1, fuse allreduce with RMSNorm in tensor parallel mode.

DeepSeek-style

Variable Type Default Description
ATOM_ENABLE_DS_INPUT_RMSNORM_QUANT_FUSION bool 1 (true) If set to 1, fuse RMSNorm with quantization.
ATOM_ENABLE_DS_QKNORM_FUSION bool 1 (true) If set to 1, use the fused Q/K RMSNorm path (fused_qk_rmsnorm) in the DeepSeek MLA attention module when Q-LoRA is enabled and QK norm+quant fusion is not used. If set to 0, apply separate RMSNorm for the Q and KV branches instead.
ATOM_ENABLE_DS_QKNORM_QUANT_FUSION bool 1 (true) If set to 1, fuse QK norm with quantization in MLA attention module.
ATOM_DUAL_STREAM_MOE_TOKEN_THRESHOLD int 1024 Upper bound on MoE token count (num_tokens in the MoE forward) for using the dual-stream path: shared experts on a secondary CUDA stream while routed experts run on the default stream. If num_tokens exceeds this value, that forward uses single-stream MoE instead. Set to 0 to disable dual-stream setup entirely (no alt stream, no maybe_dual_stream_forward registration).
ATOM_DUAL_STREAM_PIECEWISE bool 0 Opt-in: allow a PIECEWISE-captured graph piece to hold the MoE dual-stream fork/join (shared experts on alt_stream overlapping routed experts). Capture is not the obstacle — set_forward_context runs inside graph_capture(), so the main stream the fork waits on is the stream capture runs on — and vLLM and SGLang both keep this overlap on inside piecewise graphs (SGLang runs dual-stream only inside a graph). Measured on V4-Pro-DSpark under AF_PIECEWISE: the fork survives capture, hides 77.5% of shared-expert time, and leaves GSM8K and MTP acceptance unmoved. Off by default only because no throughput win has been demonstrated, and because each replayed piece then carries its own driver-allocated stream (368 vs 2 distinct streams on a tp8 rank trace). The dispatcher is shared, so this moves V2/V3.2/K3 as well. Eager (NONE) and whole-model FULL are unaffected.

DSpark block sampling

DSpark drafts a num_speculative_tokens-wide block in one backbone pass, then samples it left-to-right with a low-rank first-order Markov head (logits_k = base_logits_k + W1[x_{k-1}] @ W2ᵀ, x_k = argmax(logits_k)). The unfused loop casts the whole [V, r] W2 table to fp32 on every iteration and materializes two [B, V] fp32 tensors that only an argmax reads. See atom/model_ops/dspark_markov_sample.py.

Variable Type Default Description
ATOM_DSPARK_FUSED_MARKOV_SAMPLE bool 1 (true) Sample the DSpark block with a fused Triton kernel that computes the rank-r bias GEMV, adds the base logits in the GEMM epilogue and reduces to token ids in registers — so W2 stays bf16 and is read exactly once per block position, and no [B, V] intermediate exists. Covers both native DSpark block samplers, Kimi-K3 (r=256) and DeepSeek-V4 (r=512); the op is shape-generic and hands anything it cannot index back to the reference, but only K3 has been run on hardware. Tie-breaking matches torch.argmax (lowest index). The bias moves from an fp32 matmul to bf16 MFMA with an fp32 accumulator: every product is exact in fp32 either way, so the result is equal to the reference up to accumulation order. Measured on Kimi-K3 (MI355X, TP8, fp8 KV, full GSM8K 5-shot at 64 concurrency): acceptance 87.08% against 87.06% unfused with the accept-length distribution equal to within 0.1pp, and flexible-extract inside the run-to-run band. Saves 145 µs per drafting step at B=1 and ~235 µs at B=64. Set to 0 to force the reference spelling if an acceptance-rate regression is suspected — the two paths are not bit-identical by construction, so this is the fastest way to rule the kernel in or out. Read at Markov-head construction, so set it before the server starts.

Qwen3 style

Variable Type Default Description
ATOM_ENABLE_QK_NORM_ROPE_CACHE_QUANT_FUSION bool 0 (false) If set to 1, fuse QK norm, RoPE, and cache quantization into one kernel for Qwen3 dense and MoE models.

Llama-style

Variable Type Default Description
ATOM_LLAMA_ENABLE_AITER_TRITON_FUSED_RMSNORM_QUANT bool 1 (true) If set to 1, use Triton kernel to fuse RMSNorm with quantization.
ATOM_LLAMA_ENABLE_AITER_TRITON_FUSED_SILU_MUL_QUANT bool 1 (true) If set to 1, use Triton kernel to fuse SiLU and mul with quantization in MLP module.

Draft CUDAGraphs (all drafter flavors)

A drafter declares its forward passes as DraftGraphs (atom/spec_decode/drafter.py). At the end of CUDAGraph capture the runner runs each one once per captured batch size, so the per-shape JIT — aiter's flydsl builds an hgemm per tile config, in-process — is paid at startup instead of stalling a serving step. At serve time a pass runs at the batch the target just ran, which ForwardMode.decide picks out of those same capture_sizes — that is what makes a warmed shape and a reachable shape one set rather than two lists that drift. A pass must support capture; the switch below then decides whether its warmup also records. V4.1 DSpark currently warms and drafts eagerly because its request windows and expert dispatch are outside the whole-block capture contract. Its target supports PIECEWISE tensor-stage graphs.

Variable Type Default Description
ATOM_DRAFT_CUDAGRAPH bool 1 (true) Capture each declared draft pass into a per-capture_sizes CUDAGraph as it is warmed, so a draft pass replays instead of relaunching every kernel. 0 keeps the warmup (and therefore the JIT saving) but drafts eagerly. Only declared passes with capture support are recorded; every drafter ATOM ships declares at least one pass, including the separate-draft Kimi-K3 path, whose block pass builds its paged metadata at warmup so nothing host-side is left inside the recording. EPLB no longer declines the padding: the target pads on every cudagraph decode step and its rows reach the same expert-load recorder, so declining on the draft protected nothing. A DP-sync dummy DOES replay, in lockstep with the ranks holding work — is_dummy_run is per-rank, so gating on it splits one DP group across two collectives. Measured on V4-Flash-DSpark tp1: GSM8K 0.9527 / acceptance 65.25% captured against 0.9497 / 65.21% eager, i.e. indistinguishable; on tp4 with the LM head inside the capture, draft kernel launches went 30 → 0 per pass and draft wall time 915.8 → 118.9 µs. Read per pass at warmup time, so set it before the server starts. Grep a trace for a trailing graph in a propose_* label to confirm which passes replayed.

DSpark drafting

The Kimi-K3 DSpark draft writes the target's context rows into its own paged MLA cache once per draft layer per drafting step; the switch below shortens that path. The first write of each process logs which path it took, and logs again if that ever changes, so a fusion left inert by an unrecognised layout says so.

Variable Type Default Description
ATOM_DSPARK_FUSED_CTX_KV bool 1 (true) Write the context rows with one Triton kernel (RMSNorm + RoPE + concat + paged store) instead of four launches plus a throwaway empty_like for the RoPE's query side. Falls back per call when the cache layout or the RoPE is not the plain one the kernel understands (seg / shuffled-KV layouts keep their own write kernels), and until the RoPE's cos/sin cache has reached the device. Measured on Kimi-K3 (MI355X, TP8, fp8 KV): one 4.65 µs kernel replaces a 14 µs three-kernel chain, saving ~39 µs per drafting step at B=1 and ~36 µs at B=64. Set to 0 to force the per-op chain; that chain is the fallback above rather than debug code, so it stays reachable either way (it runs the first write of every layer).
ATOM_DSPARK_DISABLE_COMPILE bool 0 (false) Run the DSpark draft eager while the target stays compiled. Prefer it over --level 0, which drops compilation for both models; --enforce-eager does not reach it, because support_torch_compile keys off compilation_config.level alone. Flips the decorator's own bypass rather than handing the draft a cloned config, so the shared static_forward_context registry stays one object.

Speculative acceptance

Variable Type Default Description
ATOM_ENABLE_RELAXED_MTP bool 0 (false) Accept a draft token when it lands in the target's top 10 within 0.6 of the top logit, instead of requiring the argmax. Intended for quantized MTP heads, whose drafts are right about the region and wrong about the exact winner often enough that strict acceptance throws away usable tokens. Read once at rejection_sampler import, so it must be set before the server starts.

Engram (DeepSeek-V4.1)

The n-gram tables are per-layer and large enough that where they live, and whether they are rebuilt, both show up at startup. Both switches below are all-or-nothing on purpose: a half-registered set would keep the host path for some layers and the device path for others, which is the confusing state.

Variable Type Default Description
ATOM_ENGRAM_UVA bool 1 (true) Page-lock this rank's shard of the hash tables in place and let a device kernel read the rows it needs across the bus, dequantizing there. No copy and no HBM for the table. 0 falls back to gathering the rows on the host, which returns the same rows but costs ~50 ms of CPU per decode step with the GPU idle behind it. Anything that would make the device path unsafe — no CUDA, more TP ranks than hash heads, a registration that will not fit — falls back on its own, so the switch is for taking the host path deliberately. The fallback is the whole TP group's: the lookup ends in an all-gather, so one rank that cannot register turns every rank around rather than leaving the others in a collective it never enters.
ATOM_ENGRAM_CACHE_DIR path ~/.cache/atom/engram Where the compressed-vocab table is cached between runs. The table is reproducible from the tokenizer, so this only trades startup time for disk; point it at shared storage to let several servers build it once. A truncated or stale cache is rebuilt rather than raised.

Attention side streams (DeepSeek-V4.1)

A layer's compressor reads the hidden row and its own arena state, and its indexer reads the normed query latent; neither reads what the query projection and the fused rope/window launch produce, so the three are branches of one dependency graph that a single stream serializes.

Variable Type Default Description
ATOM_DSV41_SIDE_STREAMS int (0/1/2) 0 How many of a layer's branches leave the main stream. 0 — none. 1 — the compressor, on the MoE's alt_stream, waited at the scorer rather than at attention, because visible counts the index row a boundary crossed in this same forward writes, so the scorer is its first reader a whole top-k chain earlier. It borrows that stream because the MoE joins it inside its own forward, a sublayer after the compressor was joined, so the two are never live at once. 2 — the same, plus the indexer on a stream of its own; it is the only branch live beside both others, so it is the only one worth a queue. A value outside 0–2 raises rather than rounding.

Level 0 is the default because forking measured slower, not faster. Median steady-state layer period, MI355X TP4 bf16 KV DSpark-5, 1024/1024 at concurrency 64, ~10k sampled layers per trace:

level layer period vs 0
0 636.40 µs (repeat: 636.72) —
1 645.76 µs +1.47%
2 640.32 µs +0.62%

The anchor is the gap between consecutive topk_gating launches, so a branch that moves to another stream cannot drop out of the sample. The two level-0 traces were taken either side of the other two, which puts the floor — session drift included — at 0.05%, making those deltas 12× and 29× the noise; the unfiltered medians (604.0/605.9 against 618.8 and 607.6) order the levels the same way. End-to-end throughput cannot resolve this: one wall-clock number per run carries 2–6% spread, which is why an earlier A/B called the same arrangement a wash. Whoever revisits it should start from these numbers and from GPU_MAX_HW_QUEUES, not from where the forks sit.

V4 attention backend (Migration)

Selects between the legacy per-seq Python dispatch path in atom/models/deepseek_v4.py and the new batched V4AttentionBackend (atom/model_ops/v4_attention_backend.py). The new backend removes ~256 GPU→CPU .item() syncs per forward and is required to enable CUDAGraph capture for V4. Legacy stays available during PR-A migration for byte-equal A/B verification via dump-bisect; it is removed once all phases land. See atom/model_ops/v4_backend_gate.py for the selector.

Variable Type Default Description
ATOM_V4_BACKEND str legacy legacy keeps the per-seq dispatch loop. new routes through V4AttentionBackend. Layer-restricted by ATOM_V4_BACKEND_LAYERS if set.
ATOM_V4_BACKEND_LAYERS csv int "" (= all) Comma-separated layer ids that use the new backend (others stay legacy). Empty means: apply ATOM_V4_BACKEND uniformly. Used for layer-by-layer bisect during migration (e.g. 0,3,15,30).

State checkpoints

For models carrying per-request recurrent state (GDN: Qwen3-Next / Qwen3.5; Kimi-K3's KDA; DeepSeek-V4's compressor ring), a checkpoint lets a later prefix hit resume mid-prompt instead of recomputing from zero. Where they are placed is a policy, set by --state-checkpoint-interval-tokens (three regimes carried by the sign — see the configuration guide) and the flag below. Details in the state-checkpoint section of the scheduling & KV cache guide.

Variable Type Default Description
ATOM_STATE_CHECKPOINT_DEMAND bool 1 (true) Set to 0 to stop a prefix hit that was refused for want of a checkpoint from placing a rung of its own, leaving the prompt-end anchor as the only placement. Overrides --state-checkpoint-demand, so the policy can be A/B'd without editing a launch script. The rung is most of the checkpoint write traffic and little of the read-back, and every write evicts something — StateSlotPool.mark_speculative carries the measurement.

LMCache offload tier

ATOM's own offload knobs are defined in atom/utils/envs.py; their behavior is documented in atom/kv_transfer/offload/README.md. Where a kv_connector_extra_config key also exists, it takes precedence over the env var. LMCACHE_EC_PIN_TIMEOUT_SEC belongs to LMCache and is read only to derive a bound.

Variable Type Default Description
ATOM_KV_OFFLOAD str "" (off) Enables LMCache KV offload without --kv-transfer-config, so a launcher that owns that flag for P/D can still add offload. lmcache selects the in-process lmcache_offload connector, lmcache_mp the standalone-server lmcache_mp connector (start lmcache server first). With a P/D connector in --kv-transfer-config, both run behind a multi connector. Setting it while --kv-transfer-config already names an offload connector is an error.
ATOM_KV_OFFLOAD_EXTRA_CONFIG JSON object "" The offload connector's kv_connector_extra_config: lmcache.<field> overrides (e.g. {"lmcache.chunk_size": 256}) and, for lmcache_mp, lmcache.mp.* options (e.g. {"lmcache.mp.port": 5556}). Requires ATOM_KV_OFFLOAD.
OFFLOAD_COPY_WORKERS int 1 Save executor threads per offload worker. Also scales the default OFFLOAD_MAX_PENDING_SAVES.
OFFLOAD_LOAD_WORKERS int 1 Load executor threads per offload worker. DSV4's in-process path ignores it (its SLOT load path needs a serial load executor).
OFFLOAD_MAX_PENDING_SAVES int unset: max(2, 2 × OFFLOAD_COPY_WORKERS) for connectors, 2 for the scheduler's state tier Bound on running-plus-queued saves. KV and state saves share it because both pin the same pool. A non-integer raises on the connector path and warns (using 2) on the state-tier path. Overridden by max_pending_saves in kv_connector_extra_config.
OFFLOAD_MIN_LOAD_TOKENS int 8192 Smallest external-tier hit worth loading; shorter hits are recomputed. Negative values clamp to 0.
OFFLOAD_MIN_SAVE_TOKENS int 8192 Shortest prefix native lmcache_mp stores, as an absolute boundary for normal and late saves. With the default equal to OFFLOAD_MIN_LOAD_TOKENS, a shorter prefix could never be loaded back. Other connectors ignore it.
OFFLOAD_LOOKUP_MEMO_STEPS int 32 Scheduler steps a memoised tier-lookup answer is replayed before the tier is asked again.
OFFLOAD_LOOKUP_RETRY_STEPS int 32 Scheduler steps a failed tier lookup suppresses the next attempt.
OFFLOAD_PROFILE bool 0 Emit [OFFLOAD-SAVE-PROF] / [OFFLOAD-LOAD-PROF] per-transfer records. An empty value reads as off.
OFFLOAD_SINGLE_STREAM bool 0 Experimental: run the staging pack and copy legs on one stream.
OFFLOAD_GPU_STAGING_CHUNKS int derived from KV geometry GPU staging buffer size in LMCache chunks (≥ 1).
OFFLOAD_GPU_STAGING_MAX_BYTES int unset Upper bound on the GPU staging buffer in bytes; must hold at least one chunk.
OFFLOAD_RELEASE_GPU_STAGING_AFTER_TRANSFER bool 0 Free the GPU staging buffer after each transfer instead of keeping it.
OFFLOAD_SLOT_STAGING_SLOTS int 1 DSV4 in-process SLOT sidecar staging rows. Overridden by slot_sidecar_staging_slots.
OFFLOAD_COMMITTED_SIDECAR_CAPACITY int 65536 DSV4 in-process committed SLOT sidecar index capacity. Overridden by committed_sidecar_index_capacity.
OFFLOAD_PUBLICATION_TIMEOUT_S float 5.0 DSV4 in-process wait for a saved SLOT sidecar to become visible (finite, ≥ 0).
OFFLOAD_PUBLICATION_POLL_INTERVAL_S float 0.01 Poll period of that wait (finite, > 0).
LMCACHE_MP_TRANSFER_MODE str auto lmcache_mp transfer mode: auto or lmcache_driven (engine_driven is rejected). Overridden by lmcache.mp.mp_transfer_mode.
LMCACHE_EC_PIN_TIMEOUT_SEC float LMCache's own (300) LMCache's source-pin timeout. ATOM reads it only to derive the engine's save-abandon window (pin + 30s), so the two stay ordered: a lost store report is reclaimed only after LMCache would already have force-unpinned its source. Non-positive disables ATOM's reclamation. ATOM sets no default of its own and, when unset, assumes LMCache's.

KV cache events

The scheduler can publish prefix-cache changes (BlockStored, BlockRemoved, AllBlocksCleared, and BlockStored(medium=REMOTE) for KV received from a PD producer) over ZMQ so external routers and cache managers can mirror what each engine holds. Every batch carries a monotonic 8-byte sequence number; a consumer that sees a gap can ask for the missed batches over the optional replay socket. Events are advisory and never stall inference: the in-process queue drops the oldest batch when full. These variables are the only way to configure the feature today: they build KVEventsConfig (see atom/config.py) and there is no CLI flag.

Endpoint rules: each publisher binds its PUB and replay endpoints offset by its data-parallel rank (see ATOM_KV_EVENTS_ENDPOINT), so with tcp:// the two configured ports must be at least data_parallel_size apart or rank N's PUB lands on rank N-1's replay port. Under pipeline parallelism only the head stage publishes; downstream stages bind nothing. Prefill/decode disaggregation runs separate engine processes, and only decode publishes; if other engines share the host, give each its own endpoints.

Variable Type Default Description
ATOM_KV_EVENTS_ENABLE bool 0 (false) Set to 1 to publish KV cache events.
ATOM_KV_EVENTS_PUBLISHER str zmq zmq or null (accepts events and discards them).
ATOM_KV_EVENTS_ENDPOINT str tcp://127.0.0.1:5557 ZMQ PUB bind address. Under data parallelism every rank binds its own socket: tcp:// endpoints get the DP rank added to the port, ipc:///inproc:// endpoints get a _dp<rank> suffix. Rank 0 uses the configured value unchanged.
ATOM_KV_EVENTS_TOPIC str "" Subscription topic prefix sent as the first frame of every message.
ATOM_KV_EVENTS_HWM int 0 ZMQ high-water mark on the PUB socket (0 = unlimited).
ATOM_KV_EVENTS_BUFFER_STEPS int 10000 Depth of the in-process queue between the scheduler and the sender thread. When full, the oldest batch is dropped and counted in the publisher's dropped stat; the dropped batch still consumes a sequence number, so subscribers see the loss as a gap.
ATOM_KV_EVENTS_REPLAY_ENDPOINT str "" ZMQ ROUTER bind address for replay requests. Empty disables replay (PUB-only). A consumer sends an 8-byte big-endian start sequence and receives every retained batch with seq >= start, followed by a REPLAY_DONE terminal frame carrying the retained [oldest, latest] window. Offset per DP rank the same way as the PUB endpoint. Replay is serviced on the sender thread with non-blocking sends: a client that stops reading has its replay abandoned (counted in replay_aborted) rather than stalling live publication.
ATOM_KV_EVENTS_REPLAY_BUFFER_STEPS int 10000 Number of most recently sent batches retained for replay. Independent of ATOM_KV_EVENTS_BUFFER_STEPS; each entry holds an encoded payload including token ids, so size it against the event rate and memory budget. Must be >= 1 when replay is enabled.

Profiling & debugging

Variable Type Default Description
ATOM_METRICS_UPDATE_INTERVAL_S float 1.0 Shared interval in seconds for ordinary/DP/PP engine metrics pushes and API snapshot refresh. Must be finite and positive; read when each loop starts, so set it before starting every service process. Prometheus scraping is configured independently. Does not cache rendered /metrics responses or change when histogram observations are recorded.
ATOM_ENABLE_METRICS_DEVICE_TIMER bool 0 (false) Set to 1 before starting the service to collect GPU forward duration and cumulative request prefill GPU time. Uses CUDA/HIP events, a reusable pool capped at 256 pending pairs, and FIFO polling that stops at the first incomplete event. Adds event recording and query overhead; disabled services emit no GPU timing samples. Agentic dashboard CI explicitly enables it.
ATOM_TORCH_PROFILER_DIR str — When set, enables PyTorch profiler and writes traces to this directory. Create subdirectories per rank (e.g., rank_0, dp0_tp0).
ATOM_PROFILER_MORE bool 0 (false) When ATOM_TORCH_PROFILER_DIR is set and this is 1, enables detailed profiling: record_shapes, with_stack, and profile_memory. Applies to both the run-phase profiler and the CUDA-graph capture profiler.
ATOM_ENABLE_DETAILED_ANNOTATION bool 0 (false) When profiling is active, appends detailed attention aggregates to the prefill[]/decode[] trace labels: sqsq (Σ N_Q²), sqsk (Σ N_Q·N_KV), and sk (Σ N_KV), where N_Q is the scheduled query tokens and N_KV the KV length per request. Used to estimate attention FLOPs for downstream roofline analysis.
ATOM_LOG_MORE bool 0 (false) If set to 1, use verbose logging format (includes process name, PID, path, line number, function name).

Garbage collection

CPython's generation-2 pass is stop-the-world and walks every tracked container, so its cost tracks the live heap — which in a serving process is almost entirely startup state (model, compiled graph, tokenizer, KV block pool) that is never garbage. Measured on DeepSeek-V4-Flash-DSpark tp1: 242.8 ms in the EngineCore, up to 596 ms in a ModelRunner worker, while reclaiming zero objects once startup was done. See atom/utils/gc_utils.py.

Variable Type Default Description
ATOM_GC_FREEZE bool 1 (true) Move the startup heap into CPython's permanent generation once warmup is done, so collections stop scanning it. Applied in every process that outlives startup — the API server, the atomesh frontend, every EngineCore and every ModelRunner worker; undone on engine shutdown so an in-process teardown does not leak. Set 0 to keep the pre-freeze behaviour.
ATOM_GC_DEBUG bool 0 (false) Log every collection: generation, duration, objects reclaimed, objects tracked. Costly — counting the tracked set on every pass added ~90s of startup on a V4-Flash tp1 — but the only way to see these pauses, since a stall in the EngineCore idles the workers with no event in their torch trace.
ATOM_GC_THRESHOLD csv int "" (= CPython default 700,10,10) t0,t1,t2 for gc.set_threshold(). Thresholds are per-interpreter, so each process reads it independently; anything that is not three integers is logged and ignored, applying nothing. Raising these does not make a pass cheaper, it makes passes rarer — the same total scan lands in fewer, longer stop-the-world pauses, which is a trade against tail latency and not measured here. It is also not uniform across processes: at concurrency 4096 the API server's collector ran 13,956 times in twenty minutes over a set that grew to 688,646 objects and reclaimed zero, while each ModelRunner worker reclaimed thousands per pass, where spacing collections out defers real work. Read atom:gc_collected_total for the process you mean to tune before setting this — and note that only the API server exports it, so a worker has to be read with ATOM_GC_DEBUG=1.

Raising a process's thresholds is free only while its collector keeps finding nothing to free, which is a property of that one process, so it is exported rather than assumed. All three are wired into the API server only; the engine and worker processes serve no /metrics, so ATOM_GC_DEBUG=1 is what reads them there.

  • atom:gc_collected_total (/metrics, per generation — prometheus_client appends the _total) is the invariant as a series. Flat after startup is the expected shape; a rising line means the process has started building reference cycles, and spacing its collections out would defer real work into a growing heap. atom:gc_collections_total, atom:gc_uncollectable_total and atom:gc_threshold sit beside it for context. All are O(1) reads taken at scrape time, which is a bound and not a preference: /metrics renders on the loop that delivers every stream. The frozen count is deliberately not here — gc.get_freeze_count() walks the permanent generation (11.9 ms at 430k frozen, the cost freezing exists to remove) for a number that changes twice in a process's life. The startup log has it, and so does the census.
  • reclaim_watch logs one warning — once, not per check — if that line ever rises, because a counter nobody looks at is not a safeguard. It sees only what a collection reclaimed, so cyclic garbage that reaches gen-2 before it dies is invisible to it where gen-2 passes are rare; the gen-2 size in /debug/gc_census is what shows that.
  • GET /debug/gc_census breaks the scanned set down by type and by owning library. Unlike the metrics it walks every tracked object (~1s at a million), so it is asked for, never scraped, and it runs in a worker thread rather than on the loop that delivers the streams. It reports counts only: naming what a container holds would mean serialising the keys of parsed request bodies into an unauthenticated response. top and types_per_owner bound the two breakdowns.

Incremental detokenizer

Streaming decodes each delta from two tokenizer.decode calls that share a window start, so subtracting one from the other isolates the new text without emitting a half-formed UTF-8 character. These two settings govern the shortcut that removes one of those calls. Nothing here is related to the KV prefix cache, which is what "cache" means everywhere else in this repository.

Variable Type Default Description
ATOM_DETOKENIZER_DELTA_REUSE auto | on | off auto Whether the incremental detokenizer may reuse the delta it last emitted in place of one of its two tokenizer.decode calls per update. The two decodes share a window start so that subtracting one from the other isolates the new text; the first one only ever yields a length, and its span is what the previous call already emitted. Measured on DeepSeek-V4-Pro: 1.74x at one token per update, 1.40x at sixty-four. Holds only where decoding a token span does not depend on where the window started — true for the byte-level BPE tokenizers measured (DeepSeek, Qwen3.5, GLM-5.2), not for a SentencePiece-style decoder that adds or strips a leading space by position. auto therefore verifies it at startup (~15 ms) instead of assuming it from a class name, and leaves it off on any failure. Case-insensitive; an unrecognised value is logged and leaves reuse off, since off is what someone spelling this setting is reaching for. Unrelated to the KV prefix cache.
ATOM_DETOKENIZER_AUDIT_EVERY int 1000 How often a reused delta is checked against a real decode once reuse is on. A mismatch answers that call from the decode, then turns reuse off for the process and logs once. Only audited calls are checked, so at the default up to 999 deltas can ship between a tokenizer starting to disagree and the audit that notices it; 1 checks every update, which is what the startup probe runs at. An unusable value — 0, negative, or not a number — is logged and the default used; this is not how reuse is turned off, ATOM_DETOKENIZER_DELTA_REUSE=off is.

Debug dump (atom.utils.debug_helper)

Env-gated dump / compare / monkey-patch primitives for forward bisect & batch invariance investigation. All entries are no-op when their controlling *_DIR is unset, so they are safe to leave wired into production paths. See .claude/skills/dump-bisect-debug.md for the methodology and atom/utils/debug_helper/ for the implementation.

Variable Type Default Description
ATOM_FWD_DUMP_DIR str — Enables install_block_forward_hooks. Per-Block hidden state is saved to {DIR}/layer{LL}_{Cls}_rank{R}[_call{NNN}].pt.
ATOM_FWD_DUMP_LAYERS csv int "" (= all) Comma-separated layer ids to dump (e.g. 0,5,15,30). Empty string means dump every layer.
ATOM_FWD_DUMP_BLOCK_CLASS csv str Block Module class names to hook. Multiple values supported (e.g. Block,DeepseekV4Attention,MoE,Compressor,Indexer) for sub-stage bisect. Override per model.
ATOM_FWD_DUMP_LAYER_ATTR str layer_id Attribute name on the block carrying its index. Some non-DeepSeek models use layer_idx.
ATOM_FWD_DUMP_ONE_SHOT bool 1 (true) When 1, only the first call per layer is dumped (typical: warmup). Set to 0 to enumerate every call (_call000.pt, _call001.pt, …) — required when bisecting per-seq dispatch loops.
ATOM_WEIGHT_DUMP_DIR str — Enables maybe_dump_weights_and_exit. Per-rank params + buffers for selected layers dumped to {DIR}/weight_rank{R}_layer{L}.pt. Skips .experts.* (FP4 packed).
ATOM_WEIGHT_DUMP_LAYERS csv int 0 Comma-separated layer ids to dump weights for.
ATOM_WEIGHT_DUMP_EXIT bool 1 (true) When 1 (default), call sys.exit(0) after dumping. Set to 0 to continue inference after dump.
ATOM_DEBUG_TOPK int 0 Set to K > 0 to log top-K logits per row from Sampler.forward via maybe_log_topk(). Only rank 0 writes.
ATOM_DEBUG_TOPK_PATH str — Optional output file for top-K logs. Writes to stderr if unset.

CLI for comparing dumps:

python -m atom.utils.debug_helper.compare slot-invariance --dir DIR --n-slots 4
python -m atom.utils.debug_helper.compare ref-vs-target  --dir DIR
python -m atom.utils.debug_helper.compare layer-bisect   --dir DIR --threshold 0.99
python -m atom.utils.debug_helper.compare schema --a A.pt --b B.pt

Benchmarks (optional)

Variable Type Default Description
OPENAI_API_KEY str — API key for OpenAI-compatible benchmark requests.
VLLM_USE_MODELSCOPE bool false If set to true, use ModelScope for model downloads in benchmarks.
SAVE_TO_PYTORCH_BENCHMARK_FORMAT bool false If set, save benchmark results in PyTorch benchmark format.

Internal / Set by ATOM

The following variables are set internally by ATOM; users typically do not need to configure them:

Variable Description
AITER_QUICK_REDUCE_QUANTIZATION Set to INT4 for Llama models with bf16/fp16.
TORCHINDUCTOR_CACHE_DIR Set by compiler interface for inductor cache.
TRITON_CACHE_DIR Set by compiler interface for Triton cache.

Reference

Environment variables are defined and accessed via atom.utils.envs:

from atom.utils import envs

# Example: check data parallel size
dp_size = envs.ATOM_DP_SIZE

See atom/utils/envs.py for the full list of lazy-evaluated environment variables.

Mooncake PD matched rails

Variable Type Default Description
ATOM_MOONCAKE_MATCHED_RAILS str "" Use auto to discover ACTIVE local HCAs in the primary HCA's numbered name family, or set a comma-separated allowlist. Enables a lazy single-HCA engine pool on P, selected by D's advertised HCA name. Requires RDMA, a single primary HCA, and corresponding same-name rails across hosts. Unset preserves existing behavior.

See Mooncake matched rails for independent P/D rank configuration, deployment requirements, and registration lifetime.