This repo is a slow, multi-session learning project. When coming back after a break, use this page to reload the context and run the small checks before changing code.
AGENTS.mdis the short operating contract auto-loaded every session (Claude Code reads it via theCLAUDE.md → AGENTS.mdsymlink). It points here; this page holds the detail.
- Read
PROGRESS.mdfor the current milestone, last session, open decisions, and next steps. - Skim
PLAN.mdonly if you need to re-anchor the milestone map. - If working on a concept-heavy piece, re-open the relevant learning note in
docs/learnings/and the relevant reference files inreference/ds4/.
After a break, follow the reorientation route rather than reading the session log chronologically. Note numbers are not prerequisites.
A session need not stop after one helper. Choose a coherent, runnable behavior with a useful correctness boundary (for example, greedy generation end to end). Before coding, write a short contract in the owning milestone's working doc:
| Question | Required answer |
|---|---|
| What will run? | Observable artifact and boundaries/non-goals |
| What is new? | Only concepts that change implementation or verification decisions |
| What can we assume? | Basic Rust/arrays/arithmetic; links to already-taught inference concepts |
| What will make it concrete? | One worked trace with explicit shapes, units, or state transitions |
| What could be wrong? | Plausible competing interpretation and a test input that distinguishes it |
| Where does the explanation live? | Owning milestone/note and its HTML distillation |
Introduce unfamiliar inference concepts before use. Explain enough for the reader to predict an intermediate value/shape and diagnose the test's failure; do not rederive the entire prerequisite stack. Put deeper material behind optional links. Create a learning note only for a reusable concept not already owned by a note. If an explanation is requested repeatedly, improve its owner and navigation rather than adding near-duplicates. Pause at a conceptual or verification boundary, not after each function; split an arc if it accumulates unrelated lessons.
An arc closes with working code, discriminating tests, and reconciled Markdown and HTML. Record unfinished scope honestly; a working greedy loop alone will not close M3's later sampling promise. See the proposed M3 arcs.
The supported runtime target is modern Apple Silicon/macOS. Orbs and other hosts
can edit and run CPU-only host checks, but are not authoritative and cannot
validate future Metal. Normal GitHub CI should use a macOS Apple Silicon runner
for model-free fmt, build, tests, and clippy; real-model and Metal
correctness/performance belong on the development Mac.
To send an Amp web thread to that machine, follow
mac-amp-runner.md for the one-time Mac setup, runner
command, checkout synchronization boundary, and target-Mac test sequence.
rustc --version
cargo --version
uv --version
cargo fmt --all -- --check
cargo build --locked
cargo test --locked
cargo clippy --locked --all-targets -- -D warningsThe Rust edition is set in Cargo.toml: 2024. Completed
milestones must not contain executable scaffolds or broad warning suppressions.
Test helpers as they land, then validate the complete arc before closing it.
For the asset-backed gate, run these on the development Mac (release mode keeps the intentionally naive CPU math practical):
uv run --directory scripts --frozen fetch_model.py --weights
uv run --directory scripts --frozen verify_golden.py
cargo test --locked --release -- --ignoredThe verifier checks source assets and committed fixture hashes; Cargo does not
invoke it automatically. verify_golden.py --fixtures-only needs no model files.
Python is used only to fetch model assets and generate golden reference data.
Python dependency management must use uv and pyproject.toml. Do not use
pip, requirements.txt, ad-hoc virtualenv commands, or inline script metadata
for this repo. The environment is pinned by
scripts/pyproject.toml and
scripts/uv.lock.
# Fetch tokenizer/config assets into models/qwen3-0.6b/ (ignored by git)
uv run --directory scripts --frozen fetch_model.py
# Later, fetch model weights too
uv run --directory scripts --frozen fetch_model.py --weights
# Regenerate committed tokenizer golden vectors
uv run --directory scripts --frozen gen_golden.pyThe generated oracle fixtures that tests rely on live under
tests/golden/. Scratch or bulky generated files should not
be committed.
We want dependencies to stay current, but not adopt packages immediately after publication. Use a 7-day age gate for routine updates.
Check the toolchain and edition:
rustup update stable
rustc --version
cargo --version
rg '^edition' Cargo.tomlCheck what Cargo would update within the existing semver constraints:
cargo update --dry-runCargo does not currently have a built-in --exclude-newer age gate like uv. For
Rust dependency updates, the safe manual loop is:
- run
cargo update --dry-run, - inspect the proposed package versions,
- check publish dates on crates.io,
- only then run
cargo updateorcargo update -p <crate>.
For a stricter workflow, add a small helper later that queries the crates.io API for the dry-run versions and refuses any version published less than 7 days ago.
Check what is outdated:
uv tree --directory scripts --frozen --outdateduv supports the 7-day age gate directly with --exclude-newer. Always run the
--dry-run first. If the lockfile already contains a version newer than the
cutoff, uv may propose a downgrade; review that intentionally rather than
blindly applying it.
# macOS/BSD date
cutoff=$(date -u -v-7d '+%Y-%m-%dT%H:%M:%SZ')
uv lock --directory scripts --upgrade --exclude-newer "$cutoff" --dry-run
uv lock --directory scripts --upgrade --exclude-newer "$cutoff"
# Linux/GNU date equivalent
cutoff=$(date -u -d '7 days ago' '+%Y-%m-%dT%H:%M:%SZ')
uv lock --directory scripts --upgrade --exclude-newer "$cutoff" --dry-run
uv lock --directory scripts --upgrade --exclude-newer "$cutoff"Then run:
uv run --directory scripts --frozen python --version
tools/license-check.sh
cargo testThis repo is MIT licensed. Do not add GPL-family/copyleft dependencies that would complicate or pollute the license story.
Policy:
- Allowed by default: permissive licenses such as MIT, Apache-2.0, BSD, ISC, Zlib, and Unlicense.
- Prohibited by default: GPL, AGPL, and LGPL dependencies, direct or transitive.
- Manual review required: weak/file-level copyleft or unusual licenses such as MPL, EPL, CDDL, custom license text, or missing/unknown metadata.
- If a dependency is kept after manual review, document why in
PROGRESS.mdor the relevant milestone doc.
Run the license check after adding or updating dependencies:
tools/license-check.shThe check inspects transitive Rust crates via cargo metadata and Python
packages via the uv-managed scripts environment. It fails on GPL-family
licenses and warns on weak-copyleft or unknown metadata. Warnings are a prompt
for human review, not automatic approval.
Working Markdown lives in docs/*.md; the public site is hand-authored HTML in
docs/*.html plus docs/assets/. HTML is a distillation, not an automatic build.
Check whether the site has drifted from its Markdown sources with:
tools/sync-check.shAfter deliberately re-distilling a page, stamp the ledger:
tools/sync-check.sh --updateThe ledger is a commit-history reminder, not a semantic validator. It does
not inspect uncommitted prose, figures, or links. Review actual diffs and render
affected HTML, including light/dark and narrow screens. Since --update stamps
every page at HEAD, use it only after reviewing all reported sources; do not
stamp uncommitted source changes as though they were covered by that commit.
Use PENDING for a new or edited page until the reviewed changes are committed.
Learnings live as Markdown in docs/learnings/ (the source of
truth) and graduate into a dedicated Learnings section on the HTML site,
hand-distilled like every other page. Because HTML lets us do more than ASCII, a
graduated learning is the place for nicer diagrams — and sparing interactivity
in the style of diagrams.html — to illustrate the idea.
When you write or materially change a learning:
- Author/update the Markdown in
docs/learnings/NN-*.mdfirst. - Graduate it to
docs/learnings/NN-*.htmlunder the site's Learnings section (its own nav entry + index), adding diagrams/interactivity where they earn it. - Link it from the doc/milestone that references it (link the
.html, not the raw.md). - Mark changed page baselines
PENDING; re-stamp after committing and reviewing the changed sources (tools/sync-check.sh --update).
The project prefers small, understandable code and avoids dependencies that hide core inference concepts. But “zero dependencies” is not a religion.
A dependency is acceptable when it:
- handles a non-core side problem,
- improves correctness or portability,
- avoids a distracting implementation side quest,
- and is documented where the decision is made.
For example, hand-writing BPE is core to M0; hand-writing a full Unicode regex
engine is not, so M0 uses fancy-regex for Qwen's exact pre-tokenization
pattern. Small side quests are welcome when they are interesting, reasonable to
implement, and do not add much code or complexity.
Before stopping:
- Update
PROGRESS.mdwith what changed, decisions made, and the next smallest step. - Reconcile changed milestone/learning prose with its HTML and inspect the rendered result. If interrupted, record the specific outstanding site debt; do not describe that arc as fully closed.
- Leave the repo in a state where
cargo buildpasses unlessPROGRESS.mdexplicitly says otherwise.
Milestone status is shown in four hand-maintained places — bump all of them in the same commit so they don't drift:
PLAN.md— the☐ ◐ ☑legend on the milestone heading.PROGRESS.md— the Current milestone line at the top.README.md— the Status blurb and the milestone checklist.docs/index.html— the Build progress strip: flip thems--done/ms--nowclasses, theprogress-barwidth (done ÷ 8 core), and the caption.