Slipstream measures JavaScript engine performance commit by commit. It builds V8 or JavaScriptCore at each commit, runs JetStream2 or JetStream3, stores every raw score in SQLite, and delivers the series to the perf database, where the analysis lives.
It is built for unattended operation: a watch daemon benchmarks new commits as
they land, and two machines can split the work, one building and both measuring
the same artifacts.
Python 3.11.4 or newer (the bus consumer needs PEP 706 tar extraction filters),
plus a local checkout of whatever engine you want to measure and of the
benchmark suites. Building V8 needs depot_tools; building JSC needs the WebKit
build scripts. The defaults are tuned for arm64 macOS, which is what the bots
run; the bundled V8 gn_args set use_remoteexec = true, so override them if
you have no RBE backend.
For the untrusted RBE setup, each V8/Chromium checkout's .gclient solution
needs these custom_vars before running gclient runhooks:
"custom_vars": {
"download_remoteexec_cfg": True,
"rbe_instance": "projects/rbe-chromium-untrusted/instances/default_instance",
},Without the download flag, Chromium's hooks skip fetching the rewrapper
configuration. gn gen then fails with a missing rewrapper_mac.cfg, even
after a successful gclient sync. RBE credentials and service access are
also required; successful GN generation alone does not verify remote execution.
uv tool install .
slipstream config > ~/.config/slipstream/config.tomlThen edit the config: it needs out_dir, a bot_name, an [engines.*] table
per engine with its src_dir, a [benchmarks.*] table per suite with its
dir, and the [[run]] entries naming what this machine measures. The config
is closed -- an unknown key is a startup error, not a comment -- so the template
is the reference for what can be set.
For upgrades, stop all running build, watch, and deliver processes before
uv tool install --force ., then restart all three. A running process keeps
its loaded code after reinstalling; leaving an old daemon running across a
schema migration can cause incompatible writes and retained SQLite locks.
On a dedicated macOS benchmark account, disable the media and photo analysis agents; they can take a core for minutes at a time next to a running engine.
JSC archives include matching WebKit frameworks, XPC helpers and MiniBrowser,
selecting commits that touch Source/JavaScriptCore. Builds target arm64e to
use the pointer-authentication ABI and JIT paths used by Safari on Apple Silicon.
Results keep the host architecture label arm64 and the existing series. The
build command and archived set are in slipstream/data/engines.toml and can be
overridden per engine in config.toml.
slipstream bench v8 109000 109100
slipstream watchbench walks a commit range, building and measuring each commit. watch is
the same thing as a daemon: it picks up where the last run left off,
measures each new commit, and persists the scores. Run slipstream deliver
separately to send them onward.
clear forgets a range so it can be measured again, which is also how you
backfill history for a [[run]] entry added after the fact.
Chrome runs use --headless=new (Chromium 112 or newer), without opening a
browser window or requiring a display. Headless mode changes the browser's
rendering environment; compare scores against a headed baseline before joining
the two into one performance series.
build and watch can be split across a pair of boxes over a shared directory
called the bus. One box runs slipstream build, which checks out each commit,
builds it, packages the engine's run_set, and publishes it. Both boxes run
slipstream watch, which unpacks published artifacts and measures them. This
keeps the two sets of numbers comparable, since they come from the same binary,
and it keeps a slow build off the critical path of the faster box.
bus status reports what each side has done and flags a [[run]] matrix that
the two boxes do not agree on. bus pause, bus resume, and bus gc handle
the rest.
deliver continuously drains bounded cycles of local results and remote SSH
spools. push performs one bounded local cycle through the same pipeline.
Both deliver scores to every entry in [[push.targets]], tracked per commit
so a cycle only sends what is new. A target is either a Spanner database
(spanner = "project/instance/database", with the DDL in
slipstream/data/spanner_schema.sql) or a spool_dir, which appends to a
sequenced log for a machine that has no route to the database. The receiving
machine configures [[relay]] sources and runs one slipstream deliver owner
for both its local DB and remote spools. The relay command is removed.
slipstream deliver --help lists the maintenance switches (explicit replay,
rebuild, legacy reconciliation); the launchd/ example keeps it running.
uv sync
uv run pytest # parallel by default; -n0 for a single process
uv run ruff check slipstream tests && uv run ruff format slipstream tests
scripts/install-hooksinstall-hooks points core.hooksPath at .githooks, which runs ruff, checks
that every source file carries its SPDX header, and scans staged content
against the pattern files in .git/private-hooks/ (untracked, one regex per
line, empty when you have no such patterns).
The Spanner tests run against the emulator and skip unless
SPANNER_EMULATOR_HOST is set.
MIT. See LICENSE.