trapstreet.run · Docs · Discord
Whatever the job — reading documents, debugging a pipeline, reviewing code, answering domain questions — there is more than one way to do it, and picking between them is usually guesswork.
Trapstreet turns that into a measurement. A task declares its inputs and the answers it
expects; any solution runs against the same cases, is scored by the same judge, and lands on a
public board with its score, latency and cost. Your solution runs as a subprocess — no SDK, no
instrumentation — so a Claude Code skill, a Python pipeline and a Rust binary compete on one
board. Every ranked row is pinned to a public repo@commit, so anyone can re-run it.
| trapstreet-skills | Start here. Three agent skills — set up the CLI, build a solution, author a task — from plain language. npx skills add trapstreet/trapstreet-skills |
| trap | The tp CLI, if you would rather drive it yourself. uv tool install trap-cli |
| trapstreet-tasks | 36 reference tasks with judges and gold cases, plus what they have measured so far. |
On pdf-mixed-scan — four ways to read the same
half-scanned PDF, same model throughout — the best parser scores 0.90 and handing the document
straight to the model scores 0.85, for $0.026 a run against $0.68, in a fifth of the time.
The same parser that places last there ties for first on a different document. Which is the argument for measuring on your documents rather than trusting anyone's ranking, including ours.
Tasks live in their author's own repository — publish from anywhere public and register it on the site.
