Skip to content
@trapstreet

trapstreet.run

Public benchmarks for AI workflows. Compare agents, skills, models and tools on the same task — reproducible runs, public leaderboards.

Trapstreet

Find the best AI solution for your task

trapstreet.run · Docs · Discord

Whatever the job — reading documents, debugging a pipeline, reviewing code, answering domain questions — there is more than one way to do it, and picking between them is usually guesswork.

Trapstreet turns that into a measurement. A task declares its inputs and the answers it expects; any solution runs against the same cases, is scored by the same judge, and lands on a public board with its score, latency and cost. Your solution runs as a subprocess — no SDK, no instrumentation — so a Claude Code skill, a Python pipeline and a Rust binary compete on one board. Every ranked row is pinned to a public repo@commit, so anyone can re-run it.

Repositories

trapstreet-skills Start here. Three agent skills — set up the CLI, build a solution, author a task — from plain language. npx skills add trapstreet/trapstreet-skills
trap The tp CLI, if you would rather drive it yourself. uv tool install trap-cli
trapstreet-tasks 36 reference tasks with judges and gold cases, plus what they have measured so far.

One result, as an example

On pdf-mixed-scan — four ways to read the same half-scanned PDF, same model throughout — the best parser scores 0.90 and handing the document straight to the model scores 0.85, for $0.026 a run against $0.68, in a fifth of the time.

The same parser that places last there ties for first on a different document. Which is the argument for measuring on your documents rather than trusting anyone's ranking, including ours.

Tasks live in their author's own repository — publish from anywhere public and register it on the site.

Popular repositories Loading

  1. trapstreet-skills trapstreet-skills Public

    Agent skills for trapstreet.run — set up the tp CLI, build solutions against an eval task, and author new tasks, from plain language. Claude Code, Cursor, Codex and 70+ more.

    Python 6

  2. trapstreet-tasks trapstreet-tasks Public

    Open evaluation tasks for AI agents, skills and tools — inputs, expected outputs and a judge, so any solution can be measured and ranked on trapstreet.run

    HTML 2

  3. trap trap Public

    tp — score any agent, skill, model or script against an eval task, then submit it to a public leaderboard. Non-invasive: your solution runs as a subprocess, nothing to import. Powers trapstreet.run.

    Python 2

  4. dsh-trapstreet dsh-trapstreet Public

    Check which DeepSeek Harness plugins actually loaded, and look up public evaluation boards on trapstreet.run

    JavaScript 1

  5. decision-layer-bench decision-layer-bench Public

    A benchmark for decision-layer models: can a small typed model stand in for an LLM classification call?

    Python 1

  6. .github .github Public

    Organization profile for trapstreet.run

Repositories

Showing 9 of 9 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…