Winnow reads what developers complain about in public and turns it into project ideas that carry their evidence — every idea quoting the threads it came from, with a link back.
A local-first Electron desktop app: React 19, TypeScript strict, one SQLite file. One runtime dependency, no server, no account, and no API key on the default provider.
Setup · Architecture · CLI · Gate · Spec · Handoff
Ask any model for "10 developer tool ideas" and you get plausible guesses about what somebody wants. Winnow starts from what somebody already said, out loud, in public:
- Toggle topics — each shows what is already in the ledger and how much unused signal is banked.
- Fetch — eight public sources (Hacker News, GitHub and GitLab issues, Stack Exchange, Lobsters, Discourse, Jira, Bugzilla), rate limited per host, stored verbatim, then read by a model that pulls out one-sentence complaints with the human's own words attached.
- Generate — complaints become candidates, a second model judges them, and the funnel below decides what is allowed into the ledger.
- Generate a spec — any idea becomes a full technical specification you can export and hand on.
The idea in the screenshot above, GITINT-006, is not a suggestion — it is a claim with a
citation. It was built from 2 harvested problems, and here is one, verbatim from a Hacker News
thread that is one click away in the app:
"It rebased wrong, stashed changes, forgot it had stashed them, merged a stash it claimed did not exist, then reset to HEAD. By the end, I had lost the code we had just worked on."
The second complaint on that card reads said 2× — two people describing the same problem in
different words, folded into one by cosine similarity rather than by matching text. That number
is the whole point: it is how an idea says how many people actually have the problem, and it is
checkable, because the threads behind it are still linked.
Ask ten times and the same six ideas come back in different clothes. Winnow treats that as an engineering problem rather than a prompting one — every candidate passes three tiers, cheapest first.
flowchart TD
C["Candidate idea<br/>from the model"] --> T1
T1{"Tier 1<br/>normalized title key<br/>UNIQUE index"}
T1 -- collides --> R1["Rejected<br/>exact restatement"]
T1 -- survives --> RETR
RETR["Retrieve neighbours<br/>FTS5 BM25, not a table scan<br/><i>never scoped by topic</i>"] --> T2
T2{"Tier 2<br/>64-bit SimHash of the<br/>core mechanism<br/>Hamming distance"}
T2 -- "within threshold" --> R2["Rejected<br/>same mechanism,<br/>different marketing"]
T2 -- survives --> T3
T3{"Tier 3<br/>embedding cosine<br/>similarity"}
T3 -- "above threshold" --> R3["Rejected<br/>same concept,<br/>no shared vocabulary"]
T3 -- survives --> LEDGER["Ledger"]
R1 --> LOG
R2 --> LOG
R3 --> LOG
LOG["dedup_decisions<br/>tier + score + what it collided with"]
Two details decide whether this works at all:
- Matching is on the core mechanism, not the title. "Branch Graveyard" and "Stale Branch Reaper" are one idea, and only the mechanism text says so.
- Candidate retrieval is never scoped by topic. The same idea arriving under a different topic is still a duplicate — scoping that query is the mistake that makes dedup quietly stop working as the topic list grows.
When a batch returns 6 instead of 10, the UI shows the 4 that were killed and what each one collided with. Full detail in Architecture.
Left: the ledger — every row survived all three tiers. Right: the runner, showing how much signal is banked per topic, so you can see what a run has to work with before spending anything.
Requires Node 22+ — Node 24 is what CI runs. No C++ toolchain: better-sqlite3 v12 ships
prebuilt binaries for both the Node and Electron ABIs.
setup.bat :: once — dependencies, icons, Ollama models if present
dev.bat :: run with hot reload
build.bat :: typecheck, test, bundle, package into release\chmod +x *.sh # if the executable bit was lost in transit
./setup.sh # once — dependencies, icon, Ollama models if present
./dev.sh # run with hot reload
./build.sh # typecheck, test, bundle, package an arm64 .dmg into release/The default provider is Ollama, so a fresh install needs no API key and makes no outbound request to any AI service. Nothing is uploaded anywhere unless you point it at a bucket you own.
The same engine runs without a window — engine.bat health, fetch, batch --verbose,
dedup --batch 1, spec GITINT-001 --out spec.md, and the storage commands. Progress goes to
stderr and results to stdout, so redirecting a spec into a file stays clean. Every command is in
docs/CLI.md; install notes, provider setup and GPU sizing are in
docs/SETUP.md.
Warning
If the app starts and nothing happens, check ELECTRON_RUN_AS_NODE. VS Code sets it in the
terminals it spawns, and any Electron binary that inherits it runs as plain Node: it exits
instantly, prints nothing, and opens no window — indistinguishable from a crash. Every launcher
in this repo clears it first. If you invoke the binary yourself, clear it too.
The ledger lives in a vault: one folder, fixed layout, wherever you put it.
<vault>/
├── config.json policy for this vault
├── db/ the live database, and nothing else
├── backups/ snapshot_2026-08-11_193000_predanger.sqlite
└── exports/ 2026-08-11_195205/ideas.csv · ideas.json · …
db/ holding nothing but the database is load-bearing rather than tidy: it makes the stale-WAL
rule mechanical, because anything else in there is by definition not ours.
| Mode | Where the ledger lives |
|---|---|
local |
Your vault. No remote is contacted for any reason. |
hybrid |
Your vault, pushed to a bucket you own after every fetch and batch, on demand, and on quit. |
cloud |
A cache with a lifecycle: pulled on open, pushed while running, removed once its contents are provably in the bucket. |
Snapshots happen without being asked — daily, and always immediately before anything irreversible. Moving the vault copies, verifies, repoints and restarts, then deletes the original on the next launch that opens the copy successfully — never at the moment of copying, because a verification made by the code that wrote the copy is not evidence that the new location works.
Important
Two machines that diverge are told, never merged. Winnow will not guess which side of a
split ledger you meant to keep, so it stops and makes you choose. Your API keys are in no backup
and no bucket either: safeStorage ciphertext is bound to your OS account and would decrypt nowhere else.
An idea that reads 9.0/10 and "10 days" is a claim. Winnow's own MLTOO-001 proposed rebuilding
torch.cuda.graph(pool=...) — which PyTorch has shipped for years — and stated a build flag that
does not exist as a hard requirement. The quality judge has an ALREADY SOLVED rule and did not
fire, because it did not know the API existed. A knowledge gap does not close with a better
prompt.
So the check moved to the moment it is worth paying for. Generating a dead-end idea costs a fraction of a cent; the loss starts when you commit days. A check running over every candidate in every batch must be cheap — and cheap is why the judge missed it. A check running on the handful of ideas you are serious about can afford to be slow and networked.
flowchart LR
A["Plan it"] --> B{"checked<br/>in the last 90 days?"}
B -- yes --> Q["Build plan"]
B -- no --> C["Heat · critic · unchecked claims<br/><i>free, from the ledger</i>"]
C --> D["Run check<br/><i>optional, one model call</i>"]
D --> E["Model NAMES a<br/>lookup-able identifier"]
E --> F["HTTP decides<br/>whether it exists"]
F -- found --> G["flagged<br/>+ link + issue title"]
F -- "looked, absent" --> H["clear"]
F -- "could not look" --> I["inconclusive<br/><i>never 'clear'</i>"]
G --> Q
H --> Q
I --> Q
The model proposes a name; HTTP decides whether it exists. Asking a model "is this already solved?" invites a confident guess. Asking for an identifier that can be looked up makes the answer mechanically checkable — measured against the live API:
| identifier | result |
|---|---|
graph_pool_handle |
14 hits in pytorch/pytorch |
--filter=blob:none |
10 hits in git/git |
CUDA caching allocator hook support |
0 hits — the invented flag |
Three properties keep it honest. It reports and never rejects — a false positive that killed a
good idea would cost more than the check saves. inconclusive is never rendered as clear,
because "the provider was down" and "nothing already does this" are opposite facts. And every
flag carries a URL and an issue title, so a finding is audited in ten seconds rather than
believed. Exported documents carry the same discipline: a ## Verification block listing what
nothing has checked, and a line saying the scores are the generating model's own.
npm run check:ecosystems proves the 19 lookup routes still resolve — it caught a renamed repo
whose REST endpoint still returned 200 while the search API it actually uses returned 422, and two
read-only mirrors that returned zero hits for every term and would have manufactured a false
all-clear. Full design in docs/SPEC-VERIFICATION.md.
Local inference via Ollama is free and nothing leaves the machine. Harvesting is free too — every
adapter reads a public API, and the optional GitHub token only raises a rate limit. That leaves the
cloud provider as the only billable component, and the figures below come from Winnow's own
llm_calls table rather than an estimate.
Reading sources is 67% of all calls and 84% of input tokens while making no judgement at all,
so models are chosen per job in Settings → Model. Measured on a live key — 48 documents read,
ideas written and judged:
| Stage | Model | Calls | Input | Output | Cost |
|---|---|---|---|---|---|
| Reading | Cheap | 8 | 31,591 | 2,216 | $0.011 |
| Writing | Strong | 1 | 11,593 | 8,766 | $0.083 |
| Judging | Strong | 1 | 1,802 | 3,056 | $0.026 |
| Total — split models | 10 | 44,986 | 14,038 | $0.120 | |
| Same run — single model | $0.173 |
Net saving 31%, and 82% on the reading stage alone. That reading figure has been measured twice: a second run on a different corpus (97 sources searched, 48 new documents read in 8 calls) came to $0.0114 against $0.0643 — 82% again. Output tokens bill at several times the input rate and writing is output-heavy, which caps the total below the reading stage and is exactly why the split hands the cheap model the stage that is almost entirely input. Monthly estimates and the sources of variance are in docs/SETUP.md.
Warning
Cost figures are indicative only. No liability is accepted. They were measured on 2026-08-12 against the provider's published rates; pricing, model behaviour, reasoning-token volume and retry counts all change without notice. The maintainers of Winnow accept no responsibility for API charges incurred through use of this software — monitoring your provider's usage, quotas and billing is yours to do.
- The renderer holds no capability. It renders text scraped from the open internet, so it has
no Node, no filesystem, no network, and a CSP of
connect-src 'self'. The preload bridge exposes one named channel per operation and no genericinvoke()escape hatch. - One runtime dependency,
better-sqlite3, and that is the list. The S3 client is 122 lines ofnode:cryptosigning SigV4 against the documented algorithm, checked against AWS's own published test vector — rather than tens of megabytes of SDK for four HTTP verbs. - 321 tests across 22 suites, run by
npm testin about a second and gated in both build scripts. The migration suite applies every migration from zero and asserts row counts, on Node's built-innode:sqlite— which proves the SQL is valid SQLite rather than proving our wrapper can execute it. - Every rejection is recorded, with the tier that made it, the score, and what it collided with. A batch that returns fewer ideas than asked can always say why.
- Every gate check is recorded too — what it found, what it showed you, and whether you built anyway. The override reason is an enum rather than a boolean precisely so "the check was wrong" and "the check was right and I built it regardless" cannot average into one meaningless number.
- Windows and macOS from one tree. Prose that genuinely differs is fenced with
<!-- platform:… -->, andscripts/release.mjsstrips one side — so a shipped copy never mentions the OS it is not for. - Unsigned, deliberately no auto-update. An unsigned update channel is remote code execution; signing comes first or it does not happen.
Also in
docs/: the pre-build gate spec, the storage spec, the model-split spec with the measurements behind the table above, seventeen feedback rules earned from real bugs, and the brand assets. Read HANDOFF.md before changing anything — it separates what is verified from what is merely written.
Built by Itamar Dahan · MIT · © 2026

