deanonymizer is a command-line system for defensive OSINT exposure measurement. It estimates re-identification risk from public Reddit and Hacker News corpora by aggregating weak signals, scoring identity hypotheses, and emitting evidence-linked remediation guidance.
The design follows the inference setting discussed in:
Operational premise: low-entropy disclosures that appear non-identifying in isolation may become identifying under cross-post and cross-platform fusion.
Given a subject handle set H and public artifact set D, produce a risk report R containing:
- identity-relevant feature extractions
- evidence-backed linkage claims
- calibrated confidence labels
- prioritized mitigation actions
- Observer model: passive adversary with access to publicly available text and metadata only
- Data boundary: no private APIs, credentialed access, or hidden datasets
- Attack primitive: probabilistic entity linkage via feature composition
- Security goal: minimize attributable identity surface from public traces
- Acquisition
- Reddit artifacts from Arctic Shift API
- Hacker News artifacts from HN Algolia Search API
- Canonicalization
- Heterogeneous source records mapped into a unified item schema
- Temporal and textual normalization for bounded-context inference
- Feature extraction and attribution
- Detection of location, affiliation, temporal routine, self-disclosed demographics, cross-platform handles, external URLs, and stylometric cues
- Attribution binding from claim to quote-level evidence and permalink
- Risk synthesis
- Confidence-calibrated findings: low, medium, high
- Explicit exact-user section and public proof URL set
- Finding-level remediation recommendations
- Human-readable report with ranked findings and rationale
- JSON serialization for longitudinal tracking and downstream analytics
- Optional strict validation: fail if no external proof URL exists beyond audited platform profile endpoints
npm install
export ANTHROPIC_API_KEY=sk-ant-...
# default model is the fast claude-haiku-4-5
# optional: export ANTHROPIC_MODEL=claude-sonnet-4-6 # slower, higher quality# Reddit only
npm run audit -- my_reddit_handle
# Reddit + Hacker News
npm run audit -- my_reddit_handle --hn my_hn_handle
# Hacker News only
npm run audit -- --hn my_hn_handle
# JSON output
npm run audit -- my_reddit_handle --json -o report.json
# Strict proof validation
npm run audit -- my_reddit_handle --require-external-proof
# Faster wall-clock analysis (parallel chunk workers)
npm run audit -- my_reddit_handle --concurrency 3| Flag | Default | Description |
|---|---|---|
| [reddit-username] / --reddit | none | Reddit user to audit (accepts u/name) |
| --hn | none | Hacker News user to audit |
| -n, --max | 300 | Maximum items fetched per platform |
| --max-chars | 120000 | Maximum analysis transcript budget |
| --concurrency | all (≤8) | Number of chunk workers processed in parallel |
| --json | false | Emit JSON instead of text report |
| --require-external-proof | false | Fail if no proof URL exists beyond audited profile pages |
| -o, --out | stdout | Write output to file |
| --i-am-authorized | false | Skip interactive authorization prompt for scripted runs |
- Increase -n to expand retrieval depth
- Increase --max-chars to reduce context truncation
- Pin ANTHROPIC_MODEL to control inference backend variance
- Store JSON outputs for temporal diff and regression analysis
npm run build- Findings are probabilistic and should not be interpreted as identity proof
- Recall is upper-bounded by source completeness and truncation constraints
- Stylometric separability is population- and domain-dependent
- Confidence calibration depends on evidence density and artifact quality