Two independent eval suites live under faq_automation/evals/ — see
faq_automation/evals/README.md for the
full methodology (how cases are picked, check predicates, failure analysis).
This doc is just the terminal commands. All commands run from the repo root.
uv run --project faq_automation python -m faq_automation.evals.run_search_evalNo API key needed. Prints recall@k / MRR@k / hit_rate@k over the 25 retrieval challenge cases.
Plain version — runs all 61 cases on the flex tier, prints PASS/FAIL per case plus a pattern/tag failure breakdown, exits nonzero on any failure:
uv run --project faq_automation python -m faq_automation.evals.runner
uv run --project faq_automation python -m faq_automation.evals.runner --case 289 # one case only
uv run --project faq_automation python -m faq_automation.evals.runner --batch # Batch API, same price, hours not minutesNeeds OPENAI_API_KEY.
Opik-tracked version — a smaller, cheap "Friday-demo" subset (6 of the 61
cases, deterministic action_match + placement_match scoring, no judge-model
cost) run through opik.evaluate() so before/after comparisons (e.g. changing
num_results) land as comparable Experiments on Comet Opik Cloud instead of
only a terminal report. This does not run the full 61-case suite or the
per-case content-quality checks (code correctness, formatting, etc.) that
runner.py enforces — it's a fast, deterministic sanity check for prompt/
retrieval changes, not a replacement for the full suite.
source .env # OPENAI_API_KEY, OPIK_API_KEY
# one-time: push the demo cases to the Opik dataset (skips insert if already there)
OPIK_URL_OVERRIDE=https://www.comet.com/opik/api OPIK_WORKSPACE=default \
OPIK_PROJECT_NAME=faq-automation-ci \
uv run --project faq_automation python -m faq_automation.evals.opik_eval --push-dataset
# before/after
uv run --project faq_automation python -m faq_automation.evals.opik_eval \
--experiment friday-before --num-results 1
uv run --project faq_automation python -m faq_automation.evals.opik_eval \
--experiment friday-after --num-results 5(OPIK_URL_OVERRIDE / OPIK_WORKSPACE / OPIK_PROJECT_NAME only need setting
once per shell — put them in .env if you run this often. See
docs/opik.md for the full platform/workspace table and the rest of
the Opik integration — tracing, prompt library, history backfill.)
- Terminal:
runner.pyprints aPASS/FAILline per case, a summary count, and pattern/tag failure breakdowns.opik_eval.pyprints a metrics table and, at the very end,experiment '<name>' done: ...with anexperiment_urlyou can open directly. - Comet Opik Cloud: https://www.comet.com/opik → workspace → project
faq-automation-ci→ Datasets tab (faq-triage-friday) for the raw cases, or Experiments tab to select multiple runs (e.g.friday-beforevsfriday-after) and diff theiraction_match/placement_matchscores side by side.