From 984327828f4f83756c61328f91354e77e7992db0 Mon Sep 17 00:00:00 2001 From: Aswin Prakash Thiyagarajan Date: Mon, 21 Sep 2026 17:15:53 -0400 Subject: [PATCH] Vendor opik-verify: bump CANON_REF to opik-mcp main after #200 Ten skills now, not nine. opik-verify (the ship/hold verdict, opik-mcp#193) landed along with the eval fixtures and the corrected skill text (#200), so the pin moves from 0baa5ae to 8fb9c4b and `opik-verify` joins SHARED. Re-vendoring also picks up the review-driven text fixes that shipped with those PRs: `generate_report=False` in opik-compare and opik-evaluate, the graded-metric note in opik-optimize, and compatibility / allowed-tools / a pinned source_commit on opik. README skills table gains the verify row; plugin.json 0.4.0 -> 0.5.0. Co-Authored-By: Claude Fable 5.1 --- .claude-plugin/plugin.json | 4 +- README.md | 1 + scripts/sync-shared-skills.sh | 4 +- skills/opik-compare/SKILL.md | 7 +- skills/opik-evaluate/SKILL.md | 5 +- skills/opik-optimize/SKILL.md | 2 +- skills/opik-verify/SKILL.md | 165 ++++++++++++++++++++++++++++++++++ skills/opik/SKILL.md | 9 +- 8 files changed, 185 insertions(+), 12 deletions(-) create mode 100644 skills/opik-verify/SKILL.md diff --git a/.claude-plugin/plugin.json b/.claude-plugin/plugin.json index adc731b..94d2e8f 100644 --- a/.claude-plugin/plugin.json +++ b/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "opik", - "version": "0.4.0", - "description": "Opik tracing for Claude Code sessions, plus the Opik agent skills: instrument, diagnose, explain, test, compare, evaluate, online-eval, optimize", + "version": "0.5.0", + "description": "Opik tracing for Claude Code sessions, plus the Opik agent skills: instrument, diagnose, explain, test, compare, evaluate, verify, online-eval, optimize", "author": { "name": "Comet ML", "url": "https://comet.com" diff --git a/README.md b/README.md index dcce215..7979178 100644 --- a/README.md +++ b/README.md @@ -190,6 +190,7 @@ The plugin ships the Opik skill pack — the same skills published as [`opik-ski | `opik-evaluate` | "evaluate my agent", "build an eval", "write an LLM judge" | | `opik-online-eval` | "score production traces", "take this judge live" | | `opik-optimize` | "optimize this prompt", "run the prompt optimizer" | +| `opik-verify` | "is this safe to ship", "go/no-go on this change" | They are vendored at a pinned `opik-mcp` commit (see `skills/SHARED.md`); edit them upstream, not here. diff --git a/scripts/sync-shared-skills.sh b/scripts/sync-shared-skills.sh index d835fe9..edd6446 100755 --- a/scripts/sync-shared-skills.sh +++ b/scripts/sync-shared-skills.sh @@ -12,11 +12,11 @@ set -euo pipefail CANON_REPO="${CANON_REPO:-https://github.com/comet-ml/opik-mcp.git}" -CANON_REF="${CANON_REF:-0baa5aecf0224701473efd968be4aa5a166ba99f}" # opik-mcp main after #191 (nine skills) +CANON_REF="${CANON_REF:-8fb9c4b09f8854581e7b34a37049773f5fedd869}" # opik-mcp main after #200 (ten skills) SRC="src/opik_mcp/skills" DEST="skills" # Every skill the pack ships. Order matches the published index. -SHARED=(opik opik-compare opik-diagnose opik-evaluate opik-explain opik-instrument opik-online-eval opik-optimize opik-test) +SHARED=(opik opik-compare opik-diagnose opik-evaluate opik-explain opik-instrument opik-online-eval opik-optimize opik-test opik-verify) tmp="$(mktemp -d)" trap 'rm -rf "$tmp"' EXIT diff --git a/skills/opik-compare/SKILL.md b/skills/opik-compare/SKILL.md index 4f6c7fa..67bc49d 100644 --- a/skills/opik-compare/SKILL.md +++ b/skills/opik-compare/SKILL.md @@ -1,6 +1,6 @@ --- name: opik-compare -description: Run a candidate against the baseline over an Opik test suite and read the numbers back — which cases broke, which got fixed, the per-metric deltas, worst rows, and whether the two runs are comparable — with the Opik compare-view link. Runs via the SDK; reads results via the MCP when connected. Does not issue a ship/no-ship verdict. Use for "did my fix work", "compare against the baseline", "run the regression suite", "why did quality drop", "which cases regressed", "compare these two experiments". Not for live production triage (use diagnose), building an evaluation from scratch (use evaluate), or capturing a single case (use test). +description: Run a candidate against the baseline over an Opik test suite and read the numbers back — which cases broke, which got fixed, the per-metric deltas, worst rows, and whether the two runs are comparable — with the Opik compare-view link. Runs via the SDK; reads results via the MCP when connected. Does not issue a ship/no-ship verdict. Use for "did my fix work", "compare against the baseline", "run the regression suite", "why did quality drop", "which cases regressed", "compare these two experiments". Not for the ship/hold decision itself (use verify), live production triage (use diagnose), building an evaluation from scratch (use evaluate), or capturing a single case (use test). compatibility: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python or TypeScript project with Opik configured and a test suite (from the test or evaluate skill) or two existing experiments. Install the `opik` skill alongside this one — it holds the shared test-suite and experiment references; without it, this skill falls back to the public docs. allowed-tools: - Read @@ -69,6 +69,7 @@ result = opik.run_tests( experiment_name="candidate-", experiment_tags=["compare", ""], model="", + generate_report=False, # default True writes opik_test_suite_reports/ into cwd — the repo stays untouched ) candidate_id = result.experiment_id # result.experiment_url is the single-run link @@ -109,7 +110,7 @@ rows = client.get_experiments_client().find_experiment_items_for_dataset( 8. **Flaky cases** — items that flip across repeated runs of the *same* code (the suite's execution policy exposes `runs_passed`/`runs_total`). 9. **Trend** — when more than two runs exist, the pass rate across the last few, so a one-step delta has context. -Report what the numbers say. **Do not decide ship or hold** — that decision has its own policy and is a later step. +Report what the numbers say. **Do not decide ship or hold** — that is `/opik-verify`, which applies an explicit release policy to these numbers. ### 6. Link and hand off Build the compare link with **both** ids — a URL-encoded JSON array — so the user lands on the side-by-side view: @@ -118,7 +119,7 @@ import json, urllib.parse ids = urllib.parse.quote(json.dumps(["", candidate_id])) compare_url = f"{ui_base}/{workspace}/experiments/{suite.id}/compare?experiments={ids}" ``` -`ui_base` is the Opik UI origin (the configured URL minus `/api`; `result.experiment_url` shows the exact host and workspace to reuse). Then one next step (see **Output**). +`ui_base` is the Opik UI origin (the configured URL minus `/api`; `result.experiment_url` shows the exact host and workspace to reuse). Then one next step (see **Output**) — when the user's question is whether to ship, that step is `/opik-verify`. ### 7. Record (only when asked) If the user wants the finding kept beside the data: a comment on a regressed case's trace via `client.rest_client.traces.add_trace_comment(trace_id, text=…)`, or a human score beside the judge's via `client.log_traces_feedback_scores([...])`. Never by default. diff --git a/skills/opik-evaluate/SKILL.md b/skills/opik-evaluate/SKILL.md index 2a9430d..0caa40d 100644 --- a/skills/opik-evaluate/SKILL.md +++ b/skills/opik-evaluate/SKILL.md @@ -1,6 +1,6 @@ --- name: opik-evaluate -description: Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link. Picks a test suite with judge assertions or a dataset with metrics, sources cases from traces or synthetic data, scores heuristics-first then one-failure-mode judges, runs client-side via the SDK or server-side for prompt-only targets, and reads the scores back. Covers RAG evaluation, error analysis, and validating a judge against human labels. Use for "evaluate my agent", "measure quality", "build an eval", "write an LLM judge", "how good is my RAG", "set up evals for this". Not for before/after on an existing suite (use compare), one regression case (use test), or scoring production traffic (use online-eval). +description: Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link. Picks a test suite with judge assertions or a dataset with metrics, sources cases from traces or synthetic data, scores heuristics-first then one-failure-mode judges, runs client-side via the SDK or server-side for prompt-only targets, and reads the scores back. Covers RAG evaluation, error analysis, writing and validating LLM judges against human labels, and auditing an existing eval pipeline. Use for "evaluate my agent", "measure quality", "build an eval", "write an LLM judge for hallucinations", "audit our evaluation pipeline", "how good is my RAG", "set up evals for this". Not for before/after on an existing suite (use compare), one regression case (use test), scoring production traffic (use online-eval), or the ship/hold decision (use verify). compatibility: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python or TypeScript project with Opik configured. Install the `opik` skill alongside this one — it holds the shared test-suite, dataset, and metric references; without it, this skill falls back to the public docs. allowed-tools: - Read @@ -74,7 +74,8 @@ Write the task adapter as a temp file **outside the repo** (needs the app's prov ```python # Test suite results = opik.run_tests(test_suite=suite, task=lambda item: {"input": item["input"], "output": str(app(item["input"]))}, - experiment_name="baseline-", model="") + experiment_name="baseline-", model="", + generate_report=False) # default True writes opik_test_suite_reports/ into cwd — keep the repo clean # Dataset from opik.evaluation import evaluate res = evaluate(dataset=dataset, task=task, scoring_metrics=[...], experiment_name="baseline-", diff --git a/skills/opik-optimize/SKILL.md b/skills/opik-optimize/SKILL.md index abe0fb1..d374781 100644 --- a/skills/opik-optimize/SKILL.md +++ b/skills/opik-optimize/SKILL.md @@ -54,7 +54,7 @@ dataset = client.get_dataset(name="", project_name="") ### 3. Define the metric A function `(dataset_item, llm_output) -> float`, higher is better. **Give it a real name** (`def refund_answer_similarity(...)`) — its `__name__` becomes the Optimization run's objective name in the UI and `result.metric_name`; a function called `metric` shows up as "metric". -- `expected_output` present → heuristic (`LevenshteinRatio`, `Equals`, or a task-specific check) — deterministic and free. +- `expected_output` present → heuristic (`LevenshteinRatio`, `Equals`, or a task-specific check) — deterministic and free. Prefer a **graded** metric over exact match: when the baseline scores 0.0 on every item (observed with `Equals` on a strict output format), every candidate also scores 0.0 and the optimizer has nothing to climb — five trials of flat zeros is a metric problem, not a prompt problem. - Otherwise → **one** binary judge for the failure mode being optimized (`../opik-evaluate/references/write-judge-prompt.md`), wrapped to return its score `.value`. Multi-objective → `MultiMetricObjective`. Never optimize against a judge nobody validated: an unvalidated judge is the easiest thing to overfit. diff --git a/skills/opik-verify/SKILL.md b/skills/opik-verify/SKILL.md new file mode 100644 index 0000000..bcf004f --- /dev/null +++ b/skills/opik-verify/SKILL.md @@ -0,0 +1,165 @@ +--- +name: opik-verify +description: Decide ship or hold for a candidate from the compare skill's numbers, against an explicit release policy — regressions, pass rate, safety-tagged cases, subgroup consistency, latency and cost budgets, flakiness, evidence size, and whether the judge is validated. Reads two experiments on an Opik test suite via the SDK (or the MCP when connected) and returns a verdict with every criterion shown pass/fail. Use for "is this safe to ship", "can I merge this", "go/no-go on this change", "gate this release", "should I roll this out". Not for producing the numbers (use compare), building an evaluation (use evaluate), or deploying anything. +compatibility: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires Opik configured and a test suite with a baseline and a candidate experiment (from the compare skill). Install the `opik` skill alongside this one — it holds the shared test-suite and experiment references; without it, this skill falls back to the public docs. +allowed-tools: + - Read + - Grep + - Glob + - Bash + - Write +metadata: + last_updated: "2026-09-17" + source_commit: "2.0.0" + argument-hint: "[suite, or baseline and candidate experiment ids; optional --policy path]" +--- + +# Verify — Ship or Hold, Against a Policy You Can Read + +**Definition of done:** one verdict — **`ship`**, **`hold`**, **`needs_review`**, or **`insufficient_evidence`** — computed from a **declared policy** over the baseline-vs-candidate numbers, with **every criterion listed with its threshold, the observed value, and pass/fail**, the cases behind any failure named, and the compare-view link. The policy is either the repo's `opik-release-policy.yaml` or the documented defaults, and the report says which. If the two runs can't be read or aren't comparable, stop at the **first** genuine blocker and return **exactly one** next step. "Looks good to me" is not a verdict; a verdict without its criteria is not one either. + +Operate: **apply the policy mechanically, show your arithmetic, refuse to ship on a judge nobody validated, and change no application code.** The only file this skill may write is the policy file, and only when the user says so. It never deploys. + +## Inputs + +The entry point is `/opik-verify` right after `/opik-compare` (its baseline and candidate), `/opik-verify ` (the two most recent runs on the suite), or `/opik-verify `. Infer the rest; treat these as **optional overrides**: + +- policy (default: `opik-release-policy.yaml` at the repo root or under `.opik/`, else the defaults below) · which experiments (default: as above) · `--record` (default: off — write the verdict into the candidate experiment's config). + +Ask only at a genuine, non-inferable blocker (see **Blockers**). + +## The policy + +Every key is optional; missing keys take these defaults. Say in the report which source applied. + +```yaml +# opik-release-policy.yaml — repo root or .opik/. Versioned with the code so the gate is reproducible. +min_items: 10 # fewer scored items than this -> insufficient_evidence, never ship +max_regressions: 0 # pass -> fail cases allowed (flaky items excluded when flaky_policy: exclude) +pass_rate: not_below_baseline # or a number in 0..1; candidate pass rate must satisfy it +safety_tags: [safety] # a regression on an item whose data.tags contains one of these -> hold, always +subgroup_key: null # a data key (e.g. "category"); no subgroup's pass rate may fall +latency_p90_max_increase: 0.25 # candidate p90 duration vs baseline (experiments expose p50/p90/p99) +cost_per_item_max_increase: 0.25 # candidate mean cost per item vs baseline, as a fraction +flaky_policy: exclude # exclude | count — an item that flips between runs of the SAME code is flaky +judge_validated: false # set true once the suite's judge has been checked against human labels +``` + +`judge_validated: false` is the **human-review gate**: until someone has confirmed the judge agrees with people (`/opik-evaluate`'s `validate-evaluator` reference), a passing run yields `needs_review`, not `ship`. Flip it to `true` in the file once that is done — deliberately a human edit, never something this skill sets on its own. + +## Activation — the only in-scope work + +### 1. Load the policy +Look for `opik-release-policy.yaml` at the repo root, then `.opik/`. Parse it; unknown keys → **Blocker** (name the key). No file → defaults, and say so. Never invent thresholds not in the file or the defaults. + +### 2. Resolve the two runs +Take them from `/opik-compare`'s output when it just ran. Otherwise: +```python +import opik +client = opik.Opik() +runs = sorted(client.get_test_suite_experiments(name="", project_name=""), + key=lambda e: e.get_experiment_data().created_at) +baseline, candidate = runs[-2], runs[-1] # or the two ids the user gave +``` +Skip a **failed-judge run** (a run whose judge had no credential is not a candidate — `/opik-compare` explains how it happens). `scoring_failed` does not survive the read path; the read-back signal is: **every item failed and every assertion `reason` mentions a missing credential or an LLM infrastructure error**. Say which run you skipped and why. When the hosted MCP is connected, `list('experiment', name=…)` shows each run's averages and pass rate to pick from; the item-level read below stays on the SDK. + +### 3. Read both runs, item by item +An experiment holds **one item per run**: with `runs_per_item: 3` a dataset item appears three times, same `dataset_item_id`, different `trace_id`. Group — a dict keyed on `dataset_item_id` silently keeps one run and loses the counts. +```python +from collections import defaultdict +def by_item(exp): + groups = defaultdict(list) + for i in exp.get_items(): + groups[i.dataset_item_id].append(i) # each: dataset_item_data (tags / subgroup key), + return groups # assertion_results [{passed, reason}], trace_id +b, c = by_item(baseline), by_item(candidate) + +def run_passed(i): return bool(i.assertion_results) and all(a.get("passed") for a in i.assertion_results) +def counts(runs): return sum(run_passed(r) for r in runs), len(runs) # runs_passed, runs_total +thresholds = {it["id"]: (it.get("execution_policy") or suite.get_global_execution_policy() or {}).get("pass_threshold", 1) + for it in suite.get_items()} # the suite, not the experiment, holds the policy +def passed(item_id, runs): return counts(runs)[0] >= thresholds.get(item_id, 1) +# experiment level (client.rest_client.experiments.get_experiment_by_id(id)): pass_rate, +# duration (p50/p90/p99), total_estimated_cost_avg, dataset_version_id +``` +**Comparability first:** same `dataset_version_id`, same item set, same judge model (experiment config). Different → **Blocker** ("rerun the candidate on suite version X with judge Y, then `/opik-verify`") — a verdict on non-comparable runs is not a verdict. + +### 4. Evaluate every criterion, in this order +Compute all of them even after the first failure — the report shows the whole table. +1. **Evidence size** — scored items (items with assertions) ≥ `min_items`. Below → the verdict is `insufficient_evidence` regardless of the rest, **unless a gate criterion (2–7) also failed — then it is `hold`**: a known safety regression outranks thin evidence. Still report every criterion. +2. **Regressions** — items `passed` in baseline and not in candidate. An item is **flaky** when `runs_passed` is strictly between 0 and `runs_total` in either run (the counts from step 3), or it flips between two runs of the same code if you have them. Under `flaky_policy: exclude` a flaky item is dropped from the regression count and listed separately with its counts; under `count` it stays in. When every item has `runs_total == 1`, flakiness is **not observable** — report the flaky check as `not_evaluated`, state that `exclude` excluded nothing, and suggest `runs_per_item: 3` on the suite if the user wants the protection. Count ≤ `max_regressions`. +3. **Safety** — any regression whose `data.tags` intersects `safety_tags` → fail, no exceptions, no exclusions. +4. **Pass rate** — candidate `pass_rate` vs baseline, or vs the number given. +5. **Subgroups** — when `subgroup_key` is set, pass rate per value of that key must not fall. +6. **Latency** — candidate p90 duration ≤ baseline p90 × (1 + `latency_p90_max_increase`), from the experiments' `duration` percentiles (`p50`/`p90`/`p99` on the experiment record) (or per-item `duration` from the REST experiment items). +7. **Cost** — candidate mean `total_estimated_cost` per item ≤ baseline × (1 + `cost_per_item_max_increase`). Skip and say "no cost data" when neither run carries costs. + *Aggregates lag.* Right after a run finishes, the experiment record's `duration` can read `0.0` and `total_estimated_cost_avg` `None` for a few seconds while the backend aggregates (observed). A zero or missing aggregate on **one** side is not data — re-read after a short wait, or compute p90 and mean cost from the per-item `duration` / `total_estimated_cost` fields on the REST experiment items; never let a `0.0` pass or fail the gate. +8. **Evidence strength** — a paired sign test on the flips: with `f` fixes and `r` regressions, the two-sided binomial p-value under 50/50. Report it; it is **not** a gate. With `f + r < 6` say "too few flips to call it more than noise". +9. **Judge** — `judge_validated` from the policy. False → cap the verdict at `needs_review`. +10. **Attribution** — flips whose `reason` reads as judge hesitation on an unchanged output (see `/opik-compare` step 5.6) are listed for the human under `needs_review`, never silently counted either way. + +### 5. Decide +Precedence, top to bottom — the first line that applies wins: +- Any of criteria 2–7 failed → **`hold`** (even when criterion 1 also failed). +- Criterion 1 failed → **`insufficient_evidence`**. +- All gates pass but `judge_validated: false`, or attribution flagged items → **`needs_review`**, naming exactly what a person should look at. +- Otherwise → **`ship`**. + +Never round a `hold` up because the deltas are "mostly positive"; never round a `ship` down because of a hunch. The policy is the judgment; changing it is the user's move. + +### 6. Report, and record only on request +The table (criterion · threshold · observed · pass/fail), the regressions named with their assertion and trace link, the compare URL with both ids, the policy source, and one next step. With `--record`, write the verdict into the candidate experiment's config — read the existing config first and merge, `update_experiment` replaces it: +```python +exp = client.rest_client.experiments.get_experiment_by_id(candidate.id) +cfg = dict(exp.metadata or {}); cfg["verdict"] = {"status": "hold", "policy": "opik-release-policy.yaml", "failed": ["regressions"], "at": ""} +client.update_experiment(id=candidate.id, experiment_config=cfg) +``` +Offer — do not do — writing `opik-release-policy.yaml` with the defaults when no file existed, so the next verdict is reproducible. + +## Blockers + +Stop at the **earliest** blocker and return **exactly one** next step: +- "Run `opik configure`, then rerun `/opik-verify`." +- "Suite `` has fewer than two comparable runs — run `/opik-compare ` first." +- "Baseline and candidate are on different suite versions (v3 vs v4) — rerun the candidate on v3, or re-baseline on v4, then `/opik-verify`." +- "`opik-release-policy.yaml` has an unknown key `` — fix or remove it." +- "The candidate run's judge failed (every item `scoring_failed`) — set the judge's provider key and rerun `/opik-compare`." + +## Output + +**User-facing:** the verdict in one line, the criteria table, the regressions (case, assertion, why, link), the policy source, the compare link, and the single next step. Not a narrative, not JSON. + +**Underneath** (for composition / evals), one shape: +- `status`: `ship` | `hold` | `needs_review` | `insufficient_evidence` | `blocked` +- `policy`: `source` (`file` | `defaults`), `path`, `values` (the effective policy) +- `suite`: `name`, `id`, `version` +- `baseline` / `candidate`: `experiment_id`, `name`, `url`, `items`, `pass_rate` +- `criteria`: list of `{name, threshold, observed, passed, note}` — always all of them +- `regressions`: list of `{dataset_item_id, input, assertion, reason, trace_url, safety: bool, flaky: bool}` +- `flaky`: list of `{dataset_item_id, baseline_runs, candidate_runs}` excluded or counted per `flaky_policy`, or the string `not_evaluated` on a single-run suite +- `review_items`: list of `{dataset_item_id, why}` (when `needs_review`) +- `evidence`: `{items, fixes, regressions, sign_test_p}` +- `compare_url` +- `recorded`: `true|false` +- `next_step`: exactly one + +Invariants: `ship` requires every gate criterion `passed` **and** `judge_validated: true`; `hold` carries at least one failed criterion and, when the failure is regressions, a non-empty `regressions`; `insufficient_evidence` carries `criteria` with `min_items` failed; `needs_review` carries a non-empty `review_items` or `judge_validated: false` in `policy.values`; `criteria` is never partial; the codebase is never modified; nothing is deployed. + +## Examples + +**Ship.** `/opik-verify` after compare: 24 scored items, 3 fixes, 0 regressions, pass rate 0.79 → 0.92, p90 latency +4%, cost +2%, sign test p = 0.25 ("too few flips to be more than noise — but nothing regressed"), policy file present with `judge_validated: true`. → **`ship`**; next step = "merge; `/opik-online-eval` watches the refund assertion in production". + +**Hold.** Same, but the "does not give legal advice" item flipped pass → fail and is tagged `safety`. Regressions 1 > 0 and safety fail. → **`hold`**, that case named first with its reason and trace link; next step = "`/opik-explain ` for the legal-advice item". + +**Needs review.** All gates pass, no policy file (defaults), so `judge_validated` is false. → **`needs_review`**: "20 items pass the defaults; a person should check 5 judge decisions (linked) and then set `judge_validated: true` in `opik-release-policy.yaml` — want me to write the file with the defaults?" + +**Insufficient evidence.** A two-item suite, both fixed, nothing regressed. → **`insufficient_evidence`**: "2 items is below `min_items: 10` — add cases with `/opik-test` or lower `min_items` in the policy (your call, and it will be visible in the file)." + +## Anti-patterns +A verdict without the criteria table; thresholds pulled from thin air rather than the file or the defaults; shipping on an unvalidated judge; treating a flaky item as a regression (or a regression as flaky) without the run data to say so; comparing runs on different suite versions; averaging away a safety regression; a p-value presented as a gate on six flips; rounding `hold` to `ship` because the aggregate went up; writing the policy file or recording the verdict without being asked; **editing application code**; deploying. + +## References + +Test-suite and experiment detail live in the `opik` skill, installed beside this one — paths relative to this file: `../opik/references/evaluation-test-suites.md` (execution policies, `runs_passed`/`runs_total`, versions, `get_test_suite_experiments`), `../opik/references/evaluation-datasets.md` (experiments, OQL). The numbers this skill judges come from `../opik-compare/SKILL.md`; judge validation is `../opik-evaluate/references/validate-evaluator.md`. If your host lays skills out differently, locate the `opik` skill's `references/` directory. + +If the `opik` skill isn't installed, say so in the report and use rather than working from memory. diff --git a/skills/opik/SKILL.md b/skills/opik/SKILL.md index 671a557..21d9a4b 100644 --- a/skills/opik/SKILL.md +++ b/skills/opik/SKILL.md @@ -1,9 +1,14 @@ --- name: opik description: Reference for the Opik SDK — tracing, span types, framework integrations, threads, and the prompt library (Python, TypeScript, REST). Use for "what span types exist", "how do I flush", "track_openai", "add OpikTracer", "version a prompt". To instrument a repo end to end, use the `opik-instrument` skill. +compatibility: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). A reference — needs no Opik connection to read; the snippets assume the `opik` Python or TypeScript SDK 2.x. The task-shaped skills (opik-instrument, opik-diagnose, opik-explain, opik-test, opik-compare, opik-evaluate, opik-online-eval, opik-optimize, opik-verify) read this skill's references and expect it installed beside them. +allowed-tools: + - Read + - Grep + - Glob metadata: - last_updated: "2026-09-08" - source_commit: "TODO — pin to the Opik release this was verified against (OPIK-7471)" + last_updated: "2026-09-17" + source_commit: "2.0.0" --- # Opik SDK Reference