Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"name": "opik",
"version": "0.4.0",
"description": "Opik tracing for Claude Code sessions, plus the Opik agent skills: instrument, diagnose, explain, test, compare, evaluate, online-eval, optimize",
"version": "0.5.0",
"description": "Opik tracing for Claude Code sessions, plus the Opik agent skills: instrument, diagnose, explain, test, compare, evaluate, verify, online-eval, optimize",
"author": {
"name": "Comet ML",
"url": "https://comet.com"
Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -190,6 +190,7 @@ The plugin ships the Opik skill pack — the same skills published as [`opik-ski
| `opik-evaluate` | "evaluate my agent", "build an eval", "write an LLM judge" |
| `opik-online-eval` | "score production traces", "take this judge live" |
| `opik-optimize` | "optimize this prompt", "run the prompt optimizer" |
| `opik-verify` | "is this safe to ship", "go/no-go on this change" |

They are vendored at a pinned `opik-mcp` commit (see `skills/SHARED.md`); edit them upstream, not here.

Expand Down
4 changes: 2 additions & 2 deletions scripts/sync-shared-skills.sh
Original file line number Diff line number Diff line change
Expand Up @@ -12,11 +12,11 @@
set -euo pipefail

CANON_REPO="${CANON_REPO:-https://github.com/comet-ml/opik-mcp.git}"
CANON_REF="${CANON_REF:-0baa5aecf0224701473efd968be4aa5a166ba99f}" # opik-mcp main after #191 (nine skills)
CANON_REF="${CANON_REF:-8fb9c4b09f8854581e7b34a37049773f5fedd869}" # opik-mcp main after #200 (ten skills)
SRC="src/opik_mcp/skills"
DEST="skills"
# Every skill the pack ships. Order matches the published index.
SHARED=(opik opik-compare opik-diagnose opik-evaluate opik-explain opik-instrument opik-online-eval opik-optimize opik-test)
SHARED=(opik opik-compare opik-diagnose opik-evaluate opik-explain opik-instrument opik-online-eval opik-optimize opik-test opik-verify)

tmp="$(mktemp -d)"
trap 'rm -rf "$tmp"' EXIT
Expand Down
7 changes: 4 additions & 3 deletions skills/opik-compare/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: opik-compare
description: Run a candidate against the baseline over an Opik test suite and read the numbers back — which cases broke, which got fixed, the per-metric deltas, worst rows, and whether the two runs are comparable — with the Opik compare-view link. Runs via the SDK; reads results via the MCP when connected. Does not issue a ship/no-ship verdict. Use for "did my fix work", "compare against the baseline", "run the regression suite", "why did quality drop", "which cases regressed", "compare these two experiments". Not for live production triage (use diagnose), building an evaluation from scratch (use evaluate), or capturing a single case (use test).
description: Run a candidate against the baseline over an Opik test suite and read the numbers back — which cases broke, which got fixed, the per-metric deltas, worst rows, and whether the two runs are comparable — with the Opik compare-view link. Runs via the SDK; reads results via the MCP when connected. Does not issue a ship/no-ship verdict. Use for "did my fix work", "compare against the baseline", "run the regression suite", "why did quality drop", "which cases regressed", "compare these two experiments". Not for the ship/hold decision itself (use verify), live production triage (use diagnose), building an evaluation from scratch (use evaluate), or capturing a single case (use test).
compatibility: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python or TypeScript project with Opik configured and a test suite (from the test or evaluate skill) or two existing experiments. Install the `opik` skill alongside this one — it holds the shared test-suite and experiment references; without it, this skill falls back to the public docs.
allowed-tools:
- Read
Expand Down Expand Up @@ -69,6 +69,7 @@ result = opik.run_tests(
experiment_name="candidate-<sha>",
experiment_tags=["compare", "<sha>"],
model="<same judge model as baseline>",
generate_report=False, # default True writes opik_test_suite_reports/ into cwd — the repo stays untouched
)
candidate_id = result.experiment_id # result.experiment_url is the single-run link

Expand Down Expand Up @@ -109,7 +110,7 @@ rows = client.get_experiments_client().find_experiment_items_for_dataset(
8. **Flaky cases** — items that flip across repeated runs of the *same* code (the suite's execution policy exposes `runs_passed`/`runs_total`).
9. **Trend** — when more than two runs exist, the pass rate across the last few, so a one-step delta has context.

Report what the numbers say. **Do not decide ship or hold** — that decision has its own policy and is a later step.
Report what the numbers say. **Do not decide ship or hold** — that is `/opik-verify`, which applies an explicit release policy to these numbers.

### 6. Link and hand off
Build the compare link with **both** ids — a URL-encoded JSON array — so the user lands on the side-by-side view:
Expand All @@ -118,7 +119,7 @@ import json, urllib.parse
ids = urllib.parse.quote(json.dumps(["<baseline_id>", candidate_id]))
compare_url = f"{ui_base}/{workspace}/experiments/{suite.id}/compare?experiments={ids}"
```
`ui_base` is the Opik UI origin (the configured URL minus `/api`; `result.experiment_url` shows the exact host and workspace to reuse). Then one next step (see **Output**).
`ui_base` is the Opik UI origin (the configured URL minus `/api`; `result.experiment_url` shows the exact host and workspace to reuse). Then one next step (see **Output**) — when the user's question is whether to ship, that step is `/opik-verify`.

### 7. Record (only when asked)
If the user wants the finding kept beside the data: a comment on a regressed case's trace via `client.rest_client.traces.add_trace_comment(trace_id, text=…)`, or a human score beside the judge's via `client.log_traces_feedback_scores([...])`. Never by default.
Expand Down
5 changes: 3 additions & 2 deletions skills/opik-evaluate/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: opik-evaluate
description: Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link. Picks a test suite with judge assertions or a dataset with metrics, sources cases from traces or synthetic data, scores heuristics-first then one-failure-mode judges, runs client-side via the SDK or server-side for prompt-only targets, and reads the scores back. Covers RAG evaluation, error analysis, and validating a judge against human labels. Use for "evaluate my agent", "measure quality", "build an eval", "write an LLM judge", "how good is my RAG", "set up evals for this". Not for before/after on an existing suite (use compare), one regression case (use test), or scoring production traffic (use online-eval).
description: Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link. Picks a test suite with judge assertions or a dataset with metrics, sources cases from traces or synthetic data, scores heuristics-first then one-failure-mode judges, runs client-side via the SDK or server-side for prompt-only targets, and reads the scores back. Covers RAG evaluation, error analysis, writing and validating LLM judges against human labels, and auditing an existing eval pipeline. Use for "evaluate my agent", "measure quality", "build an eval", "write an LLM judge for hallucinations", "audit our evaluation pipeline", "how good is my RAG", "set up evals for this". Not for before/after on an existing suite (use compare), one regression case (use test), scoring production traffic (use online-eval), or the ship/hold decision (use verify).
compatibility: Tested with Claude Code; works with any Agent Skills-compatible host (Cursor, VS Code Copilot, Codex). Requires a Python or TypeScript project with Opik configured. Install the `opik` skill alongside this one — it holds the shared test-suite, dataset, and metric references; without it, this skill falls back to the public docs.
allowed-tools:
- Read
Expand Down Expand Up @@ -74,7 +74,8 @@ Write the task adapter as a temp file **outside the repo** (needs the app's prov
```python
# Test suite
results = opik.run_tests(test_suite=suite, task=lambda item: {"input": item["input"], "output": str(app(item["input"]))},
experiment_name="baseline-<sha>", model="<judge model>")
experiment_name="baseline-<sha>", model="<judge model>",
generate_report=False) # default True writes opik_test_suite_reports/ into cwd — keep the repo clean
# Dataset
from opik.evaluation import evaluate
res = evaluate(dataset=dataset, task=task, scoring_metrics=[...], experiment_name="baseline-<sha>",
Expand Down
2 changes: 1 addition & 1 deletion skills/opik-optimize/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@ dataset = client.get_dataset(name="<dataset>", project_name="<project>")

### 3. Define the metric
A function `(dataset_item, llm_output) -> float`, higher is better. **Give it a real name** (`def refund_answer_similarity(...)`) — its `__name__` becomes the Optimization run's objective name in the UI and `result.metric_name`; a function called `metric` shows up as "metric".
- `expected_output` present → heuristic (`LevenshteinRatio`, `Equals`, or a task-specific check) — deterministic and free.
- `expected_output` present → heuristic (`LevenshteinRatio`, `Equals`, or a task-specific check) — deterministic and free. Prefer a **graded** metric over exact match: when the baseline scores 0.0 on every item (observed with `Equals` on a strict output format), every candidate also scores 0.0 and the optimizer has nothing to climb — five trials of flat zeros is a metric problem, not a prompt problem.
- Otherwise → **one** binary judge for the failure mode being optimized (`../opik-evaluate/references/write-judge-prompt.md`), wrapped to return its score `.value`. Multi-objective → `MultiMetricObjective`.
Never optimize against a judge nobody validated: an unvalidated judge is the easiest thing to overfit.

Expand Down
Loading
Loading