Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
187 changes: 169 additions & 18 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,13 +4,48 @@
[![Codecov](https://img.shields.io/codecov/c/github/waitasecant/toolsnap?logo=codecov&label=Coverage&color=neongreen)](https://codecov.io/gh/waitasecant/toolsnap)
[![License: MIT](https://img.shields.io/badge/License-MIT-neon.svg)](LICENSE)

*Record your LLM agent's tool calls once. Replay them deterministically in every test run — no API keys, no network.*
*Zero-dependency, SDK-agnostic recorder and replayer for LLM agent. Record once, test the trajectory forever.*

## Why toolsnap?

LLM agents are non-deterministic at the model level, but their tool call trajectory i.e. which tools they call, in what order, with what arguments is the real testable surface. Existing approaches either hit live APIs on every test run (slow and expensive) or require writing brittle mocks upfront (tedious and disconnected from real behavior).
LLM agents are non-deterministic at the model level, but their tool call trajectory i.e. which tools they call, in what order, with what arguments is the real observable behavior.

`toolsnap` takes a different approach: **record first, assert after**. Run your agent once against the real APIs, capture every tool call to a JSONL fixture, then replay that fixture in CI forever **deterministically, offline, without credentials.**
Existing approaches' shortcomings:
- **Live APIs in every test run**: Slow and expensive. Worse: varying tool responses mean the agent takes different paths each run — trajectory assertions are unreliable.
- **Hand-written mocks**: You write the return value before seeing what the real agent produces. If the agent never calls the tool, the mock never fails. You're testing a fiction.
- **Network-level recording**: Records all HTTP — including every LLM request. Fixtures balloon in size and break on any SDK update, header change, or streaming format shift.

`toolsnap` takes a different approach: **record the real trajectory once, then replay and assert on it forever.**

---

## What `toolsnap` is/isn't?

### It is
- A trajectory recorder and assertion library for LLM agent tool calls
- A way to pin what the agent *does* so prompt or code changes that alter behavior surface immediately
- Most valuable when tools call external services you cannot or should not hit in CI

### It isn't
- A mock framework: you never invent return values by hand
- A network recorder: it operates at the Python function boundary, not HTTP
- A way to eliminate LLM calls: the LLM still runs in replay mode; only the tool backends are fixed
- An evaluation framework: it does not measure response quality or semantic correctness

---

## When to use `toolsnap`?

### Use when
- Tools that call external APIs, databases, clocks, or any service you don't want in CI
- Multi-step agents where tool call order matters
- Teams that want to catch when a prompt change silently alters agent behavior
- Any scenario where "does the agent still call the right tools?" is the key question

### Don't use when
- Tools that are already pure functions with no external calls — just unit test them directly
- Testing LLM output quality or reasoning — use evals for that
- HTTP-level fidelity (exact headers, status codes) — use vcrpy or pytest-recording

---

Expand All @@ -22,37 +57,153 @@ pip install toolsnap

## Quick Start

### Step 1: Decorate your tool with `@snap`
### Step 1: Record a real agent run

```python
# main.py — run once against live APIs
from toolsnap import snap
from strands import tool

@tool
@snap # auto-saves to fixtures/search.jsonl
# auto-saves to fixtures/search.jsonl
@snap
def search(query: str) -> list[str]:
"""Search the document store."""
return real_search_api(query)

# auto-saves to fixtures/get_weather.jsonl
@snap
def get_weather(city: str) -> dict:
return real_weather_api(city)

agent.run("what's the weather in london and find llm docs")
# fixtures/search.jsonl and fixtures/get_weather.jsonl written
```

### Step 2: Run your agent once to record real calls
```bash
python main.py
# fixtures/search.jsonl written
```

### Step 3: Replay in tests with `@replay`
### Step 2: Assert the trajectory in tests

```python
# test_agent.py — tool backends don't run; the LLM still runs
from toolsnap import replay
from strands import tool

@tool
@replay # reads from fixtures/search.jsonl — no API call
# reads from fixtures/search.jsonl
@replay
def search(query: str) -> list[str]: ...

# reads from fixtures/get_weather.jsonl
@replay
def get_weather(city: str) -> dict: ...

def test_research_agent_trajectory():
agent.run("what's the weather in london and find llm docs")
# search() and get_weather() returned their recorded responses
# assert on the agent's output or behaviour here
```

```bash
pytest test_agent.py # tool backends free, LLM still runs, trajectory deterministic
```

---

## Three ways to test

### 1. `@snap` / `@replay` — single-tool, decorator style

Best when you have one tool and want the simplest possible setup.

```python
# Record
@snap("fixtures/search.jsonl")
def search(query: str) -> list[str]:
"""Search the document store."""
...
return real_search_api(query)

# Test
@replay("fixtures/search.jsonl")
def search(query: str) -> list[str]: ...

def test_agent_uses_search():
result = agent.run("find llm docs") # search() returns recorded response
assert result is not None
```

### 2. `SnapSession` — multi-tool, with trajectory assertions

Best for agents that coordinate several tools. Wraps them all under one fixture and provides the full assertion API.

```python
from toolsnap import SnapSession, contains

def test_multi_tool_agent():
with SnapSession.replay("fixtures/session.jsonl") as s:
s.wrap(search)
s.wrap(summarize)
agent.run("find and summarize llm docs")

s.assert_called("search", times=1)
s.assert_called_with("search", query=contains("llm"))
s.assert_call_order(["search", "summarize"])
s.assert_no_errors()
```

### 3. `toolsnap_session` pytest fixture — CLI-controlled record/replay

Best for teams. Record and replay modes are controlled from the command line — no code changes needed.

```python
# conftest.py
pytest_plugins = ["toolsnap.pytest_plugin"]

# test_agent.py
@pytest.mark.toolsnap_fixture("fixtures/session.jsonl")
def test_agent_trajectory(toolsnap_session):
toolsnap_session.wrap(search)
toolsnap_session.wrap(summarize)
agent.run("find and summarize llm docs")
toolsnap_session.assert_called("search", times=1)
toolsnap_session.assert_call_order(["search", "summarize"])
```

```bash
pytest tests/ # replay — tool backends free, LLM still runs
pytest tests/ --toolsnap-record # re-record after prompt/code changes
pytest tests/ --toolsnap-strict=false # allow unexpected calls to fall through
```

---

## Assertion predicates

All assertion methods accept predicate objects for structural matching:

| Predicate | Matches when |
|---|---|
| `contains("llm")` | value contains the substring |
| `matches(r"\d{4}-\d{2}")` | value matches the regex |
| `any_of("london", "paris")` | value is one of the given options |
| `gt(0)` / `lt(100)` | value is greater / less than threshold |

```python
s.assert_called_with("search", query=contains("london"))
s.assert_called_with("embed", n_tokens=lt(512))
```

---

## CLI: inspect and diff fixtures

After a prompt change, `toolsnap diff` shows exactly what shifted in the agent's trajectory:

```bash
# No LLM API key. No database connection. The fixture drives everything.
pytest test.py # fast, offline, deterministic
toolsnap diff fixtures/before.jsonl fixtures/after.jsonl
# Diff: fixtures/before.jsonl → fixtures/after.jsonl
# ────────────────────────────────────────────────────
# search call 0 args unchanged result CHANGED (3 items → 2 items)
# + summarize call 0 ADDED

toolsnap show fixtures/session.jsonl # pretty-print all records
toolsnap stats fixtures/session.jsonl # call counts, avg/p95 latency, errors
toolsnap validate fixtures/session.jsonl
toolsnap list
```
41 changes: 41 additions & 0 deletions examples/google_adk/agent.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
"""
Google ADK agent definition — no toolsnap here.

toolsnap is applied in main.py (recording) and tests (replay).
This file contains only the tool and agent that would exist in production.

Note: ADK tools are plain Python functions — no SDK decorator is needed.
"""

import asyncio
from datetime import datetime, timezone

from google.adk import Agent
from google.adk.runners import InMemoryRunner


def get_current_time() -> str:
"""Return the current UTC date and time."""
return datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M:%S UTC")


async def _run(query: str, tools: list) -> str:
runner = InMemoryRunner(
agent=Agent(
name="Google ADK Example Agent",
model="gemini-2.0-flash",
tools=tools,
)
)
result = await runner.run_debug(query)
return str(result) if result is not None else ""


def make_agent(tools=None):
"""Return a callable(str) -> str that runs the ADK agent synchronously."""
_tools = tools if tools is not None else [get_current_time]

def _agent(query: str) -> str:
return asyncio.run(_run(query, _tools))

return _agent
45 changes: 45 additions & 0 deletions examples/google_adk/conftest.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
"""
Shared test configuration for the Google ADK example.

All tests in this directory require:

1. GEMINI_API_KEY — the LLM is live; toolsnap only fixes tool responses.
2. fixtures/session.jsonl — produced by running: python main.py

If either is missing every test is skipped with a clear explanation.
"""

import os
import sys
from pathlib import Path

import pytest

# Make agent.py importable when pytest is run from the project root.
sys.path.insert(0, str(Path(__file__).parent))

# Load the toolsnap pytest plugin so toolsnap_session is available in test_plugin.py.
pytest_plugins = ["toolsnap.pytest_plugin"]

# Single fixture file used by all three test routes.
FIXTURE = str(Path(__file__).parent / "fixtures" / "session.jsonl")


def pytest_collection_modifyitems(items: list) -> None:
if not os.getenv("GEMINI_API_KEY"):
_skip_all(
items,
"GEMINI_API_KEY not set — the LLM is live in these tests. "
"Run: export GEMINI_API_KEY=<key>",
)
elif not Path(FIXTURE).exists():
_skip_all(
items,
"trajectory fixture not recorded yet — run: python main.py",
)


def _skip_all(items: list, reason: str) -> None:
marker = pytest.mark.skip(reason=reason)
for item in items:
item.add_marker(marker)
1 change: 0 additions & 1 deletion examples/google_adk/fixtures/get_current_time.jsonl

This file was deleted.

52 changes: 23 additions & 29 deletions examples/google_adk/main.py
Original file line number Diff line number Diff line change
@@ -1,45 +1,39 @@
"""
Google ADK agent example — recording tool calls with @snap.
Record the agent's tool-call trajectory.

In ADK, tools are plain Python functions (no SDK decorator needed),
so apply @snap directly:
python main.py

@snap # toolsnap auto-saves to fixtures/{fn_name}.jsonl
def my_tool(...): ...
Wraps every tool with SnapSession before the agent runs so that each call —
function name, arguments, return value, duration — is saved to a single fixture.

Record once, then run tests freely:
python main.py # captures a real tool call
pytest test.py # replays it, zero API calls
"""

import asyncio
fixtures/session.jsonl ← written here, read by both test files

from google.adk import Agent
from google.adk.runners import InMemoryRunner
Re-run whenever you change the agent prompt, tools, or model.
The LLM call is live (requires GEMINI_API_KEY); only the tool backends are
captured and will be free to replay in tests.
"""

from toolsnap import snap
import os

from toolsnap import SnapSession

@snap # auto-saves to fixtures/get_current_time.jsonl
def get_current_time() -> str:
"""Return the current UTC date and time."""
from datetime import datetime, timezone
from agent import get_current_time, make_agent

return datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M:%S UTC")
FIXTURE = "fixtures/session.jsonl"


root_agent = Agent(
name="Google ADK Example Agent",
model="gemini-3-flash-preview",
tools=[get_current_time],
)
def main() -> None:
if not os.getenv("GEMINI_API_KEY"):
raise SystemExit("GEMINI_API_KEY is not set")

with SnapSession.snap(FIXTURE) as s:
wrapped = s.wrap(get_current_time)
agent = make_agent(tools=[wrapped])
response = agent("What is the current time?")

async def record():
"""Run the agent once against the real API and capture tool calls."""
runner = InMemoryRunner(agent=root_agent)
await runner.run_debug("What is the current time?")
print(response)
print(f"\nTrajectory saved to {FIXTURE}")


if __name__ == "__main__":
asyncio.run(record())
main()
Loading
Loading