Short answers about the product, scoring, verification, and execution controls.
Veridex runs and compares trading agents on TxLINE data. It is not a general model-evaluation tool.
Agents run strategies on live or replayed TxLINE markets. Separate code recalculates their scores and checks the saved evidence. The product reads TxLINE feeds, runs a defined strategy, and applies the safety checks needed for trading operations.
The agent chooses and proposes a trading action.
It can read TxLINE state, inspect allowed context, call approved tools, choose a strategy type, emit a constrained AgentAction, and explain its reason. Examples include value-vs-venue, stale-line, momentum, contrarian, spread, and market-making strategies.
The agent does not grade itself. It should not be the authority for CLV, executable edge, score rows, policy approval, proof checks, leaderboard rank, settlement, or payout.
Veridex agents should be configurable enough that the config can make or break the strategy.
The model is:
AgentTemplate + AgentConfig + PolicyEnvelope = AgentInstance
The template is the strategy family: value-vs-venue, stale-line, sharp momentum, arb scanner, market maker, or a future researcher/model-originator agent.
The config is the deployed strategy instance: market universe, signal thresholds, warmup/lookback windows, confirmation rules, quote freshness, liquidity/spread requirements, minimum executable edge, stake sizing, risk caps, cooldown, source mode, and execution mode.
So two users can deploy the same template with different configs and get different CLV, PnL, hit rate, and drawdown. That is the point of Agent Studio.
But configs cannot change Veridex's trust rules. They cannot bypass law/recompute, policy, evidence integrity, Checks, receipt separation, runtime/proof separation, or scoring immutability. The agent can trade differently; it cannot grade itself differently.
Because each layer catches a different failure mode:
| Layer | Job | Failure it prevents |
|---|---|---|
| Agent | Proposes a trading action | No intelligence or strategy |
| Deterministic recompute | Recomputes edge, CLV, and score from sealed inputs | Agent hallucinates or inflates numbers |
| Policy | Allows, denies, or pauses execution under risk limits | Good signal becomes unsafe trade |
| Proof | Produces public, tamper-evident receipts/checks | Dashboard numbers become trust-me screenshots |
Each part has one job:
Agent proposes action
Deterministic recompute verifies the math
Policy allows or denies execution
Proof card shows the trail
The deterministic/backtestable path is the part Veridex controls and can replay:
- TxLINE fixture normalization into
MarketState - recorded venue quotes and scored tool observations
- deterministic strategy code
- fair-value, executable-edge, CLV, and scoring math
- policy checks and capped sizing
- evidence hashes, proof checks, manifest roots, and leaderboard rank
LLM proposals are not treated as strictly deterministic, even with low temperature. Provider behavior, tool timing, hidden model updates, and context shape can drift. For LLM agents, Veridex records the action and evidence, then verifies the run by recomputing from sealed inputs.
reproducible means the same strategy code and replayed inputs regenerate the same actions and scores.
verified means the action was produced by an LLM or external runtime, then sealed, recomputed, policy-checked, and proof-checked. The run can be verified from evidence, but the LLM is not trusted to reproduce byte-identical behavior.
partial means the run was sealed and recomputed, but its proof is incomplete (a required check is pending/not_applicable, or evidence is missing). A partial run is still shown for transparency, but it is not eligible for ranking. Eligibility requires reproducible or verified (see derive.isEligible). partial is the third value of the shipped ProofMode enum (reproducible | verified | partial), and it is what drives the NOT-ELIGIBLE state on the leaderboard.
This distinction lets Veridex record LLM actions without claiming that an LLM will produce the same bytes on every run.
Only if the tool outputs are deterministic or recorded.
| Tool type | Example | Backtestable? | Proof treatment |
|---|---|---|---|
| Pure deterministic tool | Kelly calculator, edge calculator, line-move math | Yes | Can stay reproducible if the agent is deterministic |
| Recorded data tool | recorded TxLINE tick, recorded venue quote | Yes | Replay the recorded observation |
| Live external tool | live quote, news/context lookup, wallet exposure | Only if recorded | Usually verified |
| Execution tool | submit order, cancel order, transfer funds | Not scoring evidence | Must be policy-gated and non-scoring |
Rule: if a tool output affects a scored decision, that output must be sealed as decision evidence. Runtime telemetry such as latency, tokens, retries, or traces stays in the ops channel and never enters the evidence hash.
Not automatically. It depends on the agent and tool class.
Deterministic code agent plus deterministic or replay-recorded tools can remain reproducible.
An LLM agent that uses tools should usually be tagged verified, not reproducible. Veridex
checks the recorded actions instead of assuming that another model run will produce the same result.
Unrecorded tools are not acceptable for scored runs because they can introduce hidden, unreplayable inputs.
Separate the data source from the execution mode.
| Term | Meaning | Real live TxLINE? | Real venue/funds? | Purpose |
|---|---|---|---|---|
| Replay | Play recorded TxLINE ticks in order | No | No by default | Recreate a market window |
| Backtest | Replay plus scoring, checks, and leaderboard | No | No by default | Compare strategy performance before deployment |
| Paper | Agent acts, but no execution path submits anything | Maybe | No | Strategy evaluation only |
| Dry run | Full policy/execution lifecycle with simulated receipt | Maybe | No | Test production flow safely |
| Live | Consume current TxLINE feed | Yes | Depends on execution mode | Process current data |
| Live guarded | Real venue submit under policy, auth, caps, and kill switch | Yes | Yes | Production/live-money mode |
Implementation should compose two fields:
source_mode = "replay" | "live"
execution_mode = "paper" | "dry_run" | "live_guarded"Examples:
- Backtest:
source_mode=replay,execution_mode=paper - Execution backtest:
source_mode=replay,execution_mode=dry_run - Live paper:
source_mode=live,execution_mode=paper - Live dry run:
source_mode=live,execution_mode=dry_run - Live guarded:
source_mode=live,execution_mode=live_guarded
Yes, but they share the same engine.
Replay is the raw act of feeding recorded ticks back through the system. Backtest is replay plus evaluation: closing snapshots, CLV, simulated PnL, Brier, drawdown, proof checks, and leaderboard.
User-facing language should prefer "Backtest" because traders understand it. Internally, the system can still use source_mode="replay".
PnL is useful, but it is noisy and can over-reward luck, stake size, and outcome realization. CLV is the primary rank metric because it asks whether the agent beat the later market price, which is a cleaner signal of trading edge.
Veridex can show PnL, hit rate, Brier, and drawdown as performance metrics. They should not replace CLV as the primary rank for the hackathon product.
Executable edge decides whether an agent should act now.
For venue execution, edge should compare TxLINE de-margined fair value against the actual executable venue price:
mispricing_gap_bps = txline_fair_probability_bps - venue_implied_probability_bps
executable_edge = txline_fair_probability * venue_decimal_odds - 1
The first line is the probability-space dislocation. It is useful for explanation, but it is not executable edge. Executable edge is the EV form used for action/risk decisions.
CLV is different. CLV measures whether the entry beat the later closing line:
clv = closing TxLINE probability - entry TxLINE probability
Edge gates action. CLV ranks performance.
Veridex started with fair-value dislocation. This strategy compares two market prices instead of trying to predict a match result. It is one strategy in the Arena.
The agent compares TxLINE de-margined consensus fair value with an executable venue price:
executable_edge = txline_fair_probability * venue_decimal_odds - 1
If the edge clears threshold and policy allows it, the agent can propose an action. Later, Veridex proves whether the action beat the close with CLV.
This test makes a narrower claim than a match-prediction model. It asks whether two current prices disagree enough to cover trading costs.
No. TxLINE is the fair-value/reference feed.
To claim executable edge, Veridex also needs an executable venue price, or a replay/paper venue quote. Without that second price, the agent can still produce a signal and CLV record, but it cannot claim live execution edge.
No.
Proof Checks and performance metrics must stay separate:
- Checks answer: can this run be trusted?
- Metrics answer: how well did the agent perform?
Evidence integrity, manifest binding, receipt separation, and anchor status are eligibility/trust guarantees. They should not add performance points. CLV remains the primary rank metric, with PnL, Brier, drawdown, hit rate, and sample size as supporting metrics.
No. Dry-run receipts prove the execution path would have fired under policy. They do not prove skill.
Skill comes from the saved decision, recalculated math, CLV, and proof checks. Execution receipts are non-scoring records.
No. The odds behind every sealed result are checkable against TxLINE's Merkle-anchored root. Veridex records a proof-status stamp for them, and in a live check 269/270 of our sampled World Cup odds returned valid TxLINE inclusion proofs. So the law recompute proves our math is faithful to the sealed inputs, and the Merkle check proves those inputs are authentic TxLINE data we didn't edit.
The demo should show:
- Agents ingest TxLINE data.
- Agents autonomously propose strategy actions.
- Veridex recomputes edge/CLV instead of trusting the agent.
- Policy allows or denies execution.
- Proof card verifies the run.
- Backtest/replay and live modes are clearly labeled.
- Dry-run/live-guarded execution never gets confused with scoring.
The sharp one-liner:
Trading agents cannot verify their own results.