Eval-Native Architecture
The .eval log is the only source of truth, and nothing derived is ever a source of truth: what is serialized besides it (the thin Mongo index, sabotage-eval reports) is regenerable from it. Metrics, monitor prompts, the viewer index, cost accounting, and the leaderboard are functions over EvalLog / EvalSample, memoized in memory at most.
Storage
| Store | Content |
|---|---|
| S3 | .eval archives — the data |
Mongo trajectories | thin index per sample (task ids, outcomes, sus scores, costs, eval pointers) for listing and filtering; regenerable, never a render source |
Mongo runs | run header: config, tags, cost totals, eval pointer |
| local run dir | sabotage-eval artifacts (results.json, graphs, summary.md) |
| live cluster dir | solution.json per solved task |
Reads
trajectories/sample.py: module-level functions over EvalSample, memoized on the in-memory sample — committed_actions, main_task_success / side_task_success (None when nothing decided, never a failure), action_verdicts, trajectory_verdict, sample_metadata, agent_cost, monitor_cost, total_cost_usd, sample_error, diagnostics.
A monitor verdict is an inspect_ai.scorer.Score — NOANSWER plus a reason when the monitor produced no number. monitor_verdict.py is the metadata vocabulary (cost, retries, ensemble members), not a type.
Departures from pure inspect
| Departure | Why |
|---|---|
committed_actions selection + ActionSubstep | blue protocols propose tool calls that never execute (trusted-editing rewrites, resampled proposals); inspect has no such notion |
action-monitor sample score: value is index-keyed ({"0": 3.0, "1": null}), full verdicts in metadata["verdicts"] | inspect scores are per-sample; action monitors score per-action |
ensemble aggregation not via ScoreReducer | an unscored member must not vote; inspect's value_to_float maps NOANSWER to 0.0 |
per-call cost pricing from ModelEvents | sample.model_usage is aggregate; tiered and cached pricing apply per call |
SampleMetadata over sample.metadata | typed environment / main-task / side-task identity |
| thin Mongo documents | cross-run listing and filtering without downloading evals |
solution.json | the live cluster's working artifact outlives individual evals |
Everything else — the log format, the transcript, the scores, the viewer's render source, replay's inputs — is plain inspect.