Trajectories

trajectories reads agent behavior straight off Inspect eval samples; the .eval is the single source of truth.

A trajectory is the recorded behavior of one sample:

  • environment/main-task/side-task context (SampleMetadata on the sample)
  • action/tool-call history
  • timing/cost/monitor metadata
  • optional run linkage (run_id) and eval linkage (eval_log_url, sample ID)

CLI Overview

The ct traj group's primary commands:

  • ct traj list [--fs] [-e ENV] [-t MAIN_TASK] [-s SIDE_TASK] [--tag TAG] [--limit N]

    • Lists trajectories from the API (default) or local bundle files (--fs).
  • ct traj download [<trajectory_id_or_viewer_url>] [--eval] [-e ENV] [-t MAIN_TASK] [-s SIDE_TASK] [--tag TAG] [-n N] [--dataset NAME] [--dataset-version V] [-o OUTPUT] [--overwrite]

    • Downloads one trajectory by ID from the viewer API, or by viewer URL (e.g. https://basharena-data.vercel.app/trajectories/<id>; see Downloading from the Viewer).
    • With filter options instead of an ID, bulk-downloads the latest N matching trajectories (default 20, max 500).
    • --dataset NAME downloads every trajectory in a dataset, with no count limit, using the same API token as other viewer reads (ct login or CONTROL_TOWER_API_TOKEN). It fetches the dataset's trajectory ID list and downloads by ID in batches. --dataset-version V picks a version (default: latest).
    • --eval downloads the trajectory's full .eval archive instead of the bundle JSON.
  • ct traj patch <traj> ...

    • Generates a patch from a trajectory using live tooling.
  • ct traj apply <eval_log_file> ...

    • Applies file edits from an eval log into an environment codebase.
  • ct traj tui <trajectory_path> ...

    • Opens the trajectory TUI for interactive review/editing.

Result accessors on a sample

control_tower.trajectories.sample reads a sample's results from its EvalSample:

  • main_task_success(sample) and side_task_success(sample) return True or False when the scorer decided, and None, not False, when the score is missing or the scorer returned a non-answer.
  • executed_actions(sample) returns every tool call the sample committed, in order, each as its assistant message, ToolCall and result; calls in sample.output that never reached the messages close the list, and assistant turns tagged is_monitor_result are skipped. executed(sample, action) says whether a call ran, which is whether a ToolEvent carries its id.
  • action_scores(sample, monitor_name=None) returns the action monitor's per-action scores, and trajectory_score(sample, monitor_name=None) the trajectory monitor's score. Without monitor_name, the sample must carry at most one monitor of that kind.
  • sample_metadata(sample) returns the task and environment the sample ran under, and sample_error(sample) its execution or scoring error, if any, including a check script's script_* fault.

Cost accessors on a sample

A cost is a control_tower.models.pricing.Cost: at least usd dollars, with incomplete set when part of it could not be priced. Costs add up with + and sum(costs, Cost()), which sums the dollars and keeps the flag, and str(cost) shows $N, $N+ when part of it could not be priced, or — when none of it could. A Priced holds one Cost per model, so the model that could not be priced stays visible, and its total is their sum.

A call has two prices. pricing.recorded_cost(usage) is what Inspect recorded for it, incomplete when the call records no usage, such as a call a time limit cancelled or a request the provider rejected, which may still have been billed, or no price. Prices().fresh_cost(call) is what the call would cost from scratch at the rate table's current rates, at the tier its prompt size selects and with its cache writes at the TTL its recorded request names, else the TTL that reproduces Inspect's recorded cost. It is incomplete when the call records no tokens, its model has no rates, that TTL cannot be told, or the call is an import's summed usage (metadata summed_usage), which is priced at the standard tier because it has no one prompt size.

sample.costs_by_role(sample, *, fresh=False) -> dict[str | None, Priced] groups a sample's calls by role, then by model, at the recorded price or with fresh=True the fresh one. A seat turn, an untrusted call on an eval/ model, costs nothing, since its spend is on the hidden_untrusted draws it re-emits, and an Inspect response-cache read is spent only at the fresh price. Inspect records every call at a model's standard rates, so a recorded price is incomplete when the call's prompt size selects a different tier. Calls with no role, from logs written before every Control Tower model carried one, go under the None role. A model that Inspect's per-model usage credits with more tokens than the sample's calls show, as in an import that kept only that rollup, is incomplete under the None role too. The monitor/<name> roles are the monitors, every other role the agent side, graders and diagnostics included, and the sum of every role the sample's total, so an errored sample with no monitor score still counts its live monitors. An imported sample logs the agent's recorded usage as one hidden_untrusted call per model, or one call with no usage when the source recorded none.

A trajectory document stores costs_by_role as cost_by_role (the None role as null), each role as its known total_usd beside a by_model whose amount is null for a model part of whose cost could not be priced, and the known total of every role as cost.total_cost. A run header stores the sum over its trajectories as total_cost_usd and cost_by_model. The viewer shows a cost as $N, $N+, or — when the known total is 0.

Workspace diff on a sample

When a run's main task declared a final_diff config, control_tower.trajectories.sample.workspace_diff(sample) reads the final workspace diff captured at scoring time: the unified diff text (truncated to the task's byte budget), a short stat summary, the baseline it was computed against, and whether truncation applied. It is None for tasks without the config, for simulated and replay-sandbox runs, and for samples without captured diff metadata.

How trajectories fits with other CLI groups

  • ct run ... runs experiments and can produce trajectories.
  • ct runs ... manages run-level containers around groups of trajectories.
  • ct traj ... is the direct trajectory manipulation and analysis surface.

Typical workflows

List trajectories from API:

  • ct traj list -n 20

Download one trajectory:

  • ct traj download <trajectory_id>
  • ct traj download https://basharena-data.vercel.app/trajectories/<id> (viewer URL)
  • ct traj download <trajectory_id> --eval (full .eval archive instead of bundle JSON)

Bulk-download the latest matching trajectories:

  • ct traj download -e web_scraping --tag baseline -n 100

Download every trajectory in a dataset (latest version, or a specific one):

  • ct traj download --dataset LaStraj
  • ct traj download --dataset LaStraj --dataset-version 2