EC2 Fleet (Distributed Evaluation)

--ec2 runs an evaluation on a fleet of EC2 workers instead of local Docker. Jobs are distributed over SQS; each worker runs Docker locally through the standard lb-docker sandbox.

ct run eval --ec2, or sandbox: {use_ec2: true} in a --run-config YAML, runs one fleet job per sample and epoch. Each job runs the same control_eval task with the same task arguments as the local run, narrowed to its sample with Inspect's sample_id filter, and the controller merges the job logs into one .eval under --log-dir. ct run sabotage-eval --ec2 runs its evals the same way. ct run mtgen's live run needs --no-ec2, because workers never see the tasks it stages locally. Workers boot in about 65 seconds from the pre-baked worker Amazon Machine Image (AMI), or in about 4 minutes from stock Ubuntu 22.04 when no AMI is available.

Authentication

Use AWS single sign-on (SSO) where you can, since it has no long-lived secrets; static Identity and Access Management (IAM) keys also work.

AWS SSO. You need an Identity Center assignment to the fleet account (326120904690) with the LinuxbenchFleet permission set; ask a fleet account admin in #devops. This permission set is enough to run evals. If you already have a profile for that account and permission set, reuse it. Otherwise run aws configure sso, using the organization's start URL from #devops and a profile name of your choice.

List your profiles with aws configure list-profiles, then replace your-profile below with the one you configured:

export AWS_PROFILE="your-profile" aws sso login aws sts get-caller-identity

If your default AWS profile already selects the fleet account and permission set, omit AWS_PROFILE. Re-run aws sso login when the selected profile's session expires. In-flight runs refresh credentials only while the SSO session is valid, so log in fresh before a long run.

Control Tower also reads AWS_PROFILE from .env if the shell does not set it; this only selects an existing local profile and does not configure SSO. The AWS CLI and the ct-box scripts do not read .env, so exporting the variable in your shell is the way to select the same profile for all three.

Remove any AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY lines from .env when using AWS_PROFILE. Environment-variable keys take precedence over the profile, so fleet commands refuse to start when both are set.

Static IAM keys. Request a scoped fleet IAM key pair from a fleet account admin, who follows EC2 Fleet Administration, then put it in .env:

AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=...

Fleet Mode

Prerequisites

One-time setup:

  1. API keys: set ANTHROPIC_API_KEY and/or OPENAI_API_KEY in .env, plus GH_TOKEN if a dependency is a private GitHub repo. The controller passes them to workers as per-instance AWS Systems Manager (SSM) SecureStrings that only that instance can read, and deletes them at teardown.
  2. Docker registry auth for private env images: either
    • set DOCKER_ORG_NAME and DOCKER_ORG_PAT in .env, which reach workers like the API keys, or
    • store them in SSM under /linuxbench/docker-org-name and /linuxbench/docker-org-pat (overridable with DOCKER_ORG_NAME_SSM_PARAM / DOCKER_ORG_PAT_SSM_PARAM, staying under /linuxbench/docker-org-). The helper script stores the PAT as SecureString.

To store the Docker auth from your current .env into SSM:

./scripts/put_docker_org_auth_ssm.sh

Usage

# Fleet mode on EC2 ct run eval --ec2 --task-file medium --untrusted-policy honest --epochs 10 # Scale up with more workers for faster runs ct run eval --ec2 --task-file medium --untrusted-policy honest --epochs 10 --ec2-arg max_workers=20 # Cost estimate only (no launch) ct run eval --ec2 --task-file medium --untrusted-policy honest --epochs 10 --ec2-arg estimate=true # Skip confirmation prompt (CI-friendly) ct run eval --ec2 --task-file medium --untrusted-policy honest --epochs 10 --ec2-arg auto_confirm=true # Force a fresh fleet (skip auto-resume) ct run eval --ec2 --task-file medium --untrusted-policy honest --epochs 10 --ec2-arg new_fleet=true

Every fleet knob is an --ec2-arg key=value (repeatable); the keys are the fields of Ec2Config in src/control_tower/sandbox/cli.py, listed with their defaults in --help. A sabotage-eval config sets the same keys in an ec2_args: mapping at run level or per attack; the attack's keys win.

Workers run the repo code at a pinned git SHA: HEAD when the worktree is clean, or a tree SHA of the dirty worktree, so uncommitted changes reach workers. Untracked files ship only after git add; the controller warns at launch if any are present.

Workers have only that code, the shipped settings and --extra-src, so the run refuses a task file that is not tracked in the repo or setting it lives in, and a --trajectory-id or replay traj_like that is a local file; upload the trajectory and pass its ID or URL instead.

Settings on workers

Workers get the settings their environments come from, plus the setting a --task-file lives in, as they are on your machine, under the same rule as the repo code: the setting's repo and the repo holding each environment the run uses ship at their HEAD, or at a tree SHA of their dirty worktree, and the controller warns about untracked files it leaves out. Workers never run ct settings pull, so unpushed commits, environments off their locked commit, local-path settings and private repos all ship as they are locally. To run a branch of a setting or an environment, check it out locally. An environment whose directory resolves outside its setting, or holds no files tracked by git, stops the launch. Local changes to an environment's Dockerfile or image contents do not travel, since workers run its prebuilt image: pin ENV_IMAGE_TAG_<ENV> to an image built from your checkout. A change to a shipped repo starts a new fleet rather than resuming one. The fleet manifest and the merged log's setting_repos metadata record each shipped repo's worker directory and the SHA it shipped at, with +dirty appended when that differs from its HEAD.

Shipping external code to workers (--extra-src)

Workers get only a git archive of the control-tower checkout, so an external --untrusted-policy mypkg.foo:bar raises ModuleNotFoundError on the worker. --extra-src DIR also ships DIR as an import root, like PYTHONPATH=DIR:

ct run eval --ec2 --extra-src ~/experiments --untrusted-policy mypkg.policies:my_policy -e load_balancer -s ...

The flag is repeatable, and each directory needs a distinct basename. It applies only to --ec2, since local runs resolve external refs from the working directory. Build artifacts, caches, and VCS metadata (__pycache__, .git, node_modules, .venv/venv, build, dist) are not shipped, so source that lives in a directory named build or dist will be missing on the worker.

An external environment passed as -e my_project.envs:my_env ships the same way: pass --extra-src with the directory that holds my_project, and no setting ships for it. Without it the run stops before launch. That whole directory ships, siblings included, so keep the package in a small directory of its own.

External sandboxes on the fleet (--sandbox with --ec2)

--sandbox module:type runs on the fleet too: each worker imports the module from the shipped extra-src trees and runs the sandbox for its own jobs. If the sandbox needs host-side installation (a microVM runtime, a CLI login), name a script inside an extra-src tree with --ec2-arg worker_setup=<basename>/<path>. The worker runs it as root after extracting extra-src and before taking jobs, with the venv at /opt/linuxbench/.venv and the worker's IAM role, so it can read its own SSM parameters. Put environment variables the sandbox needs at run time in /opt/linuxbench-worker-setup.env, which the worker sources after the script. Pick the hardware with --ec2-arg instance_type=...; a microVM sandbox needs /dev/kvm, which only *.metal types expose. The setup script and instance type are part of the fleet ID, so a run that differs only in them does not resume another's results.

ct run eval --ec2 --ec2-arg instance_type=c5.metal \ --ec2-arg worker_setup=agent-glovebox/evals/control_tower/ct-worker-setup.sh \ --extra-src ~/agent-glovebox --sandbox evals.control_tower.sbx_gb:sbx-gb -e file_indexing ...

Isolated settings

Fleet jobs automatically use each isolated setting's dependencies and the shipped Control Tower checkout, so no custom worker_setup script is needed. A fleet run splits into isolated settings and in-process environments exactly as a local run does, and each job runs in its part's venv.

Pricing on workers

The controller uploads the driver's rate table beside the code artifact and each worker prices from it, so a fleet run costs what the same run would cost on the driver; CONTROL_TOWER_PRICING=off on the driver turns pricing off on the workers too.

Scaling

The main knob for throughput is --ec2-arg max_workers=N. Each worker starts at 10 concurrent evals and adapts to rate-limit pressure up to concurrency_per_worker. A fleet first launches up to 24 workers, fewer when max_workers or the pending work is smaller, and adds workers up to max_workers while it is CPU-bound.

WorkersConcurrent evalsEC2 cost/hrGood for
550~$3.50Small runs (<100 samples)
12120~$8.50Medium runs (100-300 samples)
24240~$17Large runs (300-600 samples)
40400~$28Very large runs (600+ samples)

Resume

Fleet runs resume automatically. The fleet ID is a hash of the git SHA, policy, tasks, and epochs, so after a controller crash, credential expiry, or network failure, re-run the same command. The controller finds the existing manifest in S3, lists the completed .eval files, enqueues only the remaining work, and launches fresh workers for it. When some jobs produce no sample, the run fails with the count of failed jobs, and rerunning the same command retries only those. Rerunning with --no-fail-on-error (or fail_on_error: false in a run config) also resumes the fleet: it reuses the finished jobs, keeps a sample that errors in the log instead of failing its job, and writes the partial log. A rerun without --no-fail-on-error refuses errored samples that an earlier --no-fail-on-error run kept; --ec2-arg new_fleet=true runs every job again from scratch.

After a three-minute grace period and once at least five jobs have been seen, the controller aborts the fleet and terminates its workers if every job so far has failed, or if more than 20% of jobs have failed persistently (failed on their third delivery, at least five of them); the queues are deleted and only the S3 results are kept for a resume. Each worker terminates itself after five minutes with no work, so workers shut down on their own once the queue drains.

Monitoring Fleet Resources

# Show your running instances and queues ct cloud status # Show all control-tower resources (all users) ct cloud status --all-users

Fleet Streaming Logs

With --ec2, the controller prints only a one-line progress counter ([N/M] X done, …). To follow per-trajectory results live, tail logs/fleet/<fleet_id>/events.jsonl; the startup banner prints the exact path. Each line is one JSON object; types include:

  • fleet_start — fleet_id, expected jobs, worker count
  • worker_online — instance_id, bootstrap_s
  • result — job_id, task, model, per-scorer scores, error, duration_s, instance_id. Emitted as each sample finishes, so concurrent samples on one worker arrive as separate events.
  • job_done — job_id, status, instance_id, duration_s, failure_category, error_summary. Emitted at batch end from the SQS done queue.
  • fleet_summary — final counts

Recipes:

# Watch failures land: jq -c 'select(.type=="job_done" and .status!="completed")' \ logs/fleet/<fleet_id>/events.jsonl # Live per-task scores as each trajectory finishes: jq -c 'select(.type=="result") | {task, status, scores, duration_s}' \ logs/fleet/<fleet_id>/events.jsonl # Running pass-rate across completed trajectories: jq -s '[.[] | select(.type=="result")] | {n: length, passed: ([.[] | select(.scores.linux_scorer.value.main_task_success=="C")] | length)}' \ logs/fleet/<fleet_id>/events.jsonl # Pull a failed .eval for inspection (workers also upload full .eval at batch-end): aws s3 cp s3://<bucket>/<fleet_id>/<job_id>.eval /tmp/job.eval uv run inspect view /tmp/job.eval

Debugging

  • Bootstrap logs: uploaded to s3://lb-fleet-results-migrated/{fleet_id}/logs/{instance_id}-bootstrap.log when bootstrap exits, with timing at each step
  • EC2 console: link printed after launch
  • SSH: via EC2 Instance Connect (emergency debugging)
  • Stale warning: the controller warns after 10 minutes with no progress
  • Hard timeout: the controller terminates the fleet after 3 hours with no progress while workers are silent (or 1.5 × the sample time limit, if longer), and after 6 hours with no completed sample even while workers are active