Run

How to run TimelineBench 1.0.0

01Install

TimelineBench runs in Harbor 0.23. Use a Linux VM with Docker and a large disk: task inputs range from 0.29 GB to 74.6 GB.

uv tool install harbor==0.23.0

Then get the repository and install the scoring package (CLI timelinebench, alias tlb). During review the code is the anonymized mirror at /code.

curl -LO https://timelinebench.vercel.app/code/timelinebench-1.0.0-37e21e6.zip
unzip -q timelinebench-1.0.0-37e21e6.zip && cd timelinebench-1.0.0-37e21e6
uv venv && uv pip install -e '.[all]'
source .venv/bin/activate

02Have an agent evaluated

The source packs are not distributed. To have an agent run on the benchmark, fill in the access form with the agent you want evaluated and the collections the run should cover. The maintainers will get in touch about how they will run it internally. You do not receive a download of the benchmark files.

03Smoke-test with the oracle

The oracle is a mechanical render that meets the delivery contract; it should score reward 1. If it does not, the problem is the environment or the inputs, not an agent. Trials stage from the verified mirror, mounted read-only.

unset TLB_INPUTS_URL TLB_INPUTS_TOKEN   # the token stays on the host
export TLB_INPUTS_SOURCE=dir            # stage from the mounted mirror
MOUNT="[{\"type\": \"bind\", \"source\": \"$TLB_MIRROR\", \"target\": \"/mnt/timelinebench-inputs\", \"read_only\": true}]"
harbor run -p tasks/whisper-spot -a oracle --job-name oracle-whisper-spot -y --mounts "$MOUNT"

04Run an agent

One attempt per task and the default 300-minute time limit, as on the leaderboard. The paper ran each model at its provider's maximum reasoning setting.

harbor run -p tasks -a claude-code -m anthropic/claude-opus-5 \
  --ak reasoning_effort=max --job-name my-run -n 4 \
  -y --mounts "$MOUNT"

harbor view jobs opens Harbor's viewer. The paper's native-CLI and computer-use setups are described in Running agents.

05Score

Harbor's reward is the four delivery tests only. A task is resolved when the cut also passes the six content guards and the brief's must-haves, and clears the quality bar. Scoring the last two calls three judges (Gemini 3.8 Flash, GPT-6 Astra, Claude Opus 5.5).

timelinebench score --task whisper-spot \
  --video jobs/my-run/<trial>/steps/solve/artifacts/output.mp4 \
  --sources "$TLB_MIRROR/whisper-spot" --out scores/whisper-spot.json
timelinebench aggregate scores/ --label "my-agent / my-model"

The delivery tests run in Docker, in the image timelinebench-verifier:1.0.0, which score builds from the task's environment/ on first use (--build rebuilds it; --mode local uses a local ffprobe and pytest instead). aggregate reports the resolution rate over the 56 tasks with its exact 95% interval; a task without a report counts as unresolved. See Scoring.

06Submit

The source packs are not distributed: fill in the access form and the maintainers will get in touch about running your agent internally.

Keep media private

The source packs and the output videos stay with the maintainers. They are never put in a pull request, an issue or a public bucket. Access terms.