TimelineBench 1.0

56 real video-editing jobs, from raw footage to final cut. The best agent setup reaches the professional bar on about a quarter of them.

By TimelineBench Team

Today we are releasing TimelineBench 1.0: 56 video-editing jobs that turn raw production material and a client brief into a finished video, a verifier that scores the result against the professional cut of the same project, and the results of 16 agent setups on every task.

What it is

A good shot can make a poor cut: its value depends on the statement it supports, the sound beneath it and the shots around it. Video editing therefore tests an agent's ability to make interdependent decisions across a complete work product, which is what TimelineBench asks for.

Each task hands the agent a brief and a frozen source pack (camera originals, recorded voice-over and dialogue, music, graphics, scripts, shot lists) and asks for one file, /workspace/output/output.mp4, cut to the brief's delivery specification. The jobs come from four collections: EditStock stock productions (11), Cinestudy film scenes (15), 9:16 UGC ads (15) and commissioned Commercial work (15). The brief fixes the obligations and leaves the editorial solution open. Every task is a Harbor task.

The headline metric is the task resolution rate. A task is resolved when the cut passes all of its tests (4 delivery tests, 6 content guards against degenerate cuts, and the brief's must-haves) and clears a quality bar: the panel gap (a three-lab judge panel's score of the cut, minus the professional cut) must reach a per-collection tie margin calibrated on a 43-editor blind study (2,589 judgments; 2,582 assessable). The margins are positive, so the agent cut must out-score the professional cut.

Headline results

We ran 16 setups (ten models, three coding harnesses plus computer use, and one curated-guidance condition) on all 56 tasks, one attempt each: 896 slots and 863 delivered videos. Agents resolve 125 of 896 slots, 14.0% [11.7, 16.4].

Resolution rate, exact 95% CI
  1. GPT-6 Astra · Codex / curated26.8%
  2. Claude Opus 5 · Claude Code / off23.2%
  3. GPT-6 Astra · Codex / off21.4%
  4. GPT-6 Astra · OpenCode / off21.4%
  5. Claude Fable 5.1 · Claude Code / off17.9%
  6. Claude Fable 5.1 · OpenCode / off17.9%
  7. Claude Opus 5 · OpenCode / off17.9%
  8. GPT-5.6 Sol · OpenCode / off17.9%
  9. Grok 4.6 · OpenCode / off12.5%
  10. Gemini 3.8 Flash · OpenCode / off10.7%
  11. GLM 5.3 Flash · OpenCode / off10.7%
  12. GPT-5.6 Sol · Codex / off8.9%
  13. DeepSeek Flash · OpenCode / off7.1%
  14. GPT-6 Astra · Codex computer use / off3.6%
  15. Qwen 3.8 Max · OpenCode / off3.6%
  16. Gemini 3.1 Pro Preview · OpenCode / off1.8%
Share of the 56 tasks resolved in one attempt (bar) with its exact Clopper–Pearson 95% interval (whisker). With 56 binary tasks each interval spans about ±11 points: 12 setups cannot be separated from the best.
Task resolution rate on TimelineBench 1.0: 16 setups × 56 tasks, one attempt each, exact 95% CI. Ranked by resolution rate; ties share a rank. Each interval spans about ±11 points, so neighbouring ranks are often not statistically separable.
RankSetupTests passedResolvedResolution rate [95% CI]Editors W/T
1GPT-6 AstraCodex / curated52/5615/5626.8% [15.8, 40.3]24.4%
2Claude Opus 5Claude Code / off50/5613/5623.2% [13.0, 36.4]23.2%
3 (tied)GPT-6 AstraCodex / off48/5612/5621.4% [11.6, 34.4]23.8%
3 (tied)GPT-6 AstraOpenCode / off48/5612/5621.4% [11.6, 34.4]22.3%
5 (tied)Claude Fable 5.1Claude Code / off49/5610/5617.9% [8.9, 30.4]23.5%
5 (tied)Claude Fable 5.1OpenCode / off50/5610/5617.9% [8.9, 30.4]22.9%
5 (tied)Claude Opus 5OpenCode / off46/5610/5617.9% [8.9, 30.4]19.0%
5 (tied)GPT-5.6 SolOpenCode / off45/5610/5617.9% [8.9, 30.4]17.6%
9Grok 4.6OpenCode / off39/567/5612.5% [5.2, 24.1]17.0%
10 (tied)Gemini 3.8 FlashOpenCode / off42/566/5610.7% [4.0, 21.9]13.1%
10 (tied)GLM 5.3 FlashOpenCode / off41/566/5610.7% [4.0, 21.9]9.5%
12GPT-5.6 SolCodex / off45/565/568.9% [3.0, 19.6]13.1%
13DeepSeek FlashOpenCode / off43/564/567.1% [2.0, 17.3]13.0%
14 (tied)GPT-6 AstraCodex computer use / off35/562/563.6% [0.4, 12.3]8.2%
14 (tied)Qwen 3.8 MaxOpenCode / off35/562/563.6% [0.4, 12.3]4.0%
16Gemini 3.1 Pro PreviewOpenCode / off19/561/561.8% [0.0, 9.6]9.0%
Overall 125 of 896 slots resolved, 14.0% [11.7, 16.4]; 687 slots pass all their tests. Editors W/T = the human editors' win-or-tie rate against the professional cut. With 56 binary tasks, each setup's interval spans about ±11 points.

What we learned

The best setup resolves about a quarter of the tasks. GPT-6 Astra in Codex with curated guidance resolves 15 of 56 (26.8% [15.8, 40.3]), and Claude Opus 5 in Claude Code 13 of 56. Most cuts are not broken: 687 of 896 slots pass every test. But only 18.2% of those clear the bar, and editors prefer the professional cut 83.5% of the time. The gap that remains is craft.

Computer use costs 18.9 points of editor preference. Running GPT-6 Astra through Codex's native Computer Use in DaVinci Resolve, instead of in a coding environment, resolves 2 of 56 tasks. Of the six matched contrasts it is the only one that survives Holm correction: −18.9 points of agent preference [−25.8, −11.9], p < 0.001, and editors and the automatic verifier agree.

Harness and skills contrasts are null. Codex against OpenCode and Claude Code against OpenCode, for the same model, show no detectable effect; neither does offering the curated editorial-v1 guidance, although all 56 curated runs read it. The 95% intervals reach about 14 points in either direction, so moderate effects are not ruled out, and these contrasts were not preregistered.

Brief compliance is necessary but not sufficient. 762 of 896 slots pass every must-have, and across setups compliance tracks editor preference (Spearman 0.86). Within cuts the link is weak: editors give cuts that pass every must-have 17.0% win-or-tie, against 13.7% for cuts that fail one.

Agents do best on long narrative scenes and worst on UGC. Editors' win-or-tie is 28.7% on Cinestudy film scenes and 6.8% on UGC ads. Only 21 of the 56 tasks are resolved by any setup.

Can an automatic number stand in for editors?

Per setup, yes. Held out by setup (margins refitted without that setup's own votes), the resolution rate tracks the editors' win-or-tie rate with a mean error of 3.0 points and Spearman 0.93, and the editors' own per-setup rate has reliability 0.84. Per cut, no: the panel agrees with the editors' majority at κ 0.15, about what one editor achieves against the others. Report setup-level rates.

What we are releasing

  • 56 Harbor tasks with byte-identical briefs, delivery tests and oracles from the frozen release the paper evaluated, a dataset manifest and a registry entry.
  • The timelinebench package: input staging and verification, delivery tests, content guards, the brief-compliance engine with 56 rubrics, the quality bar with frozen constants, and aggregation with exact intervals.
  • Every per-slot result of the 16 setups, the sanitised judge outputs and the human-study aggregates, with scripts that recompute the paper's numbers.
  • This site: the leaderboard, every task with its brief, rubric and results, and the docs.

No video, still or thumbnail of the footage is published. The source packs are not distributed: fill in the access form and the maintainers will get in touch about running an agent on the benchmark internally (how). During review the code is an anonymized mirror at /code. To have your own agent evaluated, see Submit an agent.