Paper

Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut

Anonymous authors · Under review at ICLR 2027

Code and tasks: an anonymized mirror of the repository for review, at /code.

TL;DR

Timeline-Bench asks agents to turn raw production material and a brief into a finished video across 56 real editing tasks; the best of 16 agents resolves 15, and most unresolved runs pass every test but the quality test, which is calibrated on blind judgments by professional editors.

Abstract

AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from selecting dialog takes and shaping interview footage into a story to cutting commercials from product shots, voiceovers and graphics. Every task provides a brief, source assets, a container and a set of tests. A task is resolved when the output passes every test. The tests check the delivery format, the content and the brief’s explicit requirements, and include a quality test calibrated on 2,582 blind judgments by 43 video editors.

We evaluate 16 agents that pair frontier models with coding-agent harnesses such as Codex, Claude Code and OpenCode. The best, GPT-6 Astra in Codex with curated editorial guidance, resolves only 15 of the 56 tasks (26.8%), and the average agent resolves 14.0%. Human editors prefer the reference edit in 83.5% of judgments. Most unresolved runs (562 of 771) fail only the quality test: agents perceive footage through stills and transcripts and check their renders for defects, not craft. We release the tasks, verifier and per-run results at https://timelinebench.vercel.app.

  • video editing
  • AI agents
  • multimodal benchmarks
  • agent evaluation
  • human preference
  • editorial quality

Key numbers

Tasks
56
4 collections, 5,003 frozen input files (672.7 GB); source packs are not distributed; no footage published
Runs
896
16 agents × 56 tasks, one run each; 863 videos delivered
Resolved
14.0%
125 of 896 runs, exact 95% CI [11.7, 16.4]; best agent 15 of 56
Editors
43
2,589 blind judgments; the reference edit is preferred in 83.5% of the 2,582 assessable

Figures

Two-panel workflow diagram. A, task review and preparation: review the editorial assignment (brief, source footage, reference edit, editor review), check the agent input package against the versioned manifest with the reference held out, and exercise the task environment with a mechanical render and delivery checks. B, verifier checks: ground tests in stated requirements, probe known positives and negatives with complete and omission controls, and retain evidence and uncertainty as pass, fail or unresolved.
How tasks were prepared and how the verifier's checks were validated. Schematic; no footage.
Overview diagram of an editing assignment (brief, source assets, reference edit, tests) and three validation tracks: specification review, execution and integrity, and test validation.
The three validation tracks behind every assignment. Schematic; no footage.

Cite

During the double-blind review period the author entity is anonymous; the entry will be updated after review.

BibTeX
@inproceedings{anonymous2026timelinebench,
  title     = {{Timeline-Bench}: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut},
  author    = {Anonymous},
  booktitle = {Submitted to The Fifteenth International Conference on Learning Representations},
  year      = {2026},
  note      = {Under review}
}

To cite a benchmark release, see Citation.

Disclosures