Leaderboard
16 agent setups on the 56 TimelineBench 1.0 tasks, one attempt each. A task is resolved when the cut passes every test and clears the quality bar. Overall 125 of 896 slots resolved, 14.0% [11.7, 16.4].
| Rank | Model | Harness | Resolved | |||
|---|---|---|---|---|---|---|
| 1 | GPT-6 AstraOpenAI | Codexguidance curated | 26.8% [15.8, 40.3] | 15/56 | 52/56 | 24.4% |
| 2 | Claude Opus 5Anthropic | Claude Codeguidance off | 23.2% [13.0, 36.4] | 13/56 | 50/56 | 23.2% |
| 3 (tied) | GPT-6 AstraOpenAI | Codexguidance off | 21.4% [11.6, 34.4] | 12/56 | 48/56 | 23.8% |
| 3 (tied) | GPT-6 AstraOpenAI | OpenCodeguidance off | 21.4% [11.6, 34.4] | 12/56 | 48/56 | 22.3% |
| 5 (tied) | Claude Fable 5.1Anthropic | Claude Codeguidance off | 17.9% [8.9, 30.4] | 10/56 | 49/56 | 23.5% |
| 5 (tied) | Claude Fable 5.1Anthropic | OpenCodeguidance off | 17.9% [8.9, 30.4] | 10/56 | 50/56 | 22.9% |
| 5 (tied) | Claude Opus 5Anthropic | OpenCodeguidance off | 17.9% [8.9, 30.4] | 10/56 | 46/56 | 19.0% |
| 5 (tied) | GPT-5.6 SolOpenAI | OpenCodeguidance off | 17.9% [8.9, 30.4] | 10/56 | 45/56 | 17.6% |
| 9 | Grok 4.6xAI | OpenCodeguidance off | 12.5% [5.2, 24.1] | 7/56 | 39/56 | 17.0% |
| 10 (tied) | Gemini 3.8 FlashGoogle | OpenCodeguidance off | 10.7% [4.0, 21.9] | 6/56 | 42/56 | 13.1% |
| 10 (tied) | GLM 5.3 FlashZhipu AI (Z.ai) | OpenCodeguidance off | 10.7% [4.0, 21.9] | 6/56 | 41/56 | 9.5% |
| 12 | GPT-5.6 SolOpenAI | Codexguidance off | 8.9% [3.0, 19.6] | 5/56 | 45/56 | 13.1% |
| 13 | DeepSeek FlashDeepSeek | OpenCodeguidance off | 7.1% [2.0, 17.3] | 4/56 | 43/56 | 13.0% |
| 14 (tied) | GPT-6 AstraOpenAI | Codexguidance offcomputer use | 3.6% [0.4, 12.3] | 2/56 | 35/56 | 8.2% |
| 14 (tied) | Qwen 3.8 MaxAlibaba (Qwen) | OpenCodeguidance off | 3.6% [0.4, 12.3] | 2/56 | 35/56 | 4.0% |
| 16 | Gemini 3.1 Pro PreviewGoogle | OpenCodeguidance off | 1.8% [0.0, 9.6] | 1/56 | 19/56 | 9.0% |
Task resolution rate on 56 TimelineBench tasks, one attempt each; exact 95% CI. Ranked by resolution rate; ties share a rank.
Overall: 125 of 896 slots resolved, 14.0% [11.7, 16.4]. 687 slots pass all their tests. A missing output resolves nothing.
VerifiedRun and scored by the TimelineBench team. All 16 paper entries are verified.
Rank: by resolution rate; setups with the same number of resolved tasks share a rank, and ranks are recomputed over the setups shown. How scoring works.
Quality bars in the grid are the frozen collection tie margins (EditStock 0.21, Cinestudy 0.30, UGC 1.54, Commercial 0.52 SD of the panel score). The paper numbers score each setup with margins refitted without that setup's own votes, so a few slots near the bar differ.
Reading the leaderboard
Resolution rate is the share of the 56 tasks a setup resolves in a single attempt, with an exact (Clopper–Pearson) 95% interval. Missing outputs count as unresolved.
Tests passed counts tasks whose cut passes the four delivery tests, the six content guards and every validated brief must-have. Editors' W/T is the share of blind editor judgments that prefer the agent's cut or call it a tie.
Rank orders the setups by resolution rate; setups that resolve the same number of tasks share a rank. With intervals about ±11 points wide, 12 setups cannot be separated from the best, so read neighbouring ranks as close. Open a row for the per-task breakdown, or switch to the grid to see every slot.
Disclosures
- One run per task.
- The reference edits are professional edits, not certified ground truth.
Add an entry
The source packs are not distributed. To have an agent evaluated, fill in the access form. The maintainers will get in touch about running it on the benchmark internally, then score the run with the frozen verifier.