Briefing edition:
Research · · reported
arXiv preprint proposes BOTTLED benchmark for LLM agents that build cheaper task-specific solutions
A preprint by Sonthalia, Puerto, Rubinstein, Gubri and Oh introduces BOTTLED, a benchmark where LLM agents must complete an entire unlabelled workload under fixed time, compute and API budgets by choosing their own approach, such as training a small model or writing a reusable program. The authors report that across ten models and three tasks, strong zero-shot performance did not reliably translate into strong bottling capability, while one reported run retained about 82% of zero-shot macro-F1 at roughly 657 times lower reported cost.
5.8/10 significance · AI confidence estimate 62%
What changed
Researchers Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein, Martin Gubri and Seong Joon Oh posted an arXiv preprint introducing BOTTLED, a benchmark in which LLM agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets.
Why it matters
If the reported pattern holds, it would suggest that agent cost savings on large repetitive workloads are not predictable from zero-shot benchmark scores alone, which matters for anyone budgeting agent deployments.
What remains uncertain
Still to verify for this briefing: when this specific development occurred; technical specifications; performance claims; independent corroboration.
What to watch
Watch for the full paper and any independent replication of the BOTTLED results, and for clarification of the benchmark's task set and cost accounting.
Sources
arxiv.org ↗
Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
Why this ranks here
The narrow action is a new arXiv preprint proposing a benchmark and reporting author-claimed results on agent "bottling"; it is relevant to builders weighing agent cost versus quality, but the full paper was not read, peer review and replication are not established, and the repository timestamps are not verified announcement times.
- impact
- 6/10
- reach
- 6/10
- novelty
- 7/10
- institutional
- 2/10
- evidence
- 6/10
- potential
- 6/10
Story development
First recorded development in this briefing.