← All AI stories

Briefing edition:

Research · · reported

arXiv preprint proposes BOTTLED benchmark for LLM agents that build cheaper task-specific solutions

A preprint by Sonthalia, Puerto, Rubinstein, Gubri and Oh introduces BOTTLED, a benchmark where LLM agents must complete an entire unlabelled workload under fixed time, compute and API budgets by choosing their own approach, such as training a small model or writing a reusable program. The authors report that across ten models and three tasks, strong zero-shot performance did not reliably translate into strong bottling capability, while one reported run retained about 82% of zero-shot macro-F1 at roughly 657 times lower reported cost.

5.8/10 significance · AI confidence estimate 62%

What changed

Researchers Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein, Martin Gubri and Seong Joon Oh posted an arXiv preprint introducing BOTTLED, a benchmark in which LLM agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets.

Why it matters

If the reported pattern holds, it would suggest that agent cost savings on large repetitive workloads are not predictable from zero-shot benchmark scores alone, which matters for anyone budgeting agent deployments.

What remains uncertain

Still to verify for this briefing: when this specific development occurred; technical specifications; performance claims; independent corroboration.

What to watch

Watch for the full paper and any independent replication of the BOTTLED results, and for clarification of the benchmark's task set and cost accounting.

Sources

arxiv.org ↗
Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?

Discovery metadata from GDELT. AI summaries and significance scores can be wrong; read the original sources.

Why this ranks here

The narrow action is a new arXiv preprint proposing a benchmark and reporting author-claimed results on agent "bottling"; it is relevant to builders weighing agent cost versus quality, but the full paper was not read, peer review and replication are not established, and the repository timestamps are not verified announcement times.

impact
6/10
reach
6/10
novelty
7/10
institutional
2/10
evidence
6/10
potential
6/10

Story development

First recorded development in this briefing.

  1. · reported
    arXiv preprint proposes BOTTLED benchmark for LLM agents that build cheaper task-specific solutions