Back to articles
Evaluation & Benchmarks

ArrivalBench: Agent-Built Data Pipelines Can Fail Over Time

3 min read

Introduction

Getting an agent to produce a data pipeline that runs is no longer the whole evaluation problem. The harder question is whether that pipeline remains correct when records arrive late, appear more than once, arrive out of order, or are delivered again after a retry. The arXiv paper ArrivalBench argues that conventional snapshot-based grading is too optimistic for this setting.

Key findings

  • From snapshots to event replay. A conventional benchmark runs a pipeline once on a fixed batch and checks the output. ArrivalBench keeps the artifact produced by the agent and executes it again under adversarial but replayable delivery schedules.
  • Correctness is defined by recomputation. The final state produced incrementally is compared with a batch recomputation over the complete log. This makes a crashed pipeline distinguishable from one that finishes but produces a wrong table.
  • One-shot certification hides substantial failures. Across 40 tasks, a reimplementation of single-execution grading certified 86% to 100% of pipelines produced by each of 11 models. Replaying those same certified artifacts found silent errors in 7.0% to 79.2% of them.
  • Snapshot repair does not solve the underlying issue. Pipelines repaired against the snapshot test failed replay at roughly similar rates to pipelines that passed the snapshot test without repair. The missing factor is the temporal environment, not simply the number of repair rounds.
  • Idempotency is the recurring weakness. Every model showed more failures from hazards involving repeated processing than from ordering hazards alone.

Why it matters

The paper broadens what “correct” should mean for agent-generated data work. A pipeline that returns the right answer for one input snapshot may still corrupt state as events continue to arrive. Duplicate writes, overwrites, or improper handling of late records can create an incorrect table without triggering an exception, making the problem harder for existing monitoring systems to detect.

The intervention results also caution against reading a single metric in isolation. For one model, a hazard warning reduced silent failures from 48.2% to 10.5%, but increased crashes from 9.0% to 37.0%. When both outcomes were counted as failures, the overall rate moved only from 51.0% to 44.0%. In other words, an intervention may convert invisible corruption into visible instability rather than eliminate the underlying risk. The authors independently reran all 11 arms, with rates moving by no more than 5.9 percentage points.

For developers, this points toward explicit tests for idempotency, ordering, retries, and convergence to a recomputed state. For benchmark designers, delivery schedules and temporal state transitions should be part of the default evaluation rather than an optional stress test. The paper’s broader message is straightforward: data agents should be judged not only by whether they run once, but by whether they remain correct as time changes the input.

arXiv

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
HyperBrowseComp Turns Web Research into a Multilingual, Multimodal Stress Test
Evaluation & Benchmarks
cctest.ai

HyperBrowseComp Turns Web Research into a Multilingual, Multimodal Stress Test

HyperBrowseComp is a challenging benchmark for web-browsing agents, spanning 13 languages and several forms of evidence. Instead of testing whether a model can retrieve a familiar fact, it tests whether the agent can persistently discover, connect, and verify clues across the open web.

Read more