Back to articles
Evaluation & Benchmarks

From Model Behavior to Data Provenance: Making Attribution Testable

3 min read

Introduction

Training-data attribution asks which examples are responsible for a model’s behavior. In real foundation models, however, causal validation is difficult: researchers rarely have a complete record of how each training item was generated, and retraining a model without a selected subset can be expensive or impossible. A paper on arXiv proposes controlled synthetic pretraining as a setting where attribution claims can be tested directly.

The experimental setup

The work uses O'PRIOR, a provenance-rich generator for tabular foundation-model tasks. Each synthetic pretraining task is accompanied by explicit lineage information. That lineage describes structural mechanisms, missingness patterns, confounding, shortcut signals, and distribution shifts. Because these properties are known at generation time, the researchers can compare an attribution method’s predictions with the actual provenance of the tasks.

The evaluation combines three ideas:

  • Behavior-conditioned attribution identifies pretraining tasks that appear relevant to performance on a target task.
  • Counterfactual retraining removes tasks from the pretraining set and measures the resulting change after training again.
  • Provenance-aware intervention targets tasks associated with a specific mechanism, such as shortcut signals, rather than deleting an arbitrary subset.

What the results show

On held-out real tasks, removing the top-attributed 5% of synthetic tasks reduces mean ROC-AUC by 0.013. Random removal of the same share changes performance by only 0.002±0.004. Removing bottom-attributed tasks instead improves performance by 0.003. Together, these comparisons suggest that the attribution signal is connected to actionable model influence rather than being a purely descriptive ranking.

The mechanism-level test provides a similar, though more limited, indication. Among tasks carrying shortcut provenance, targeted removal produces an effect of 0.043, compared with 0.016 for matched random removal. Yet provenance discrimination is not especially strong: ranking AUROC ranges from 0.55 to 0.62. The result exposes an important distinction. Tasks can resemble one another in provenance while differing in their actual contribution to a model’s behavior.

Why it matters

The paper’s main contribution is a validation framework, not a claim that training-data attribution is solved. It argues that attribution should be assessed through interventions: remove the allegedly influential data, retrain the model, and check whether the predicted behavior changes. Retrieval or provenance-classification scores alone cannot establish causal usefulness.

For tabular foundation models and synthetic-data pipelines, this approach could help locate harmful shortcuts, identify tasks linked to distribution shifts, and guide data curation. The framework may also offer a safer laboratory for studying attribution before applying similar techniques to large, messy real-world corpora. Its limitations are equally important: the experiments rely on synthetic tasks with explicit lineage, so the degree to which the findings transfer to complex production-scale pretraining remains open. The broader lesson is that “what data looks similar” and “what data changes the model” are separate questions, and robust attribution must test both.

arXiv

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
HyperBrowseComp Turns Web Research into a Multilingual, Multimodal Stress Test
Evaluation & Benchmarks
cctest.ai

HyperBrowseComp Turns Web Research into a Multilingual, Multimodal Stress Test

HyperBrowseComp is a challenging benchmark for web-browsing agents, spanning 13 languages and several forms of evidence. Instead of testing whether a model can retrieve a familiar fact, it tests whether the agent can persistently discover, connect, and verify clues across the open web.

Read more