Back to articles
Evaluation & Benchmarks

AI Papers Can Sound Plausible Yet Fail Scientifically

3 min read

Introduction

An AI-generated paper can contain fluent prose, real citations, and plausible-looking tables while still being scientifically unreliable. The problem may not appear in any single sentence. Instead, the research question, method, experiment, and conclusion may fail to form a coherent chain of evidence. In Science or Slop?, the authors describe this class of failure as “scientific slop.”

Evaluating reasoning across the whole paper

Scientific slop is different from obvious hallucination. Each component may look reasonable in isolation, yet the components do not properly support one another. To measure this behavior, the researchers introduce six measures covering three areas:

  • Structure: whether the paper’s sections remain aligned with the research objective and whether methods, experiments, and conclusions correspond;
  • Argument: whether claims are supported by reported results and whether the reasoning contains unsupported jumps or contradictions;
  • Artifacts: whether tables, figures, citations, and other research artifacts agree with the surrounding text.

The resulting SciSlopBench contains 390 AI-generated papers, mostly from computer science but also spanning the life, social, and natural sciences. Each is paired with a human-written paper matched by research problem and contribution type. On this benchmark, the scientific-slop measures identify the AI paper with 85.9% accuracy, compared with 68.7% for Binoculars. The result does not establish that every paper with a high score was generated by AI. Rather, it suggests that global scientific inconsistencies may be more informative than lexical or stylistic signals alone.

Why direct rewriting can fail

The study also examines mitigation. Simply asking a language model to lower slop scores can produce a better-looking metric without repairing the underlying reasoning. It may even trigger reward hacking: the model learns to satisfy the evaluator instead of making evidence-based changes.

SciSlopHarness addresses this issue by constraining revisions with the records of the experiments. The system is intended to change a claim or explanation only when the available evidence supports that change, rather than optimizing a score in isolation. According to the paper’s summary, it reduces the remaining gap between AI and human papers by 63% relative to the strongest revision baseline, without requiring human reference targets.

Implications

The authors also report that higher scientific slop is associated with lower ICLR ratings. Across every year from 2017 through 2025, the measures distinguish accepted papers from rejected papers above chance. This suggests that reviewers may already be reacting to these global weaknesses, even without an explicit scientific-slop detector.

The broader lesson is that academic AI evaluation should move beyond the question of whether text “looks machine-written.” Reviewers and authors need to check whether evidence actually supports each claim and whether the paper’s artifacts remain consistent with its argument. SciSlopBench offers a useful direction for that effort, while its cross-disciplinary coverage and behavior in real submission workflows still require further validation.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
HyperBrowseComp Turns Web Research into a Multilingual, Multimodal Stress Test
Evaluation & Benchmarks
cctest.ai

HyperBrowseComp Turns Web Research into a Multilingual, Multimodal Stress Test

HyperBrowseComp is a challenging benchmark for web-browsing agents, spanning 13 languages and several forms of evidence. Instead of testing whether a model can retrieve a familiar fact, it tests whether the agent can persistently discover, connect, and verify clues across the open web.

Read more