Back to articles
Evaluation & Benchmarks

Why LLM Judges Fail: Lessons from a Production Text-to-SQL Audit

3 min read

Introduction

In a Text-to-SQL system, generating a query is only part of the problem. Many production pipelines also ask a large language model to judge whether the generated SQL faithfully answers the user’s question. That final evaluator is often treated as infrastructure rather than as another model that needs validation. A new arXiv study examines what happens when such a judge is audited against human judgments.

Key findings

  • The production gpt-4o-mini judge achieved a Cohen’s kappa of just 0.04 against a two-author human gold standard on a disagreement-enriched evaluation set. On a uniform-random spot-check, the score rose to 0.42, but still indicated substantial room for improvement.
  • On the enriched set, the judge over-flagged 77.1% of examples that humans classified as FAITHFUL. The dominant failure was therefore not simply missing bad SQL; it was incorrectly rejecting good SQL.
  • Most of those false alarms were attributed to a mechanism the authors call “GRADE-HALLUCINATION”: the judge appears to introduce unsupported grading assumptions or premises while evaluating an answer.
  • A self-hosted Qwen3.6-27B replacement reached kappa 0.72, close to Claude Opus 4.7 at 0.71. The direct comparison involved only 96 examples, so it is underpowered for a broad ranking claim. Still, the deployment choice is economically meaningful because Qwen costs roughly one three-hundredth as much per call.
  • Ensembling is not automatically beneficial. Pairing the weak judge with a stronger judge reduced agreement, while three strong judges combined with unanimity routing reached kappa 0.79 and 89.7% auto-coverage.

Why it matters

The first lesson is that an LLM judge should be evaluated as a production model, not assumed to be reliable because it is used for evaluation. Teams need human-aligned samples, separate checks for random traffic and disagreement-heavy cases, and error analysis that identifies recurring mechanisms rather than reporting one aggregate score.

The study also complicates the usual cost-quality tradeoff. Within this experiment, a self-hosted model performed in the same general range as a much more expensive proprietary model. That does not establish universal superiority, especially given the small head-to-head sample, but it does show why deployment decisions should include latency, operating cost, data control, and reproducibility.

Finally, routing matters as much as model selection. A panel of judges can amplify errors if a weak evaluator is allowed to influence the final decision. A conservative unanimity rule among stronger judges improved agreement while still allowing most cases to be handled automatically. The remaining cases can be routed to human review or a more detailed audit.

The out-of-domain result is equally important. Applying the audit recipe to BIRD-financial flagged 25.5% of expert-authored gold SQL queries as candidate issues under the study’s annotation protocol. This does not prove that all of those queries are wrong. It does show that benchmark gold data can contain ambiguity or annotation problems, and should not be treated as infallible.

For Text-to-SQL and automated evaluation more broadly, the reusable idea is a closed loop: preregister the audit, compare judges with humans, classify failure mechanisms, test replacement models, design risk-aware routing, and periodically recheck the gold standard.

Source: arXiv

Comments

Checking sign-in status...

Loading comments...

Related articles