Back to articles
Evaluation & Benchmarks

Vals Wants to Make Real-World AI Evaluation the New Standard

3 min read

Introduction

AI models are improving faster than many traditional benchmarks can adapt. Public test sets and academic exams remain useful, but they also create an obvious weakness: developers can optimize directly against the questions. A high score may therefore signal good preparation for a benchmark rather than dependable performance in an unfamiliar workflow. Vals, a startup founded in 2024, is trying to move evaluation from abstract intelligence tests toward evidence of practical work.

Key points

  • Private test materials: Vals does not disclose its specific evaluation items, making it harder for model developers to train directly on the exam.
  • Domain-specific work: The company tests tasks connected to law, finance, and coding instead of measuring only general knowledge.
  • Positive and negative outcomes: Its research also touches mental health, cybersecurity, biosecurity, recursive self-improvement, and the law of armed conflict.
  • A procurement role: Companies pay to identify weaknesses and improve models, while buyers increasingly use evaluation results when choosing systems.

Co-founder Rayan Krishnan argues that older benchmarks often measure intelligence in an abstract form—for example, whether a model can pass a bar-exam-style test. That is not necessarily the same as producing a useful legal memo, reliable code, or a high-quality financial deliverable. Vals therefore frames the central question differently: can a model produce work comparable to what a human would deliver in a particular domain?

The approach also extends beyond capability scores. Vals wants to examine what might happen if a model were widely deployed, including possible negative consequences. That makes evaluation closer to a combination of quality assurance and risk analysis. The aim is not simply to identify the highest-scoring model, but to understand where a system works, where it fails, and what kinds of misuse or unintended effects may follow.

Business and industry implications

Vals sells evaluations to AI companies and other organizations. Paying to discover that a model performs poorly may sound counterintuitive, but reliable measurement can help teams compare versions, diagnose weaknesses, and decide whether a system is ready for a particular use case. The company says its revenue is now eight times what it was last year and that its staff has grown from eight people at the start of the year to 25. It also recently raised a $40 million Series A led by Andreessen Horowitz, following a seed round led by 8VC and Bloomberg Beta.

The startup has also launched a program focused on evaluating models for federal agencies. If AI becomes more embedded in public services, enterprise workflows, and financial markets, benchmarks could become more than marketing collateral. They may influence procurement, risk controls, and the way companies explain AI investments to stakeholders.

Vals still has to prove that it can become a genuine industry standard. Confidential testing may reduce benchmark gaming, but it also makes independent verification harder. Cross-industry comparability, risk measurement, evaluator bias, and reproducibility will all matter. A credible benchmark is not merely difficult; it must also explain its methodology clearly and demonstrate that its results correspond to behavior in the real world.

Source: TechCrunch AI

Comments

Checking sign-in status...

Loading comments...

Related articles