Harbor Builds a Common Infrastructure for Agent Evaluation
Harbor Adapters brings more than 80 agentic benchmarks into a shared evaluation layer, while Harbor-Index distills 82 challenging tasks into a more affordable test suite.
Read moreClaude API relay guides, detection insights and hands-on LLM API benchmarks
98 articles
Harbor Adapters brings more than 80 agentic benchmarks into a shared evaluation layer, while Harbor-Index distills 82 challenging tasks into a more affordable test suite.
Read morePerfReasoning evaluates LLMs on hardware performance reasoning and analytical model generation. The results show that producing plausible explanations is far easier than building reliable, executable performance models.
Read moreHarvestBench turns an abstract safety question into a costly choice: an LLM agent can drive over an animal for free or pay fuel to steer around it. Results vary dramatically across models, and moral instructions make a striking difference.
Read moreA visually polished video can still violate object counts, spatial relations, or event timing. VeriPhy turns prompts into typed physical obligations and produces verdicts that remain traceable to the evidence behind them.
Read moreSpeech brain-computer interfaces are often evaluated under incompatible vocabularies, datasets, and metrics. A new measure called open-vocabulary mutual information, or OVMI, accounts for both decoding accuracy and the share of a user’s intended language that a system can support.
Read moreVideo generators can produce visually convincing motion while violating basic laws of physics. Principia evaluates them through calibration-independent relationships between objects in the same scene.
Read moreAFAC2026 brought together 5,027 teams and nearly 20,000 participants to tackle market analysis, financial documents, automated experimentation, and long-context agents. Its challenges show how financial AI is moving from isolated models toward deployable systems.
Read moreExecRetrieval tests whether code retrievers can rank an executable solution above near-identical buggy variants. Its results expose a substantial gap between finding the right code somewhere in the list and placing it first.
Read moreS³Gym is an interactive benchmark for testing whether language-model agents can test their own behavior, judge experience, and improve later decisions. Its results show that learning from experience is highly dependent on task structure and memory design.
Read moreA new study proposes knowledge-gated task construction to separate failures caused by missing domain knowledge from failures caused by weak execution or reasoning. Its experiments show a clear artifact-dependent effect on some tasks, while stopping short of claiming any post-training benefit.
Read moreEarlyEval uses intermediate agent behavior to predict whether a task is likely to succeed or fail before execution is complete. The approach reduces unnecessary steps and token usage while keeping changes to measured resolve rates small across three benchmarks.
Read moreSWE-bench problems are typically formal, structured, and information-rich, unlike the short and incomplete requests developers often make in practice. RealSWE shows that realistic inputs lower average resolution rates and can change the ranking of coding models.
Read moreE-Commerce Bench places LLM agents in a 365-day e-commerce operation covering sourcing, negotiation, sales, fulfillment, returns, and cash management. Its results suggest that maximizing profit is not the same as being a well-rounded business operator.
Read moreAgentJudgeBench evaluates LLM judges on dependency-driven tool-calling workflows rather than open-ended text preferences. Its results suggest that workflow difficulty, not judge scale alone, sets the main reliability limit.
Read moreSpanCalib-VLM combines a generative vision-language model with a multimodal sequence tagger to locate hallucinated text and recalibrate its confidence. On the English SHROOM-Visions evaluation split, the hybrid system balances span coverage with more reliable scoring.
Read moreA new FACE-Eval study shows that where a preference cue appears can be as important as what it says. Cues delivered through tool outputs or raw artifacts are more likely to influence models without being clearly reflected in their reasoning traces.
Read moreA single benchmark score can blend safety, reasoning, knowledge, and instruction-following signals. Ai2’s BenchMIRT uses multidimensional item response theory to separate those effects at the question level.
Read moreEASEL evaluates a capability that ordinary image benchmarks often miss: turning visual evidence into precise, iterative tool actions. Its central challenge asks an agent to paint a canvas until it matches a reference image.
Read moreLoopArena isolates the control loop around a coding agent and measures whether one model can guide another fixed model through long-running software tasks. Its results show that reliable long-horizon control remains difficult.
Read moreA new benchmark called GMA expands mobile-agent evaluation with seven open-source-based apps and 300 tasks ranging from atomic actions to multi-step workflows. Its results show that current agents struggle as complexity rises, while harness choices such as context retention and explicit state tracking can improve some demanding tasks.
Read moreA commit-pinned census of Inspect Evals shows that executable evaluation code does not automatically support the historical claim attached to its metric. Of 124 mechanically eligible units, 110 stopped before deterministic inference because evidence or semantic grounding was missing.
Read more