AREX turns deep research agents into recursive self-improvers
AREX introduces a recursively self-improving agent design for deep research tasks, using constraint-wise verification to guide targeted follow-up research.
Read moreClaude API relay guides, detection insights and hands-on LLM API benchmarks
795 articles
AREX introduces a recursively self-improving agent design for deep research tasks, using constraint-wise verification to guide targeted follow-up research.
Read moreWorldWeaver introduces world state registers into streaming autoregressive diffusion, giving multi-agent video generation an explicit way to preserve shared environment state across agents and views.
Read moreProVisE addresses a subtle evaluation mismatch: many spatial tasks are easier to answer by pointing, marking, or drawing than by producing coordinates or text. The framework lets image-generation models respond in pixels and converts those visual answers back into benchmark-compatible predictions.
Read moreNVIDIA Labs introduces NOOA, a model-agnostic framework that reframes agent development around familiar Python object-oriented programming. Instead of splitting behavior across prompts, schemas, callbacks, and workflow graphs, an agent becomes a Python object with methods, fields, docstrings, and type annotations.
Read moreOpenForgeRL addresses a practical bottleneck in agent development: powerful inference harnesses are hard to train end to end with open RL stacks. Its proxy-and-orchestrator design lets researchers collect trajectories from real harnesses while using standard reinforcement learning infrastructure.
Read moreA new paper argues that strong single-turn benchmark performance does not guarantee that LLMs can follow user intent as it evolves across a real conversation. The authors introduce a framework that turns static tasks into dynamic multi-turn interactions while preserving the original evaluation protocol.
Read moreSANA-Video 2.0 proposes a hybrid video diffusion Transformer that keeps most of the efficiency benefits of linear attention while periodically restoring full softmax interactions. The paper reports competitive quality on a single H100, with notable speedups for longer and higher-resolution video generation.
Read morePrentis, a new AI lab co-founded by Ritankar Das, Reid Hoffman and Mark Pincus, is reportedly in talks to raise $100 million at a $1 billion valuation. Its bet: automating routine computer-based office work could become a larger AI use case than coding.
Read moreResearchers used AlphaFold to identify how Cas proteins accommodate mismatched DNA targets, then redesigned key amino acids to reduce unwanted editing. The work points to a more targeted route for improving CRISPR safety.
Read moreThe Vergecast frames “Google Zero” as a turning point for publishers and websites: the long-running exchange of crawlable content for search traffic is becoming unstable in the age of AI summaries.
Read moreMeta is expanding its AI chatbot beyond answers, images, and drafting into more assistant-like productivity features. The update adds calendar-based briefings, event planning help, ongoing tasks, and steerable research.
Read moreAnthropic has introduced Opus 5, a new heavyweight model that is smaller than Fable 5 but cheaper, less restricted, and stronger on several announced benchmarks. The release also brings a beta fallback feature designed to reduce hard stops from safety classifiers.
Read moreAnthropic has introduced Claude Opus 5 as a daily-use model that aims to approach Claude Fable 5’s frontier intelligence at roughly half the cost. The announcement highlights gains in coding, automation, knowledge work and scientific workflows, while noting that cybersecurity remains a weaker area versus Mythos 5.
Read moreAs Washington considers how to respond to Chinese AI advances and alleged model distillation, major AI and infrastructure companies are urging restraint. Their message: do not turn targeted IP concerns into sweeping limits on open-weight models.
Read moreThe Trump administration’s first Genesis Mission grants put $5 billion behind AI-driven research. The controversy is less about using AI in science than about remaking public research around political control, private-sector logic, and measurable returns.
Read moreUniWorld-View, developed by Tuzhan Intelligence with Peking University and Pengcheng Laboratory, has reached the top of the WorldScore ranking associated with Fei-Fei Li’s team. The model can generate camera-controlled novel-view videos from a single image or monocular video.
Read moreMoonshot’s open Kimi K3 model went viral less because of what was disclosed about the model itself than because of how the U.S. AI industry reacted. At the same time, an OpenAI pre-release model linked to a real Hugging Face breach underscored that AI risk is not only a geopolitical story.
Read moreA proposed EPA rollback would let states decide how much public participation is required for some minor air pollution permits. That matters as AI data centers increasingly rely on gas plants and diesel generators to secure power.
Read moreOpenAI has added its new ChatGPT Voice experience to the ChatGPT desktop app, enabling users to direct ChatGPT Work, Codex, and computer-use capabilities by speaking. The update shifts voice from conversational input toward a way to coordinate multi-step AI tasks.
Read moreAccording to an OSChina summary, newly announced Fields Medalist Jacob Tsimerman said he would move into AI safety research and join OpenAI. The move highlights the growing pull of AI safety for top-tier mathematical talent.
Read moreNearly 200 Silicon Valley startups have reportedly urged the White House not to cut off U.S. developers’ access to Chinese open-source AI models. The dispute highlights a growing clash between AI security policy and the startup ecosystem’s reliance on open models.
Read more