A 73% Agent Report Card Ends the Model-Only Benchmark Era
Open Agent Leaderboard evaluates full agent systems, not just standalone models, combining architecture, tools, cost, and failure behavior.
New model releases, benchmarks, evaluation methods, and research results.
Open Agent Leaderboard evaluates full agent systems, not just standalone models, combining architecture, tools, cost, and failure behavior.
OpenAI’s counterexample to the Erdős unit-distance conjecture shows both the promise of AI research automation and the reproducibility gap left by an unnamed model.
Google Pics brings Nano Banana image generation into Workspace with object-level, text-level, and collaborative precision editing.
OpenComputer shifts computer-use agent evaluation from LLM judges to reproducible desktop tasks and app-state verifiers.
IBM Research and Hugging Face’s Open Agent Leaderboard evaluates AI agents as systems, including harnesses, costs, and failure modes.
Cohere Command A+ is an Apache 2.0 open model aimed at enterprise agents, private deployment, and the practical cost of sovereign AI.
Alibaba Qwen3.7-Max is not just a model launch. It packages agents, custom chips, 128-accelerator racks, and cloud runtime into one stack.
Google added Street View grounding to Project Genie. The world-model race is moving from prompts toward real spatial data and responsibility boundaries.
Gemini 3.5 Flash is not just another fast model release. It points to the cost, latency, and routing fight behind coding agents and AI search.
Google Gemini API Managed Agents move model calls into isolated Linux sandboxes with stateful agent execution.
A May 13 arXiv study measured 55K Google searches and 98K AI Overview claims, showing where citations, ranking, and publisher economics diverge.
Mistral 3 packages a 675B MoE model with 3B, 8B, and 14B edge models under Apache 2.0, shifting open AI competition from benchmarks to deployment.