Language Models Need Sleep, Fast Weights Target the Long-Context Bottleneck
Language Models Need Sleep proposes sleep-like fast weight consolidation, shifting long-agent memory from prompt compaction toward model execution schedules.
New model releases, benchmarks, evaluation methods, and research results.
Language Models Need Sleep proposes sleep-like fast weight consolidation, shifting long-agent memory from prompt compaction toward model execution schedules.
Trump’s delayed AI executive order shows frontier model launches being reshaped around speed, security evaluation, and critical infrastructure readiness.
Google Co-Scientist and Gemini for Science shift AI research tools from answer generation toward hypothesis loops that humans can test.
NVIDIA released tri-mode diffusion LLMs that switch between AR, diffusion, and self-speculation generation in one checkpoint.
Gemini 3.5 Flash is no longer just a fast chatbot model. It reframes Flash as an agent execution engine and changes how developers calculate cost.
Open Agent Leaderboard evaluates full agent systems, not just standalone models, combining architecture, tools, cost, and failure behavior.
OpenAI’s counterexample to the Erdős unit-distance conjecture shows both the promise of AI research automation and the reproducibility gap left by an unnamed model.
Google Pics brings Nano Banana image generation into Workspace with object-level, text-level, and collaborative precision editing.
OpenComputer shifts computer-use agent evaluation from LLM judges to reproducible desktop tasks and app-state verifiers.
IBM Research and Hugging Face’s Open Agent Leaderboard evaluates AI agents as systems, including harnesses, costs, and failure modes.
Cohere Command A+ is an Apache 2.0 open model aimed at enterprise agents, private deployment, and the practical cost of sovereign AI.
Alibaba Qwen3.7-Max is not just a model launch. It packages agents, custom chips, 128-accelerator racks, and cloud runtime into one stack.