Blog
Notes and analysis on AI development.
React Doctor hits 11.1k stars as AI React audit layer
React Doctor adds a post-generation audit loop for React code written by coding agents, scanning state, effects, performance, security, and accessibility.
The same PR, a 12.5x bill: what coding agents really cost
Joule Index V0.1 adds dollars, joules, and public traces to coding-agent benchmarks, shifting the question beyond accuracy alone.
SRE agents stall at 47% on real Kubernetes incidents
Artificial Analysis ITBench-AA shows that even leading SRE agents remain below 50% on Kubernetes root-cause analysis, exposing the reliability gap in operational AI.
OpenAI Adds Live Vote Counts to Election Answers
OpenAI outlined its 2026 election safeguards, combining AP vote counts, voting information, Codex Security, SynthID, usage policy, and political bias evaluations.
620,000 attacks expose a 35-point safety gap in reasoning models
TELUS Digital tested 34 AI models with more than 620,000 adversarial attacks. The benchmark shows why enterprise AI safety is now an operating discipline.
90% of PRs are agent-built, Warp exposes the new bottleneck
OpenAI and Warp show that the coding-agent race is shifting from code generation to open-source verification, observability, and agent orchestration.
Furiosa and Broadcom are designing a 2nm token factory
FuriosaAI and Broadcom’s third-generation inference chip plan shows how agentic AI is shifting the bottleneck from raw GPU speed to token density, networking, and power.
Copilot Studio Now Clicks Apps Without APIs
Microsoft Copilot Studio computer use GA moves UI automation agents from demos into enterprise deployment, audit, and governance.
One API Call Now Boots Linux, the Line Google Drew with Gemini Agents
Google Managed Agents extends the Gemini API from model calls into sandboxed execution, shifting where agent infrastructure begins.
The 7,000-return loop behind Codex self-improving agents
OpenAI Tax AI shows why production traces, eval sets, and practitioner feedback matter more than agent automation alone.
From Inbox to DCF, Why Codex Is Moving Beyond Code
OpenAI Codex use cases now span inboxes, data, finance, QA, app automation, and collaboration, a sign that coding agents are becoming work agents.
Agent memory moved into files, and AMP shows a different path from OpenAI
OpenAI Agents SDK memory and the AMP v0.1 draft turn long-term agent memory into files, Git history, MCP resources, and auditable state.