Blog
Notes and analysis on AI development.
Composer 2.5 shows Cursor training for reward hacking
Cursor Composer 2.5 shows the coding-agent race shifting from benchmark scores toward long-task failure points, targeted feedback, and reward-hacking detection.
Google AI Overviews exposes the gap behind citation cards
A May 13 arXiv study measured 55K Google searches and 98K AI Overview claims, showing where citations, ranking, and publisher economics diverge.
ECHO makes stderr part of the coding agent world model
Microsoft Research ECHO turns terminal output into a direct learning signal so coding agents can learn from failed logs, not only final rewards.
Copilot remote control turns coding agents into an operations layer
GitHub’s May 18 Copilot updates link remote control, low-cost models, CI repair, and audit APIs into a control plane for coding agents.
Genkit Middleware Moves Agent Control Outside the Prompt
Google Genkit Middleware shows how retries, fallback, tool approval, and filesystem boundaries are moving into the runtime layer of agent apps.
arXiv one-year bans show the trust cost of AI citations
arXiv scrutiny of AI-generated manuscripts is not a blanket LLM ban. It is a warning about hallucinated citations entering research infrastructure.
Copilot LTS model makes coding agents enterprise infrastructure
GitHub Copilot has moved GPT-5.3-Codex to the default model for Business and Enterprise. The bigger story is not speed, but lifecycle, billing, and governance.
Dust Raises $40M as Enterprise AI Hits a Teamwork Bottleneck
Dust’s Series B is a signal that enterprise AI is moving from personal chatbots toward shared workspaces where people and agents operate together.
GitHub Copilot app turns coding agents into PR operators
The GitHub Copilot app technical preview moves coding agents from IDE assistance into issues, verification, pull requests, and merge follow-through.
Why OpenAI replaced certificates after signed malicious packages
The TanStack npm incident reached OpenAI Codex and ChatGPT Desktop certificate rotation, showing how AI development tools now inherit supply-chain trust risk.
PolyAI 10-Minute Voice Agents Expose the New Bottleneck in Contact Center AI
PolyAI opened its Agentic Dialog Platform. Raven, Agent Builder, and ADK show why voice-agent competition is shifting from speech quality to operations.
Mistral MCP Connectors turn agent integrations into a control plane
Mistral Studio Connectors move MCP from app-specific wrappers into central registration, direct calls, and approval-ready agent operations.