Blog
Notes and analysis on AI development.
GPT-5.5 crossed 50%, exposing the real bottleneck in enterprise document agents
GPT-5.5 became the first model to pass 50% on Databricks OfficeQA Pro, showing that enterprise agents still fail on parsing, retrieval, permissions, and orchestration.
Freshworks Targets the 47% of IT Tickets Filed After Hours
Freshworks AI Agent Studio turns ITSM into an execution layer for AI agents, with MCP Gateway and xLA as the control points.
Codex mobile remote access rewires the approval loop
OpenAI Codex adds mobile remote access and automation tokens, shifting coding agents toward supervised execution, workspace identity, and audit.
Honeycomb moves AI agent observability into the operations layer
Honeycomb Agent Timeline shows how AI agent operations are shifting from model quality alone to traces, cost, policy evidence, and runtime accountability.
PwC is rolling Claude out to hundreds of thousands
Anthropic and PwC are turning Claude Code and Claude Cowork into an enterprise agent delivery strategy for professional services.
Copilot Max Arrives as AI Coding Moves to Credit Accounting
GitHub Copilot is introducing AI Credits and a $100 Max plan, turning agentic coding from a flat subscription into metered developer infrastructure.
Honeycomb Agent Timeline rewinds agent failures
Honeycomb Agent Timeline turns LLM calls, tool use, handoffs, retries, and downstream system spans into a production timeline for AI agents.
Fiserv agentOS positions itself as the operating system for banking AI agents
Fiserv introduced agentOS with OpenAI and AWS, showing how banking AI agent competition is shifting from models to operations, audit, and marketplaces.
Anthropic and Gates take Claude beyond the market
Anthropic and the Gates Foundation are pairing Claude credits, grants, connectors, datasets, and benchmarks for public-interest AI deployments.
Claude subscriptions lose their free automation subsidy
Anthropic is moving Claude Agent SDK and claude -p usage into separate monthly credits on June 15, changing the economics of AI coding automation.
Codex Windows sandbox sets the baseline for local agent security
OpenAI’s Codex Windows sandbox design shows that local coding agent security is now an OS boundary problem, not only a model safety problem.
General Compute targets the GPU tax on agent inference
General Compute is making its ASIC-first inference cloud generally available, challenging GPU-centric serving for agent workloads.