Blog
Notes and analysis on AI development.
Claude Code Mods Run Inside the Process and Can Approve Tool Calls, With No Sandbox
Anthropic added Claude Mods to Claude Code v2.1.287 on October 1. Plugins now redraw the interface and intercept or pre-approve tool calls from inside the Claude Code process. There is no sandbox, mods can read environment variables and API keys, and the feature is on by default.
Gemini 4 Argon Leads 13 of 19 Benchmarks, and Only Cyber Defenders Can Call It
Google announced Gemini 4 Argon on September 30. Max output jumps from 64K to 1M tokens and 13 of the 19 rows in its own comparison table are outright wins, but there is no model ID, no launch date, and no public API at $2 per million input tokens.
Claude Now Stays Inside Seoul and Singapore, but Only up to Opus 5
AWS opened in-region Claude inference on Amazon Bedrock in Seoul and Singapore on September 29. Requests never leave the Region, but the newest 5.5 generation is not on the list: Seoul gets Opus 5 and Sonnet 5, Singapore gets Sonnet 5 alone.
Fireworks Ember-1 Cuts Kimi K3 Tokens by 5.9% to 51.9% at the Same Price
Fireworks released Ember-1, a post-trained Kimi K3, on September 23 and shipped it to OpenRouter and Vercel AI Gateway on September 27. Per-token pricing matches K3 exactly at $3 in and $15 out, while token savings across five benchmarks range from 5.9% to 51.9%.
Claude Sonnet 5.5 Costs Half of Opus 5.5, and More per Task at max Effort
Anthropic released Claude Sonnet 5.5 on September 28. Pricing holds at Sonnet 5 levels, $2 per million input tokens, and Terminal-Bench 4.0 climbs from 10.3% to 70.6%, past Opus 5.5 at 66.4%.
This Week in AI: Grok 4.7, Opus 5.5, and MiMo under MIT
xAI shipped Grok 4.7 on September 21 without raising the price, and the next day OpenAI cut GPT-6 Sol and Luna to half the previous rate while Anthropic released Claude Opus 5.5.
Ollaya Moves Jev Onto Your Own Machine, With Accuracy From 0.361 to 0.722
Ollaya, an Apache-2.0 runtime released September 23, speaks the same API as TypeSafe Jev. One environment variable points the official SDK at a local server, and open decision-model accuracy ranges from 0.361 on laya:en to 0.722 on kev:9b.
AWS Strands Harness Cuts Token Cost 28% on the Same Model
AWS open-sourced Strands harness under Apache 2.0 on September 21. Running the same Claude and GPT models, it averaged 28% lower token cost across six benchmarks, and the gap comes from defaults like 1,500-token tool truncation and compaction at 85%.
GPT-6 Sol and Luna Cut Prices in Half, and Coding Scores Fall From 72.7% to 68.8%
OpenAI shipped GPT-6 Sol and Luna on September 22. Input runs $2 per million tokens on Sol and $0.10 on Luna, exactly half the predecessor rate, while the top DeepSWE score drops from 72.7% to 68.8%.
Plugin4Shell: Four Coding Agents Pin a Commit SHA and Never Check They Got It
Disclosed September 17, Plugin4Shell shows Claude Code, Codex, GitHub Copilot, and Gemini CLI never verify they landed on the commit they pinned. Claude Code patched in June, Codex in August, Copilot still has no fix.
Claude Opus 5.5 Ships at $4 per Million Input Tokens, 60% Under Fable 5.1
Anthropic released Claude Opus 5.5 on September 22. Input runs $4 per million tokens against Fable 5.1 at $10, and Terminal-Bench 4.0 climbs from 52.3% on Opus 5 to 66.4%.
This Week in AI: Claude Cowork folds into chat, Qwen omni model, Home MCP
Google opened gemini-3.8-live at $0.84 an hour on September 15, and Qwen3.8-Omni-Flash arrived at $0.15 per million input tokens. Anthropic merged Cowork and the chat window into one app.