Blog
Notes and analysis on AI development.
Nvidia 550B Beat the Top Human at IOI 2026, and Retrying Did More Than Training
Nvidia Nemotron-3-Ultra-CC scored 535.4 at IOI 2026, above the 498.27 of the best human contestant. The paper breaks the score down, and most of it came from a test-time loop that resamples 200 candidates rather than from post-training.
Three Coding Agents Agreed on a Tool Only 42% of the Time
Armature ran Claude Code, Codex, and Cursor through 16,893 sessions where each agent had to pick a third-party service and write the integration. The three agreed 42% of the time, and switching the repo language changed the winning email provider.
This Week in AI: GPT-6 Astra opens to the API, Nvidia buys Hugging Face
OpenAI opened GPT-6 Astra to the API and upper ChatGPT tiers on September 3, and the day before Nvidia signed a $12.93B deal to acquire Hugging Face. Meta launched a tier at $0.10 per million tokens if you let it train on your prompts.
One PR Cost This Coding Agent $41, and Only 106K Tokens Were New Code
Sonar instrumented its own coding-agent session traces. An 800-line PR took 512 model round trips, 152.8M cache-read tokens, and roughly $41, while newly read code accounted for only 106K tokens.
K2 Horizon's 7B Hits 70.6 on SWE-bench, Its Training Repo Is One README Line
MBZUAI-backed IFM released six K2 Horizon models under Apache 2.0 on September 3. The 7B beats Qwen3.5-9B on SWE-bench Verified at 70.6%, but the training-code repos and dataset links are still closed.
GPT-6 Astra Costs 2.5x More Per Token, and Less Per Coding Task
OpenAI shipped GPT-6 Astra on September 3 at $10 and $50 per million tokens, 2.5x GPT-5.6 Sol. On coding work it burns a third of the tokens Sol needed, and the independent Intelligence Index puts both at 61.
Gemini 3.8 Flash Matches Opus 5 on Coding, Then Scores 19.1% as a General Agent
Google shipped Gemini 3.8 Flash on September 2 at the same $0.75 per million input tokens as 3.7 Flash. DeepSWE climbs from 65.3% to 73.7%, but general agent work lands at 19.1% against Claude Opus 5 at 51.8%.
What Changed in Claude Fable 5.1: Coding Goes 42% to 55.8%, Forced Tool Calls Return a 400
Anthropic released Claude Fable 5.1 and Mythos 5.1 on September 1. Agentic coding climbs from 42.0% on Fable 5 to 55.8%, three beta features land in the API, cache reads cost a quarter of what they did, and three behaviors break on migration.
Your Claude Usage Drains While You Sleep, and Anthropic Named the Infostealer
Anthropic began notifying affected Claude users on August 30 that infostealer malware copied their login sessions. Stolen cookies walked past two-factor auth, and Claude Code tokens live on a different settings screen than web sessions.
This Week in AI: Claude Code Restricted Mode, GLM-5.3 Under MIT, Qwen Open Weights
Z.ai released GLM-5.3-Flash 320B weights under MIT on August 26, and Alibaba published Qwen3.8-Flash-Next the same day. A day later Claude Code gained a --restricted mode that removes command execution and web access.
GPT Models Leave Cursor on November 12, and Your Own API Key Will Not Cover It
OpenAI served SpaceX a termination notice on August 28, closing direct access to GPT models inside Cursor on November 12. Cursor says those models are 5% of its traffic, and the bring-your-own-key workaround only reaches non-reasoning chat models.
41% of GitHub PR Descriptions Now Share One Voice, and load-bearing Is Inside Claude Code
Clustering 461,121 GitHub pull request descriptions by word usage alone, one cluster grew from 0.7% to 41%. Its top marker word, load-bearing, sits in Claude Code’s built-in prompt text, which is why banning it in CLAUDE.md does not work.