Devlery

Blog

Notes and analysis on AI development.

Nvidia 550B Beat the Top Human at IOI 2026, and Retrying Did More Than Training

Nvidia 550B Beat the Top Human at IOI 2026, and Retrying Did More Than Training

Nvidia Nemotron-3-Ultra-CC scored 535.4 at IOI 2026, above the 498.27 of the best human contestant. The paper breaks the score down, and most of it came from a test-time loop that resamples 200 candidates rather than from post-training.

Three Coding Agents Agreed on a Tool Only 42% of the Time

Three Coding Agents Agreed on a Tool Only 42% of the Time

Armature ran Claude Code, Codex, and Cursor through 16,893 sessions where each agent had to pick a third-party service and write the integration. The three agreed 42% of the time, and switching the repo language changed the winning email provider.

This Week in AI: GPT-6 Astra opens to the API, Nvidia buys Hugging Face

This Week in AI: GPT-6 Astra opens to the API, Nvidia buys Hugging Face

OpenAI opened GPT-6 Astra to the API and upper ChatGPT tiers on September 3, and the day before Nvidia signed a $12.93B deal to acquire Hugging Face. Meta launched a tier at $0.10 per million tokens if you let it train on your prompts.

One PR Cost This Coding Agent $41, and Only 106K Tokens Were New Code

One PR Cost This Coding Agent $41, and Only 106K Tokens Were New Code

Sonar instrumented its own coding-agent session traces. An 800-line PR took 512 model round trips, 152.8M cache-read tokens, and roughly $41, while newly read code accounted for only 106K tokens.

K2 Horizon's 7B Hits 70.6 on SWE-bench, Its Training Repo Is One README Line

K2 Horizon's 7B Hits 70.6 on SWE-bench, Its Training Repo Is One README Line

MBZUAI-backed IFM released six K2 Horizon models under Apache 2.0 on September 3. The 7B beats Qwen3.5-9B on SWE-bench Verified at 70.6%, but the training-code repos and dataset links are still closed.

GPT-6 Astra Costs 2.5x More Per Token, and Less Per Coding Task

GPT-6 Astra Costs 2.5x More Per Token, and Less Per Coding Task

OpenAI shipped GPT-6 Astra on September 3 at $10 and $50 per million tokens, 2.5x GPT-5.6 Sol. On coding work it burns a third of the tokens Sol needed, and the independent Intelligence Index puts both at 61.

Gemini 3.8 Flash Matches Opus 5 on Coding, Then Scores 19.1% as a General Agent

Gemini 3.8 Flash Matches Opus 5 on Coding, Then Scores 19.1% as a General Agent

Google shipped Gemini 3.8 Flash on September 2 at the same $0.75 per million input tokens as 3.7 Flash. DeepSWE climbs from 65.3% to 73.7%, but general agent work lands at 19.1% against Claude Opus 5 at 51.8%.

What Changed in Claude Fable 5.1: Coding Goes 42% to 55.8%, Forced Tool Calls Return a 400

What Changed in Claude Fable 5.1: Coding Goes 42% to 55.8%, Forced Tool Calls Return a 400

Anthropic released Claude Fable 5.1 and Mythos 5.1 on September 1. Agentic coding climbs from 42.0% on Fable 5 to 55.8%, three beta features land in the API, cache reads cost a quarter of what they did, and three behaviors break on migration.

Your Claude Usage Drains While You Sleep, and Anthropic Named the Infostealer

Your Claude Usage Drains While You Sleep, and Anthropic Named the Infostealer

Anthropic began notifying affected Claude users on August 30 that infostealer malware copied their login sessions. Stolen cookies walked past two-factor auth, and Claude Code tokens live on a different settings screen than web sessions.

This Week in AI: Claude Code Restricted Mode, GLM-5.3 Under MIT, Qwen Open Weights

This Week in AI: Claude Code Restricted Mode, GLM-5.3 Under MIT, Qwen Open Weights

Z.ai released GLM-5.3-Flash 320B weights under MIT on August 26, and Alibaba published Qwen3.8-Flash-Next the same day. A day later Claude Code gained a --restricted mode that removes command execution and web access.

GPT Models Leave Cursor on November 12, and Your Own API Key Will Not Cover It

GPT Models Leave Cursor on November 12, and Your Own API Key Will Not Cover It

OpenAI served SpaceX a termination notice on August 28, closing direct access to GPT models inside Cursor on November 12. Cursor says those models are 5% of its traffic, and the bring-your-own-key workaround only reaches non-reasoning chat models.

41% of GitHub PR Descriptions Now Share One Voice, and load-bearing Is Inside Claude Code

41% of GitHub PR Descriptions Now Share One Voice, and load-bearing Is Inside Claude Code

Clustering 461,121 GitHub pull request descriptions by word usage alone, one cluster grew from 0.7% to 41%. Its top marker word, load-bearing, sits in Claude Code’s built-in prompt text, which is why banning it in CLAUDE.md does not work.