Devlery - AI news for builders
Devlery blog
AI news for builders.
US Advisory Names Six Chinese AI Firms, Then Tells Providers to Quietly Degrade Their Answers
NSA, CISA and the FBI published advisory AA26-251A on September 8, naming DeepSeek, Alibaba and four others. The recommended mitigation is not an account ban but an unannounced drop in answer quality, and the detection signatures do not separate a distillation campaign from a production agent running around the clock.
GPT-6 Astra's 99.9% Came From OpenAI's Own Harness; the Neutral One Says 62.7%
ARC Prize ran GPT-6 Astra on ARC-AGI-3 under two harnesses and got 62.7% and 99.9%. Same model, same games. The only difference was whether reasoning state survived between calls. The leaderboard will now carry both numbers.
Nvidia 550B Beat the Top Human at IOI 2026, and Retrying Did More Than Training
Nvidia Nemotron-3-Ultra-CC scored 535.4 at IOI 2026, above the 498.27 of the best human contestant. The paper breaks the score down, and most of it came from a test-time loop that resamples 200 candidates rather than from post-training.
Three Coding Agents Agreed on a Tool Only 42% of the Time
Armature ran Claude Code, Codex, and Cursor through 16,893 sessions where each agent had to pick a third-party service and write the integration. The three agreed 42% of the time, and switching the repo language changed the winning email provider.
This Week in AI: GPT-6 Astra opens to the API, Nvidia buys Hugging Face
OpenAI opened GPT-6 Astra to the API and upper ChatGPT tiers on September 3, and the day before Nvidia signed a $12.93B deal to acquire Hugging Face. Meta launched a tier at $0.10 per million tokens if you let it train on your prompts.
One PR Cost This Coding Agent $41, and Only 106K Tokens Were New Code
Sonar instrumented its own coding-agent session traces. An 800-line PR took 512 model round trips, 152.8M cache-read tokens, and roughly $41, while newly read code accounted for only 106K tokens.
K2 Horizon's 7B Hits 70.6 on SWE-bench, Its Training Repo Is One README Line
MBZUAI-backed IFM released six K2 Horizon models under Apache 2.0 on September 3. The 7B beats Qwen3.5-9B on SWE-bench Verified at 70.6%, but the training-code repos and dataset links are still closed.
GPT-6 Astra Costs 2.5x More Per Token, and Less Per Coding Task
OpenAI shipped GPT-6 Astra on September 3 at $10 and $50 per million tokens, 2.5x GPT-5.6 Sol. On coding work it burns a third of the tokens Sol needed, and the independent Intelligence Index puts both at 61.
Gemini 3.8 Flash Matches Opus 5 on Coding, Then Scores 19.1% as a General Agent
Google shipped Gemini 3.8 Flash on September 2 at the same $0.75 per million input tokens as 3.7 Flash. DeepSWE climbs from 65.3% to 73.7%, but general agent work lands at 19.1% against Claude Opus 5 at 51.8%.
What Changed in Claude Fable 5.1: Coding Goes 42% to 55.8%, Forced Tool Calls Return a 400
Anthropic released Claude Fable 5.1 and Mythos 5.1 on September 1. Agentic coding climbs from 42.0% on Fable 5 to 55.8%, three beta features land in the API, cache reads cost a quarter of what they did, and three behaviors break on migration.
Your Claude Usage Drains While You Sleep, and Anthropic Named the Infostealer
Anthropic began notifying affected Claude users on August 30 that infostealer malware copied their login sessions. Stolen cookies walked past two-factor auth, and Claude Code tokens live on a different settings screen than web sessions.
This Week in AI: Claude Code Restricted Mode, GLM-5.3 Under MIT, Qwen Open Weights
Z.ai released GLM-5.3-Flash 320B weights under MIT on August 26, and Alibaba published Qwen3.8-Flash-Next the same day. A day later Claude Code gained a --restricted mode that removes command execution and web access.