Devlery - AI news for builders
Devlery blog
AI news for builders.
Meta Ships Muse Code, and Its 21x Cheaper Tier Trains on Your Repo
Meta opened the Muse Code terminal agent in beta on August 5. Muse Spark 1.2 lost to Claude Opus 5 on all four benchmark charts Meta published itself. What is actually new is a contributor tier at $0.20 per million output tokens.
Mistral Opens a 3B Moderation Model That Ties GPT-OSS-Safeguard 20B
Mistral released Shieldstral 1.0-3B under Apache 2.0 on August 4. It scores 84.9 aggregate F1 on text safety, tied with a 20B model, and 83.8 on multimodal, ahead of every guard model tested. There is no API yet.
Six AI-Invented SQLite CVEs Scored 9.8 on NVD, Then Got Withdrawn
JFrog tried to reproduce six SQLite CVEs and none of them held up. The functions they blamed did not exist in those versions, and the proof-of-concept code never crashed. MITRE rejected all six on July 31, but the GitHub Advisory DB still shows 9.8 Critical.
This Week in AI: OpenAI Cuts Output Price 80%, DeepSeek Lands in Codex, EU Enforcement Starts
OpenAI cut GPT-5.6 Luna output pricing from $6 to $1.20 per million tokens on July 30, and DeepSeek opened V4-Flash in public beta with official Codex integration docs the next day. EU AI Act GPAI enforcement began August 2.
OpenAI Astra Cracks 10 Open Math Problems, Machine-Checked but Not Yet Peer-Reviewed
On August 1 OpenAI published proofs for 10 math and TCS problems open for a decade or more, produced by an internal version of the unreleased Astra model. It shipped a 249-page paper plus Lean 4 certificates, priced the tokens at roughly $2,000, and left the repo marked agent-reviewed.
EU AI Act Fines for GPAI Start August 2, High-Risk Rules Pushed to December 2027
The EU AI Act's high-risk obligations slipped to December 2027, but from August 2 the Commission can fine GPAI providers up to 3% of global revenue.
DeepSeek Ships an Official Codex Setup Guide, at $0.28 per Million Output Tokens
DeepSeek moved the V4-Flash API into public beta on July 31, added native Responses API support, and published its own documentation for wiring the model into OpenAI Codex.
An AI Agent Ran a Real App Business for 24 Hours: 5 Users, $0 Revenue
Bottleneck Labs handed a GPT-5.6 Sol agent a live App Store app and $350 in cash. After 24 hours it had booked no new revenue, cut the price six times, and ended at negative $97.
Gemini Managed Agents Can Now Deny Tool Calls, and Run on a Free API Key
Google added environment hooks, a token budget cap, cron triggers, and free-tier access to Gemini API Managed Agents on July 28. Your validation script now runs inside the remote sandbox and can block a tool call outright.
132 Companies Signed the Open Weights Letter. Anthropic Asked for Distillation Enforcement Instead
Nvidia and Microsoft led an open-weights letter now signed by 132 companies, and Anthropic is the one US frontier lab missing. Amodei countered with chip export controls, industrial-scale distillation enforcement, and mandatory pre-release safety testing for every model.
Microsoft Cuts Security Agent Costs in Half, but GPT-5.4 Still Handles the Hard 10%
Microsoft shipped MAI-Cyber-1-Flash, its first cybersecurity model, on July 27. A small in-house model takes up to 90% of the work and routes the hard 10% to GPT-5.4, halving cost and scoring 95.95% on CyberGym.
OpenAI Partner Network Turns AI Deployment Into a Consultant Race
OpenAI announced a Partner Network, $150M in ecosystem investment, and a goal to train 300,000 certified consultants for enterprise AI deployment.