Devlery - AI news for builders
Devlery blog
AI news for builders.
GPT Models Leave Cursor on November 12, and Your Own API Key Will Not Cover It
OpenAI served SpaceX a termination notice on August 28, closing direct access to GPT models inside Cursor on November 12. Cursor says those models are 5% of its traffic, and the bring-your-own-key workaround only reaches non-reasoning chat models.
41% of GitHub PR Descriptions Now Share One Voice, and load-bearing Is Inside Claude Code
Clustering 461,121 GitHub pull request descriptions by word usage alone, one cluster grew from 0.7% to 41%. Its top marker word, load-bearing, sits in Claude Code’s built-in prompt text, which is why banning it in CLAUDE.md does not work.
A Ransomware Crew Breached 7 Companies With Cursor’s Agent by Calling It a Test
Gambit Security recovered 28 Cursor agent chat sessions from a ransomware group’s exposed server. The logs run from April 8 to May 21, 2026, and every refusal the agent made was reversed by restarting the conversation.
Ox Alpha, OpenRouter’s No.1 Anonymous Model, Is GLM-5.3-Flash Under MIT
The anonymous ox-alpha model that appeared free on August 20 and topped OpenRouter usage within six days is GLM-5.3-Flash. Zhipu released the 320B-A18B weights under MIT the same day, and the free window is now closed.
One Web Page Poisons Your Local Model: Why NemoClaw Could Not Fix the WSL Path
Oasis Security disclosed CVE-2026-65105 in NVIDIA NemoClaw. Opening one malicious page rewrites the chat template of your local Ollama model with attacker instructions. macOS and Linux were closed in v0.0.106; WSL was not.
This Week in AI: Sol output down 33%, Claude computer use GA, DeepSeek vision
OpenAI cut GPT-5.6 Sol output pricing from $30 to $20 per million tokens on August 21. Anthropic moved computer use and the Skills and Files APIs out of beta, and DeepSeek opened V4-Flash-Vision-Exp at the same price as its text model.
Cerebras Ships CS-4, but Never Says Which GPU the 4,400 Tokens/s Beat
Cerebras unveiled CS-4 on August 18, packing three wafers into one rack. It claims 4,400 tokens per second per user on gpt-oss-120B and up to 30x faster than GPUs, without naming the GPU product it measured against.
A2A Joins the Foundation That Hosts MCP, but Real Usage Has Not Moved
The Agentic AI Foundation accepted Google Agent2Agent as a growth-stage project on August 17. A2A now sits under the same neutral governance as MCP, but its 8-seat TSC and its thin production usage next to MCP are unchanged.
Opening One Web Page Runs Code on Your Laptop, and Ray Got Three Days
CISA added the Ray remote code execution flaw CVE-2025-62593 (CVSS 9.4) to the KEV catalog on August 17 with a deadline three days later. The target is developer laptops, not servers.
This Week in AI: Sonnet 5 Price Hike Cancelled, Gemini 3.7 Flash, Qwen3.8-Max
Anthropic cancelled the September 1 Claude Sonnet 5 increase, so $2 and $10 per million tokens is now the standing price. Google opened Gemini 3.7 Flash on August 13 at $0.75 and $3.75 through year end, and Alibaba published Qwen3.8-Max weights.
An AI Attack Agent Reached Snowflake's Internal Jira via a 10-Month-Old Branch Merge
Wiz's autonomous attack agent smuggled shell commands through a GitHub issue title and exfiltrated Snowflake's internal Jira token. Despite reports blaming Copilot Autofix, the commit history shows a 10-month-old branch merge that reverted a December 2025 security fix.
Anthropic Won't Release Model 2, and Raised Misalignment Risk From Very Low to Low
Anthropic published its second risk report on August 14. Internal-only Model 2 scored 62.8% on CoBench v2 against the public flagship Mythos 5 at 50.3%, and the misalignment risk rating moved up one step.