Devlery

Blog

Notes and analysis on AI development.

10,000 Developers Say AI Coding Winners Are Being Decided by Satisfaction

10,000 Developers Say AI Coding Winners Are Being Decided by Satisfaction

JetBrains AI Pulse says Claude Code reached 91% CSAT and an NPS of 54 while GitHub Copilot growth stalled, pushing AI coding toward a best-of-breed market.

MiniMax M2.7 brings self-evolving training to low-cost agent models

MiniMax M2.7 brings self-evolving training to low-cost agent models

MiniMax M2.7 uses a self-evolution loop around OpenClaw, activates only 10B of 230B parameters, and challenges premium coding models on price, benchmarks, and licensing.

Claude Mythos Preview turns zero-day discovery into a controlled-release problem

Claude Mythos Preview turns zero-day discovery into a controlled-release problem

Anthropic is limiting Claude Mythos Preview to Project Glasswing partners after reporting large jumps in autonomous vulnerability discovery, exploit chaining, and cyber safety risk.

Stanford AI Index 2026 shows the paradox of 53% adoption and 40-point transparency

Stanford AI Index 2026 shows the paradox of 53% adoption and 40-point transparency

Stanford HAI published the AI Index 2026 report: generative AI reached 53% global adoption in three years while model transparency fell from 58 to 40.

GLM-5.1 tops SWE-Bench Pro as Meta closes its open-source era

GLM-5.1 tops SWE-Bench Pro as Meta closes its open-source era

China-based Z.ai released GLM-5.1 under MIT terms and topped SWE-Bench Pro with a 744B MoE coding model, sharpening the open-source versus closed-model split.

Meta Muse Spark closes the Llama open-weight chapter

Meta Muse Spark closes the Llama open-weight chapter

Meta launched Muse Spark as its first proprietary frontier model after Llama 4 lost trust, shifting MSL toward closed weights, Meta-scale distribution, and unclear developer access.

Claude Code Monitor turns coding agents into live log readers

Claude Code Monitor turns coding agents into live log readers

Anthropic added Monitor to Claude Code v2.1.98, letting Claude watch background command output and react to logs, CI status, and file changes while a coding session continues.

Cursor 3 turns the IDE into an agent workspace

Cursor 3 turns the IDE into an agent workspace

Anysphere launched Cursor 3 with an agent-first workspace, parallel agents, local-cloud handoff, Design Mode, and Composer 2 as Cursor shifts from editor to orchestration surface.

Qwen3.6-Plus beats Claude on Terminal-Bench, but closes the flagship model

Qwen3.6-Plus beats Claude on Terminal-Bench, but closes the flagship model

Alibaba released Qwen3.6-Plus for agentic coding with a 1M-token context window, free preview access, and a closed-source strategy that changes how builders should read Qwen.

One prompt injection can take over a server: four CrewAI CVEs expose the agent security gap

One prompt injection can take over a server: four CrewAI CVEs expose the agent security gap

Four CrewAI CVEs chain prompt injection into sandbox escape, RCE, SSRF, and arbitrary file reads, showing why AI agent frameworks need fail-closed security.

Microsoft launched three MAI models in one day, and OpenAI dependence is no longer the default

Microsoft launched three MAI models in one day, and OpenAI dependence is no longer the default

Microsoft released MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2 across speech transcription, voice generation, and image generation, turning its OpenAI backup plan into a product stack.

StepFun Step 3.5 Flash reaches frontier-class scores with 11B active parameters

StepFun Step 3.5 Flash reaches frontier-class scores with 11B active parameters

StepFun Step 3.5 Flash activates only 11B parameters inside a 196B MoE model, posts strong math and coding benchmarks, and ships as Apache 2.0 open weights.