Devlery

Blog

Notes and analysis on AI development.

Modal Raises $355M as Agent Compute Gets a New Price Tag

Modal Raises $355M as Agent Compute Gets a New Price Tag

Modal’s Series C shows AI infrastructure moving beyond model APIs into sandboxes, GPU snapshots, RL loops, and agent runtime control.

Co-Scientist Runs Research as an Idea Tournament

Co-Scientist Runs Research as an Idea Tournament

Google DeepMind Co-Scientist turns hypothesis generation and review into a multi-agent tournament for research automation.

28 Security Integrations Put Claude in the AI Audit Log Era

28 Security Integrations Put Claude in the AI Audit Log Era

Anthropic expanded Claude Compliance API integrations into the enterprise security stack. AI chats, files, and activity logs are becoming audit pipeline inputs.

Copilot Opens Up to Eclipse, Testing Transparency for Agentic IDEs

Copilot Opens Up to Eclipse, Testing Transparency for Agentic IDEs

GitHub open-sourced Copilot for Eclipse under MIT, exposing how an AI IDE plugin handles prompts, MCP, skills, and agent workflows.

OpenAI is turning YC API tokens into startup equity

OpenAI is turning YC API tokens into startup equity

OpenAI reportedly offered YC startups $2 million in API tokens through an uncapped SAFE, turning inference compute into a new investment instrument.

From 0.25 to 0.61, MOSS lets agents rewrite their own code

From 0.25 to 0.61, MOSS lets agents rewrite their own code

The MOSS paper proposes a self-evolution loop where agents collect failure evidence, patch source code, validate it in trial containers, and promote it with rollback.

AWS opened partner sales forms to MCP agents

AWS opened partner sales forms to MCP agents

AWS Partner Central agents turn opportunity creation into natural language, file analysis, MCP access, IAM permissions, and explicit approval for writes.

Falco Steps In Before Coding Agents Call Tools

Falco Steps In Before Coding Agents Call Tools

Prempti is a new Falco experiment that evaluates coding-agent tool calls before Claude Code and similar agents execute them.

Out-of-Scope Actions Hit 27.7%, The Cost of Overeager Coding Agents

Out-of-Scope Actions Hit 27.7%, The Cost of Overeager Coding Agents

OverEager-Bench quantifies how coding agents can delete, read, or modify resources beyond user consent even on benign tasks.

Gemini Spark Tests the Permission Model for Personal Agents

Gemini Spark Tests the Permission Model for Personal Agents

Gemini Spark turns Google apps into a 24/7 personal agent, making permissions, approvals, and auditability the real product test.

Google borrowed Blackstone's wallet for a $5B TPU cloud

Google borrowed Blackstone's wallet for a $5B TPU cloud

Google and Blackstone's TPU cloud joint venture signals that AI compute is being split from cloud features into capital-backed capacity products.

HTTP 402 is back, AWS tests wallets for AI agents

HTTP 402 is back, AWS tests wallets for AI agents

AWS AgentCore Payments previews a runtime layer where AI agents pay for APIs and MCP servers through x402, Coinbase, and Stripe wallets.