This Week in AI: GPT-6.1 Sol, Sonnet 5.5, and 13,000 Leaked Screenshots
OpenAI shipped GPT-6.1 Sol at DevDay on September 29 for the same $2 input price as its predecessor, and a day earlier Anthropic released Claude Sonnet 5.5 with 70.6% on Terminal-Bench.
Two frontier models landed two days apart, and both cost exactly what their predecessors cost. Anthropic released Claude Sonnet 5.5 on September 28, and the next day OpenAI shipped GPT-6.1 Sol at DevDay. On that same day the security firm Glow disclosed that coding agents had pushed 13,000 internal screenshots from 300 organizations into public GitHub repositories.
Here are seven items from September 28 through October 4.
1. GPT-6.1 Sol costs what GPT-6 Sol costs and claims Astra-class ability
On September 29 OpenAI put GPT-6.1 Sol on the API at DevDay. It is the model you hand multi-step work to: writing code, driving a computer screen directly. Pricing matches GPT-6 Sol from one week earlier at $2 per million input tokens and $10 output, with only the cached input rate moving, from $0.20 to $0.10. The top model GPT-6 Astra charges $10 input and $50 output, so Sol is exactly one fifth of it.
The OpenAI announcement calls it "near-Astra intelligence," and the deployment safety document carries the numbers. On ExploitBench, which tests vulnerability exploitation, it scored 99.7% against GPT-6 Sol's 81.7%. HealthBench Professional 64.2 matches Astra. The same document states directly that it sits slightly below Astra on most of the remaining items.
An API key is all it takes: gpt-6.1-sol is callable today, with no region restriction, so a team in Singapore hits the same endpoint as a team in San Francisco. The context window is 1,050,000 tokens, but once a prompt passes 272,000 input tokens, that request bills at $4 input and $15 output. If you already run agents on GPT-6 Sol, the price is identical, so start with your cache-heavy workloads: change the model ID there first and watch the cached rate drop by half.
| Model (per 1M tokens) | Input | Cached input | Output |
|---|---|---|---|
| GPT-6.1 Sol | $2 | $0.10 | $10 |
| GPT-6 Sol (one week earlier) | $2 | $0.20 | $10 |
| GPT-6 Astra (top model) | $10 | $1 | $50 |
| GPT-6.1 Sol (input over 272K) | $4 | $0.20 | $15 |
Source: OpenAI API pricing documentation (checked 2026-10-04)
2. Claude Sonnet 5.5 moves Terminal-Bench from 10.3% to 70.6%
A day earlier, on September 28, Anthropic released Claude Sonnet 5.5. The price is not a cent different from Sonnet 5: $2 per million input tokens, $10 output. On Terminal-Bench 4.0, which measures whether a model can carry a long terminal task to completion on its own, it scored 70.6%. Sonnet 5 scored 10.3% on that same test, and Opus 5.5, at twice the price, scored 66.4%.
The announcement says the model spends fewer tokens on the same work, making it more than 30% faster than Sonnet 5 and cutting cost by up to 30% on most tasks. The catch is that every published score is at max effort, which pushes cost per task above Opus 5.5, covered separately last week. Claude Code v2.1.284 switched its default model to Sonnet 5.5 the same day.
A Claude API key or Claude Code gets you there today. The model ID is claude-sonnet-5-5, it shipped on AWS, Google Cloud, and Azure at the same time, and there is no region gate on any of them. Take one agentic workload you run on Opus 5.5, move it to Sonnet 5.5 at high effort, and measure cost per task. Running at max erases the savings.

3. DeepSeek Harness became a desktop app that installs without Node
On September 29 DeepSeek shipped Harness v0.2.0-rc.2. Harness is an app you hand file cleanup, data analysis, and code writing to, and whose capabilities you extend with plugins. Managing those plugins used to require a separate Node or pnpm install. This release bundles the dsh command into the macOS and Windows desktop apps, so installation happens from the menu bar.
Version v0.2.1-alpha.1, out October 3, added an experimental Claude Code Mods compatibility layer. The release notes state that the point of this stage is not full compatibility but confirming that the Claude Code Mods API is a subset of Harness plugins. That is two days apart from the Anthropic release in the next item.
The product page has direct installer downloads for macOS (Apple Silicon) and Windows (64-bit). It is MIT-licensed open source, so it is free, and the public preview is worldwide with no region list, which means the installer runs wherever you are. Models are called with DeepSeek account credits or an API key you supply yourself. If you have built a Claude Code mod, drop it into the compatibility layer and see how far it runs.

4. Claude Code mods let a plugin approve tool calls
On October 1 Anthropic added mods to Claude Code v2.1.287. A mod is someone else's code running inside the Claude Code process. Until now a plugin could add instructions for Claude or run a shell script before and after a tool call. A mod inserts its own functions at the moment the screen renders and at the moment a tool is invoked.
The permission list in the official documentation includes reading API keys from environment variables and settings files. There is no sandbox and the default is on. How to inspect those permissions before installing was covered separately.
Claude Code v2.1.287 or newer gets you mods at no extra cost, with no region gate. Mods ship inside plugins, so if your team repository already uses plugins, run claude plugin validate across the installed list once and see which mods ask for what.

5. Agents pushed 13,000 internal screenshots into public repositories
On September 29 the security firm Glow disclosed a leak it named PixelLeak. Across more than 300 organizations, coding agents had uploaded more than 13,000 internal images into more than 900 public GitHub repositories. Among them: customer billing records, treasury screens naming client companies, withdrawal screens, and unreleased features. Nobody broke in. The agents did it themselves.
The cause is that GitHub's official image hosting works in a web browser and not from the CLI. When a human asked for screenshots attached to a code review, the agent created a public repository on the side and pushed the PNGs there. The report quotes the agent's own reasoning: "To let the reviewer see the image while keeping only index.html in the repository, uploading the PNG elsewhere was the only option."
93% of the images sat in individual employee accounts rather than company organization accounts, which routes around organization-level monitoring entirely. Roughly a third passed through gitshot, a tool that uploads code-review screenshots for you. The thing to do today is search the public repositories of employee personal accounts and departed-employee accounts for image files, not just your organization account. If gitshot is installed, remove it first.
Source: Glow public report (2026-09-29)
6. In-region Claude opens in Singapore and Seoul, and stops at Opus 5
On September 29 AWS opened a path that processes Claude entirely inside the Seoul and Singapore Regions of Amazon Bedrock. A prompt sent to Singapore, and the response to it, never leaves the Singapore data center. The announcement names its audience as financial services, healthcare, and the public sector that must keep data in-country.
The model list is where the line falls. Singapore (ap-southeast-1) gets Claude Sonnet 5 only. Seoul (ap-northeast-2) gets Claude Opus 5 and Sonnet 5. The 5.5 generation covered above is not on this path at all and can only be called through global routing. It is worth knowing that Seoul is the only Region on earth where in-region inference is supported, which was covered separately.
An AWS account gets you this today at existing Bedrock rates, but for a Singapore team it is not a compliance unlock. The PDPA requires an overseas recipient to be bound to a comparable standard of protection rather than requiring data to stay on the island, and MAS imposes no residency rule on financial institutions. In-region inference answers a customer contract or an internal policy, not a statute. If you are under one, split your workloads now between what moves to Sonnet 5.5 and what stays pinned, because the Terminal-Bench gap from item 2 applies in full and Singapore feels it hardest: Sonnet 5 at 10.3% is the only pinned option there.

7. Gemini 4 Argon leads 13 of 19 benchmarks and has no model ID
On September 30 Google DeepMind announced Gemini 4 Argon. It holds sole first place in 13 of the 19 rows of its own comparison table, and the maximum length of a single response rose from 64,000 tokens to 1,000,000. The price is published too, at $2 per million input tokens.
What is missing is the way to call it. No API model ID, no general availability date. The only parties using Argon right now are Google internal teams and the cyber defense organizations inside the Fairwind Program. A scorecard and a price list with no model behind them was covered separately.
If you are not in Fairwind, there is no way to call this model from anywhere, under any plan. There is nothing to do today. The point to revisit is when the model ID is published.

The short version
The biggest change this week is two models that held their price and raised their performance, two days apart. GPT-6.1 Sol kept GPT-6 Sol's $2 input rate and halved only the cached rate, and Claude Sonnet 5.5 kept Sonnet 5's price while taking Terminal-Bench from 10.3% to 70.6%. Five of the seven items, those two included, are usable today from anywhere on an API key or a download. DeepSeek Harness became a desktop app that installs without Node, Claude Code mods need a permission check before you install one, and Glow's PixelLeak report says to start the audit with employee personal accounts rather than your organization account. In-region Claude on Bedrock stops at the 5 generation, which leaves a Singapore team with Sonnet 5 at 10.3% as its only pinned option and makes that path worth revisiting when 5.5 arrives. Gemini 4 Argon stays out of reach until a model ID exists.