DVx Signal/Builder/Issue 008

Claude Code mods, Gemini Argon and Copilot computer use

Five updates on agent controls, longer outputs, desktop tasks, and the cost of model runs.

DVx Research4 min read

Mods can stop a risky Claude Code command

October 1 release · checked October 8Official source ↗
Anthropic's Blast Radius mod pauses a git reset command and shows two files that would lose changes

Anthropic’s Blast Radius example holds a destructive command and shows the files it would affect. View official source ↗

What changed

Claude Code mods are JavaScript or TypeScript hooks in plugins. They can observe an event, rewrite a tool call, deny it, or add interface elements. Claude Code 2.1.287 or later supports them.

Why it matters

A team can put a review step in front of a specific command without changing every agent prompt. But a mod runs with Claude Code’s machine access, so its source needs the same scrutiny as any installed package.

What to do

Start with Anthropic’s Blast Radius example in a disposable repo. Trigger a harmless dry run, inspect the affected-file list, and check where its command-text filter can miss an alias or script.

Argon can output far more, but most builders cannot use it yet

September 30 announcement · limited rollout · checked October 8Official source ↗
Gemini 4 ArgonAnnounced state
Maximum output1M tokens, up from 64K
Access nowSelected cyber defenders
Google’s announced limit and first access group. A longer output limit does not guarantee a complete project or correct patch. View official source ↗
What changed

Gemini 4 Argon raises Google’s output limit from 64K to 1M tokens, about 15.6 times as much. Google is giving selected defenders early access to test vulnerability discovery and patching before a wider release.

Why it matters

A single run can carry much more generated code or analysis. Today the practical decision is access, not model migration: Google has not set a broad release date.

What to do

Keep an evaluation ready for a long coding task, with tests and review checkpoints. Run it when the paid API or AI Ultra release reaches your account; do not design a production dependency around the announced limit yet.

Haiku 5.5 changes the price of small tasks

October 7 launch and price change · checked October 8Official source ↗
API price / 1M tokens≤100K / >100K prompt
Input$0.10 / $0.50
Output$0.50 / $2.50
Cache read$0.01 / $0.05
Anthropic’s standard API rates. The 100K prompt boundary changes every rate shown; task cost also depends on retries and output length. View official source ↗
What changed

Claude Haiku 5.5 is available now across Anthropic’s platforms for short, high-volume tasks. Anthropic also cut Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens on October 7.

Why it matters

The same model can cost five times more per input or output token once a Haiku prompt crosses 100K. Anthropic’s claims of roughly 75% lower Haiku task cost and 20% lower Sonnet agent cost are vendor averages, not a bill forecast.

What to do

Replay a labeled batch of summaries or classifications. Log accepted results, prompt lengths, output tokens, cache hits, and retries before routing routine traffic to Haiku.

Copilot can operate desktop apps in public preview

October 1 release · macOS and Windows · checked October 8Official source ↗
GitHub's demo of Copilot reading a Safari window and clicking through an expense-report task

GitHub’s demo shows the app’s computer-use actions during an expense-report task in Safari. View official source ↗

What changed

GitHub Copilot computer use lets the CLI and desktop app read app windows, click controls, and enter text across macOS and Windows apps. It can reach software with no API.

Why it matters

The demo shows a real interface task, but this is a public preview. Tool permissions govern app control; saved approvals can remove later prompts. An organization can disable it.

What to do

Try a disposable workflow with test data. Name the apps and the outcome, watch each action, and check the approval list before allowing an app again.

GPT-6.1 Sol may lower the cost of harder agent tasks

September 29 launch · Work, Codex, and API · checked October 8Official source ↗
API / 1M tokensGPT-6.1 SolAstra
Input$2$10
Output$10$50
Cached input$0.10$1
OpenAI’s standard API rates. Its near-Astra quality claim comes from OpenAI’s evaluations, not ours. View official source ↗
What changed

GPT-6.1 Sol updates the Sol model covered in Signal 007. OpenAI reports better coding, computer use, and document work; its standard input and output rates are one-fifth of Astra’s.

Why it matters

A model that finishes more tasks at the same rate can be cheaper overall. OpenAI’s benchmark gains still need checking on your own work.

What to do

Run the same small set of reviewed tasks on GPT-6 Sol, 6.1 Sol, and Astra. Compare accepted results and total tokens. It is in ChatGPT Work, Codex, and the API, not Chat.

Also since the last issue

Useful updates

28 days of shipping · ongoing

OpenAI’s Developer Community tracker began October 4 and changes daily. Check the current day and your account limits before relying on a reset or speed claim.

Copilot local sandbox · general availability

GitHub’s October 7 release lets teams restrict the files, network, and credentials that agent-run tools can reach. Check the effective policy before running a real task.

Copilot workflows · preview

GitHub’s October 1 release adds coded, repeatable agent workflows. Start with one release check that has a human review checkpoint.

One workflow to try

Test the cost per accepted task

Use ten already-labeled support tickets. Run the same classification prompt on your current model and Haiku 5.5. Save the accepted labels, mistakes, token usage, and total API cost. A cheaper token rate only matters if the completed task stays reliable.

Term of the week

Cache read

A cache read reuses prompt text the provider has stored from an earlier request. It usually costs less than fresh input, but only when the request actually hits the cache. Compare bills using observed cache hits, not the advertised cache rate alone.