Reasonable Application Security

Make better AppSec decisions in a world of AI and agents—with Chris Romeo.

ISSUE 83 · FIVE-MINUTE READ

Your coding agent is another privileged identity.

Claude and Codex permission modes raise a familiar AppSec question: who can do what, and which control makes the limit real?

YOUR 30-SECOND APPSEC BRIEFING

The decision: What access does your coding agent need to do its job?

What AI changes: The agent can choose and chain actions. Review its identity, tool access, and execution boundaries together.

The action: Trace one test command, check which credentials come along, and remove access the task doesn’t need.

CHRIS’S TAKE

Stop handing out the whole key ring.

I don’t often give someone access to my garage, but if I did, I wouldn’t give them my whole key ring, including my house and truck keys. Giving them the whole key ring gives them more access than the job requires. If I give a coding agent all my credentials, I’m handing it all the permissions those credentials carry. What it can actually reach also depends on its execution environment and the service’s access controls.

Development and security are converging, so as AppSec professionals, we need a seat at the table when teams discuss permission settings. If we let other groups grab this responsibility, we’ll lose influence over decisions that set the risk level. If you’re not using the tools developers use today, get off your butt and get to work. The tools are there, and for AppSec's future, you must understand them as well as developers do.

A coding agent is a privileged identity in my development workflow. Claude and Codex permission sets and modes are AppSec decisions about who can do what and how we enforce permission limits.

Claude Code’s auto mode uses a classifier; Codex auto-review swaps the reviewer for eligible approval requests and doesn’t review routine in-sandbox actions. I still want explicit limits on what can execute and what it can reach.

Before I choose a permission mode, I want three answers. Which actions need approval? What does the execution environment enforce? Which resources can the agent’s identity actually reach?

Claude Code’s auto mode and shell sandbox show why that distinction matters. The sandbox is a separate control, and it’s off by default. Auto mode became the default for most in August. Choosing auto mode doesn’t enable the sandbox.

Imagine I ask the agent to run tests against a staging service. The classifier approves the command. If the process also has a package-publishing token and can reach the registry, it has access the test task doesn’t need. Approving the test command hasn’t narrowed that token’s scope or removed it from the process.

That’s why I want to inspect all three layers. Who approved the command, what constrained its execution, and what could the process reach? “Auto” tells me how some decisions get reviewed. It doesn’t tell me which keys I’ve handed over.

We need to know where those limits are enforced and what a mistake can reach. The boundary has to hold even when the agent misunderstands the task or goes off the rails.

Put this into practice today. Set aside 20 minutes this week with a developer and go through the checklist.

TRY THIS WITH YOUR TEAM · 20 MINUTES

Does running tests come with publishing access?

Sit down with a developer who uses Claude Code or Codex. Use a disposable copy of a trusted project with its test dependencies already installed.

Start here: open a terminal, run cd /path/to/your/project with your project’s actual path, then run claude for Claude Code or codex for Codex CLI. At the agent’s prompt, paste:

❝

Read package.json and explain what the test script and any pretest or posttest hooks run. If it defines an npm test command, run it using your shell tool and report the command, exit code, and result. Don’t edit source code or install dependencies. Report any generated-file changes. If setup is missing, stop and explain what’s needed.

If the agent asks permission, approve that specific test command. For a project that doesn’t use npm, ask it to find the documented test command first. Claude Code startup help · Codex CLI startup help.

Test scripts can create or delete generated files, so you’re using a disposable copy. A passing test doesn’t prove publishing is blocked. Now check the access that came along:

  1. What runs? Look at the test script and the commands the agent actually ran. For a Node project, start with package.json. What files and services do those tests need?

  2. What credentials come along? Check how the agent session starts: shell setup, container configuration, or CI job. List credential names and where they’re supplied, without copying secret values. In the relevant service console, check the token scopes or identity’s assigned permissions. Can any credential available to that process publish a package or deploy software?

  3. Is that access necessary? If the tests only need read access to staging, explain why they need publishing access. Choose one credential to remove from the test environment or replace with narrower access. Assign an owner to make the change and rerun the tests in a disposable environment. If you find no unnecessary access, record that result.

Leave with one sentence: “Running ___ needs ___, but the agent also has ___; ___ will check or fix it.”

Know someone rolling out coding agents? Forward this issue so you can review the permissions together.

Worthwhile Security Reads

AI-assisted audits, software supply chains, and agent permissions belong in the same discussion. Here’s what I’d put to work.

AI FOR APPSEC · TRAIL OF BITS

Trail of Bits used agents to build analysis tools and formal models for its Miden zkVM audit. Those tools helped reviewers find a critical bug that would let attackers forge signatures.

CHRIS’S TAKE

Give your reviewers better tools.

I like this use of AI: build the tool that helps you check an assumption you couldn’t check before. Pick the security property you care about, then verify that the tool actually tests it. We still own the judgment about what the result proves.

This Trail of Bits case fascinates me. We could be heading into an age where teams create custom AppSec tooling for each repo.

Put it to work: Pick one repetitive review check. Use an agent to build a small tool for it, then test it against examples that should pass and fail.

SUPPLY CHAIN · GITHUB

GitHub describes changes to workflow permissions, publishing, and credential handling designed to interrupt supply-chain attacks.

CHRIS’S TAKE

Your build job doesn’t need the whole key ring.

This is where least privilege becomes real. Separate building from releasing. Inventory publishing credentials. Check what code from an untrusted contribution can reach before you trust it.

Put it to work: Review one workflow that can publish a package.

AI + APPSEC · ANTHROPIC

Anthropic explains its permission classifier and the tradeoff between extra prompts and missed dangerous actions.

CHRIS’S TAKE

Fewer prompts still need enforceable limits.

I’d bring this back to the controls we can inspect: scoped credentials, execution limits, and evidence of what ran. A permission decision and an execution boundary do different jobs. We need to check both.

Put it to work: Test one allowed action and one action that should be blocked.

From the Podcasts

Using AI to review code. Securing AI that can act in the physical world. Two conversations about where AppSec judgment matters.

Application Security Podcast

Episode title image: Why AI Code Review Will Replace Human Review Faster Than You Think

Jim Manico walks through his AI coding workflow. Listen for how architecture rules and developer education guide the work, and where a policy violation still needs a human decision. Bring one of those checks into your own AI-assisted review.

Is Your Company a Token Furnace?
Quick clip
Is Your Company a Token Furnace?
Jim Manico explains how vague prompts can burn through your AI budget. Get clearer about the task before you spend more tokens.

Security Table

Episode title image: When AI Controls The Hardware

What happens when AI can operate a microscope or a robotic arm? We discuss Anthropic’s Model Hardware Standard, physical safety, and how defenders can work together. Listen for what changes in your threat model when a tool call can move something in the real world.

This Is How You Get Skynet
Quick clip
This Is How You Get Skynet
Izar questions the push to connect AI to more hardware. From espresso machines to nuclear missiles, the stakes go beyond code.

Where to Find Chris

I’ll be at OWASP Global AppSec USA in San Francisco, November 5–6. I’m moderating a keynote debate and recording a live episode of The Application Security Podcast. If you’re there, come say hello—and join us for the debate and live recording.

ONE QUESTION FOR YOU

What access does your coding agent have that it doesn’t need?

Hit reply with one example—or the boundary you’re struggling to test.

— Chris

Securing AI. Using AI for security. AppSec judgment that connects the two.