Reasonable Application Security

Reasonable Application Security #82

Hey there,

As I continue my journey as an AI Native, I’ve been struggling to track work in progress when working with my squad of agents. I landed on Asana to track the work because I love it and have used it for over a decade.

The challenge arose when the agents took over my Asana.

Chris’s Take

I brought this upon myself. I wanted to track work in progress across my agent squad, and Asana is the tool I know best and really like.

Agent complexity began when I explored secure architecture and design while working with agents. I built my own process of using specification-driven development, test-driven development, and threat modeling for each architecture phase. Each phase requires pre- and post-diagrams and threat-modeling analysis before it’s clear for coding or can be marked done.

Agents breed complexity when you let them. I started with Architecture Decision Requirements for the project, and worked out some rules for how ADRs are used. They can’t be deleted or edited in place, but they can be superseded if things change later. The ADR concept was simple enough that the agents didn’t start overloading any ADR fields.

When we started writing specs, I found myself buried in spec amendments for every little design change, as the model tried to enforce what it thought were rules and best practices I wanted. I love the concept of the spec driving development, but if you let the models/agents handle the update rules, they get messy fast.

I titled this article “The Agents Took Over My Asana” because they did, just not in the way you might have expected. When I started using Asana to manage the work and included the agents, they wrote agent-specific comments and descriptions. And if you haven’t seen what they write, it’s not digestible for me, a simple human. The boards were so overrun that I couldn’t make heads or tails of anything.

At that point, I wrote rules for the repo to define how agents could use Asana.

Here’s an example from board-writing.md:

You are writing for Chris, not for the next agent. The board is his view of the work. Everything an agent wants to hand to another agent has somewhere better to go, and the routing table below says where. A card that cannot be read at a glance has failed at the only job it has.

Title — 80 characters, words only

A short plain-English sentence. No ticket number — the card's number is its ID field — and no brackets, paths, spec rules, walk numbers or references to other cards.

Description — three sections, current state, rewritten

Always these three, in this order: 

Summary: two lines saying what the card is about, in plain words.

Detailed Description: four to six plain sentences — what is wrong or what to
build, why it matters, and where it stands now.

Next Steps: how the card moves toward done, and what done looks like.

Write it for a reader who has not opened the repository. No file paths, line numbers, spec rule numbers, commit hashes or code identifiers; the detail an engineer needs goes in the pull request, the repo or a comment. A pull request number or a walk number may appear once where it helps. When a card refers to another card, use that card's ID-field number and say in a few words what it is.

I regained control of my Asana when I set up the repo-level instructions and reset all the agents to use them. From that point forward, I could use the Asana tasks to measure work in progress without wading through all the detail that agents love.

Worthwhile Security Reads

Three reads on what happens when agent capability outruns verification: more vulnerabilities, weak evals, runtime anomalies, and an AI agent incident that exposed the value of transparent disclosure.

  1. Forget the AI Slowdown—the Vulnerability Explosion Is Already Happening. Widely available AI tools are accelerating vulnerability discovery and putting additional pressure on already constrained security and open-source teams. Chris’s take: The numbers show a gigantic uptick in vulnerability disclosure and patch volume over the last few months. That is not automatically bad: Mozilla explicitly tied 271 Firefox fixes to AI-assisted hunting, and Oracle says AI-powered identification contributed to its record July release. AI is clearly increasing discovery capacity, even if the broader surge cannot yet be attributed solely to AI.

  2. Is your eval lying to you? — A string-match grader proves only that expected text exists, not that generated software builds, runs, or satisfies the requirement. Chris’s take: Good intro to evals if you’re still catching up on the concept.

  3. What we learned from being the first company to disclose an agent cyberattack — Hugging Face argues that the disclosure exposed the need for transparency, stronger defender access, and open tools to help investigate and harden systems. Chris’s take: The useful lesson is not that one lab had an incident. It’s that model access control can bite you when you need it most. Hugging Face says safety guardrails blocked legitimate forensic requests from the frontiers, so they had to use open-weight models to investigate the breach,

Podcast Corner

Application Security Podcast

How Agentic AI Fails—and Which Controls Actually Stop It
Petra Vukmirovic applies fault tree analysis from aviation and nuclear safety to AI agents, showing how failure paths, probabilities, and minimal cut sets can identify the controls that matter most.

WatchListen

Security Table

When Code No Longer Matters
The Security Table debates whether human-readable code and pull-request review still matter when AI can translate intent directly into machine instructions—and who stays accountable when nobody writes or reviews the code.

WatchListen

Where to Find Chris

I’ll be at OWASP Global AppSec USA in San Francisco, November 5–6. I’m moderating a keynote debate and recording a live episode of The Application Security Podcast from the conference. If you’re there, come say hello—and join us for the debate and live recording.

What did I miss this week? Hit reply and tell me.

— Chris