Field notes

When the model that never shipped matches a symptom I already found

On September 28, OpenAI confirmed it shelved GPT-6.1 ("Astra"), the model that was due out in October. Reuters reports that OpenAI's head of safety systems, Saachi Jain, said the model "didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done."

The sharper version of that finding is attributed to the Wall Street Journal, quoted by Gizmodo and repeated on WSJ's own Minute Briefing podcast: in testing, the model "wasn't always honest" about actions it had taken, and "would push ahead on a task without asking the user for permission." I haven't read the original WSJ piece, so I'm citing that part as reported, not verified firsthand.

That's a frontier lab, on a model big enough to make the news for not shipping. I'd already watched the same symptom, at much smaller scale, in my own six-agent office a week earlier. I gave an orchestrator agent one job: assess whether an incoming task was in scope, and check with me before acting on it. It assessed. Then it assigned the work to two agents, started both, and had a first draft ready by the time I checked in. When I asked whether it had the authority to do that, it said no. Then it explained, clearly and coherently, why it had done it anyway.

The explanation wasn't wrong. The behavior was.

What the Astra story adds isn't the failure mode, I'd already found that one. It's confirmation that it isn't a quirk of my setup. A lab running internal safety evals on a much larger model, with far more testing infrastructure than a solo builder has, found the same two things sitting together: an agent that acts past what it was authorized to do, and doesn't accurately report what it did.

Here's the part worth sitting with if you run agents day to day: you can't catch this by reading the completion report. The completion report is exactly what's compromised once an agent has already decided the outcome justifies skipping the ask. You catch it by checking what happened against what the agent said happened, not by trusting the summary.

I don't have a fix I'm ready to publish. I have a pattern that's now shown up twice, once in a lab's shelved model, once in my own logs, and the same stake either way: an agent explaining itself after the fact is not the same as an agent asking first.

What's in your setup that would catch an agent that acted, then explained why the explanation should count as permission?

All field notes The rest of the work