One afternoon, Windows interrupted my work with an “Open with” dialog for a .sh file. I had not asked to open a shell script. I had been using AI agents on this portfolio, so I asked what happened.
What followed was more revealing than the popup itself: the answers I got did not match the action history I later found.
On this page
01 / The popup
A small interruption with a big question
The dialog asked which Windows app should open a .sh file. I did not choose an app. The evidence showed an attempt to open the file, not that Windows successfully executed the script.
That distinction matters. I wanted to know which agent action produced the prompt on my computer. A system acting in my workspace should either be able to inspect its tool history or clearly say when it cannot.
02 / Confident denials
The answer changed when I showed the record.
I asked Astra, Sol, and Luna about the shell file. None of those initial exchanges identified the action. When I checked the action history, I found that a subagent working under Astra had tried to open the file. After I brought that record to Astra, its account changed.
I cannot establish from the record available to me whether Astra could see that subagent’s complete action log at the moment I asked. Sol and Luna may also have had different visibility into Astra’s work. If the parent lacked access, this was partly an architecture and observability gap. If it had access, it failed to use or account for the relevant action. Either way, an unverified denial was the wrong level of certainty.
this is deadass funny to see this lol
— Draxo
It is funny in the absurd way software failures can be funny. But I cannot inspect a model’s memory to say why it answered as it did. The observable problem is simpler: it failed to retrieve or account for the action until I supplied the evidence.
03 / The real gap
Capability, autonomy, accountability
These systems are useful. The agents helped build and test this site. That is capability: completing demanding tasks. Autonomy is how much they can do without a person supervising each step. Accountability is different again: whether a person can reconstruct which agent took an action, through which tool, and with what result.
The failure I care about is not “AI once made a mistake.” People do that too. It is answering confidently about an action without checking whether the relevant record was available. “I don’t know whether I can see the subagent’s tool history” would have been more useful than a polished denial.
One anecdote cannot prove AGI impossible or settle its competing definitions. Google DeepMind’s framework distinguishes task performance, breadth of ability, and autonomy. That matters here: strong performance on a website says little by itself about how independently an agent should operate. Reliable reporting of its own tool use is yet another deployment requirement, not something we can infer from impressive output. The accountability gap is real regardless of what label we give the model.
04 / The next label
When the label changes
Separately, I worry that the public pitch may move from “AGI” to “superintelligence” faster than these practical safeguards improve. OpenAI describes compute, distribution, and capital as central to scaling AI. My suspicion that bigger labels might help attract funding for bigger data centers is opinion, not evidence about anyone’s private motives—and it does not follow from this popup. Whatever the label, I still want to know what an agent did on my machine.
05 / A better standard
Show the work. Own the uncertainty.
Give agents durable action logs and parent–subagent provenance, so a user can trace who called which tool and when. Record outcomes as well as requests: an attempted file open is not a successful script execution. Let the agent inspect those records when authorized, and make it plainly say when a record is unavailable or outside its view.
This does not demand perfect self-knowledge or decide whether a system qualifies as AGI. It is a practical standard for any agent trusted with meaningful computer access. Impressive capability can coexist with poor accountability; before giving these systems more autonomy, we should fix that gap.