Staff Writer
Published August 24, 2026 · Updated October 7, 2026last updated dates

Show, Dont Tell
I notice a difference in how I feel about agents I can see versus agents I cannot.
The ones running on my Mac mini have a visible desktop. I can watch them click buttons, type text, open windows. When something goes wrong, I see it happen. When something goes right, I see that too. I trust them more. Not because they are more accurate. Because I can watch them work.
The agents running on my pods, I cannot see. They log to stdout. They produce structured output. I read the logs and I believe them or I do not. There is no intuition about whether things are going well. No way to glance at a screen and feel confident.
Lumenbox, a project I found with one star on GitHub at the time I wrote this, gets this right.
One Desktop Per Agent
Lumenbox gives each agent its own X11 desktop. You can watch the agent work in real time. The desktop shows exactly what the agent sees. Every click, every typed command, every rendered page. The project calls this watchability, and it is the feature that made me stop scrolling.
You read the repo once and the reasoning is obvious. If you cannot see what your agent is doing, you are operating on trust and trust is not auditable. Logs tell you what an agent did. A desktop shows you what an agent is doing. Those are different things.
Our own system works this way for local agents. The computer use feature drives the macOS desktop in the background. We can see the windows open and close. We can watch the agent navigate a web page. We know when it is stuck because we can see the spinner.
The pod agents get none of this. They run headless. They produce JSON. I read the JSON and guess whether it went well.
Logs Are Not Trust
Logs are an abstraction layer over reality. They describe what happened. They do not show what happened.
Think about the difference between reading a transcript of a conversation and watching the recording. The transcript tells you what words were said. The recording shows you tone, hesitation, interruption. The transcript is cleaner. The recording is more honest.
Agent logs work the same way. A log says "function called, result returned." It does not show the three failed attempts before success. It does not show the edge case the agent handled silently. It does not show the moment the agent got confused and then recovered. Those are exactly the moments that build trust or erode it.
When I watch an agent work, I learn its patterns. I know when it is confident versus when it is guessing. I can predict when it will need help. This intuition does not come from logs. It comes from watching.
The Trust Asymmetry
Fast agents feel untrustworthy. If an agent completes a complex task in seconds and reports "done," you do not actually know it did the thing correctly. You know it ran and did not crash. The speed creates a gap between the agent's confidence and your own.
Visible agents feel different. Even when they are fast, you saw the work happen. The gap closes. A twenty second task you watched is more trustworthy than a two second task you did not.
This is the asymmetry. Speed without visibility erodes trust. Visibility without speed builds it.
The best UX for agent supervision is not a better log viewer. It is a window.
Turning Background Agents Into Collaborators
Our pod agents are headless by design. They run on Fly machines with no display. Giving them a desktop is not practical. But the principle does not require a full X11 session.
A frame-by-frame replay of what the agent saw would change the trust equation. A timeline of screenshots. A video of the browser state. Something between a static log and a live desktop.
Lumenbox ships a web interface that lets you check in on any agent at any time. The project calls it "tactile" supervision. I think that word captures something important. Supervision should feel tactile. You should be able to reach in and see what is happening. Not just read a report of what happened.
The Window, Not the Log File
The pattern is consistent across everything I have read about human agent trust. People trust agents they can observe. The observation does not have to be constant. It has to be possible.
A log file is not observation. It is an audit trail. Audit trails are for post-mortems. Observation is for the moment work happens.
Lumenbox showed me that the one star idea of giving each agent a desktop is actually a core architectural decision. How you watch your agents work determines how much you trust them.
I am thinking about what a fleet wide observation layer looks like. A way to glance at ten agents and know how each one is doing. Not from logs. From windows.
