
I was ill recently. For days on end I couldn't sleep, and I could barely focus—all at a critical point for the product and its market. I tensed up and forced myself to concentrate. I ended up working harder, getting less done, and finding it even harder to pay attention.

Agents are becoming more reliable at execution. The least reliable part of the whole system is the human. Human judgment has the highest ceiling, but its floor is tied to the body. Sleep, illness, and attention can all pull it down. And when it falls, the loop doesn't stop.
The more important point is that people have a high ceiling, but they don't stay there. Whether it's a personal agent or an agent team for work, the purpose is to support us at our worst and extend what we can do at our best.
Being the “boss” isn't so simple

Execution and judgment complement each other. The cheaper execution gets, the more valuable judgment becomes. Agents drive the cost of execution down, which gives judgment much greater weight: every decision I make sets off a whole round of work. If the direction is right, good execution amplifies it. If the direction is wrong, good execution only takes the mistake further.
That makes the decision itself more important than anything else. The more capable the people working for you, the less you can afford to make decisions casually. A judgment you haven't thought through, or an instruction without context, will still be carried out faithfully.
Where standards exist, they can carry the judgment. Where there's no precedent, the standard has to be set on the spot—and I'm the one who has to set it. There's a longstanding finding in automation research that Bainbridge called the “ironies of automation”: machines take away the easy parts, leaving people with precisely the hardest parts, the ones that resist being written as rules. In my loop, I've delegated everything I can. What I can't delegate is judgment. While I was ill, that was the part that drifted furthest off course.
Exercising judgment and going through the motions look like the same action from the outside. The difference is whether, before clicking Approve, I can say where this step is most likely to go wrong. If I can't, but let it through anyway, the loop keeps moving further in a direction that hasn't been properly weighted or corrected. Eventually, the whole project takes more time to clean up. To me, that's why jev took off: at these critical points, people don't reliably converge on a decision—or, put another way, the decision needs context or intuition. This part needs someone to design it and work it through fully. Only then can you really hand off a very long task and have it turn out as you intended.
AI is driver assistance, not self-driving

I'm working on a persistence feature: an agent loop that executes against a goal, with key checkpoints and a source of truth along the way. I want the development agent team to run on its own, without keeping a human in the loop throughout development.
Today's vibe coding feels more like driver assistance. The car can follow traffic, change lanes, and take corners on its own. Most of the time, it doesn't need intervention, but the human must always be ready to take over. When I step away, the agents don't stop. That's how the product is designed. They keep driving, possibly down the wrong road; different agents may even collide. Losing attention is like closing your eyes with driver assistance on.
When I used to do the work myself, my output fell as soon as I got tired. That was the first brake. When the models were weaker, errors were obvious: the code wouldn't run, the tests wouldn't pass. I had to stop whether I wanted to or not. That was the second brake. Now both have been released. When I'm tired, the agents don't slow down with me. They hand back just as much work as usual, looking just as complete and orderly.
The models are strong enough to pass every checkpoint. Even their mistakes are coherent. Checkpoints and the source of truth can remember how far we've gone and what the facts are. Each stretch of the road makes sense on its own; the drift only shows up in the whole. Only the person holding the goal can see which direction is more worthwhile. Inside the loop, the only thing speaking for that direction is the little bit of weighting I've put in. By the time I come to my senses, there's already a long stretch of work to undo.
Inside Shopify, AI output thrown over to a colleague without careful review is called a “slop grenade.” That's exactly what a tired person hands off. In a one-person loop, there is no colleague. The person catching it is your future self. A loop without attention or input about what matters is a truck full of slop. It looks as though work has been completed and executed, but apart from consuming compute, nothing has been gained.
An agent's “attention” can be reset with a click. A person's can't.

I've always thought quota resets are brilliant marketing. But for people, they're a terrible destroyer of attention. They turn compute into something that expires: when the allowance refreshes, whatever went unused feels like money that's been voided, and we instinctively want to use it up. The tool's supply schedule becomes our work schedule, or our schedule for using it. We started out using tools; now we live by their reset cycles. People who already struggle to find enough attention are the first to be scattered by them.
We've been deeply poisoned by resets and anxiety about compute. Living in constant FOMO, we reduce ourselves to another link in the loop, endlessly executing. The loop's to-do list decides how we spend the day. Compute can refresh once a week. Human attention has no reset button; it has to be slept back into existence, one night at a time. While I was ill, the compute kept refreshing as usual. My attention didn't.
I've written about this state on X: when I use AI, I can't get into flow. My mind keeps spinning, until what's left is no longer attention, just numb execution. I'm still sitting in front of the screen. My hands haven't stopped. But my judgment is gone. That's what I mean by AI Attention Syndrome. Once a person is folded into a tool's rhythm, they become a tool too.
We're always on guard against models gaming their metrics, yet we rarely notice that we're gaming our own usage quotas. When judgment breaks down, another round only creates rework for later. Stopping at least means we stop digging. Better to let the compute go, stop, have a good meal, and get a good night's sleep. Sometimes doing nothing is better.
People and agents set the ceiling. The system holds the floor.

Being ill showed me how low my judgment can fall.
When I'm in good shape, I can see things the agents can't: whether something is worth doing, which road will eventually become a dead end, and when it's time to turn around. That's the ceiling, and for now, humans still reach the highest. But the same person, after several nights of poor sleep, starts making worse judgments—and often can't tell how far they've slipped.
You don't have to be ill for this to happen. Vigilance research established long ago that when people watch an information source that rarely goes wrong, effective attention lasts only about half an hour. Most agent outputs are correct, and the one that's wrong looks just like all the others. To keep up with the loop, we gradually turn down our sensitivity, until confirming becomes a reflex.
So whether the person is absent, or present but already dulled, without someone truly supervising, it doesn't matter how much the system produces. The output has no value.
Human organizations faced this problem long ago. That's why they have rules and procedures. People get tired, fall ill, and become confused. Institutions protect the floor.
I want to take my judgment and boundaries and make them into a separate role. It would itself be an agent, making the trade-offs on my behalf each time: what can proceed freely, what must wait for my decision, and which lines must not be crossed. It doesn't have to be smarter than me at my best. It only has to be steadier than me at my worst. The ceiling stays with me: I decide the direction, the decisions without precedent, and the boundaries of this role itself. I hand the floor over to it, so it doesn't rise and fall with my physical condition.
An organization built to persist should preserve judgment as well as tasks and memory.
Human attention has to be paid upfront
To protect my attention, first I need to sleep well. Second, I need to break the loop down: spell out the feedback and standards for each task, then review the work checkpoint by checkpoint. That's how I can bring my attention back together. This human-in-the-loop approach works, but it takes too much out of me. I still want a complete handoff, so I can work on bigger, more ambitious things. In my recent experience, Opus 5.5 paired with Fable 5.1 can already handle most development scenarios. But the overall architecture and how to divide it into modules require deep thought and abstraction. I can't casually hand those over, or I'll still end up with a pile of shit code.
A complete handoff doesn't eliminate the need for attention. It means paying for it upfront. The key checkpoints, the source of truth, the scaffolding built from good and bad cases, and the role that makes trade-offs for me—all of these are attention I've paid into the structure while my mind is clear. Pay enough at design time, and I can let go at runtime. Pay too little, and someone still has to watch.
Whether I can hand something off now depends less on the model than on whether I have examples to work from. For something I've never done, I first have to exercise the judgment myself and record what counts as good and what counts as bad. Only then can I hand off the next round. That's also why the overall design is hard to delegate: what matters and where the boundaries belong have to be decided anew for every project. There are no ready-made cases to borrow.
Checkpoints also need to be designed for a person at their worst. Each one should ask a single concrete question, small enough that even a tired person can answer it correctly. Otherwise, as the checkpoints multiply and start looking alike, the person slips back into going through the motions.
Be quick, but don't hurry

Looking back, the first mistake was the approach itself: tensing up and forcing myself to concentrate. Be quick, but don't hurry. Being quick is a capability; I can delegate it to an agent. Hurrying is a state of mind; only I can keep it in check. Back then, I was hurrying. I was physically at the current step, but my mind was already on the next one, and on the next quota reset.
Enthusiasm and FOMO can both make it hard to stop, but they point in opposite directions. Enthusiasm is directed at the work itself; FOMO is directed at quotas and the fear of being left behind. I ran into the same thing when I was building my presence on X. The follow-for-follow and engagement routines were all about follower counts—the same thing as chasing quotas. I couldn't do it. When I turned instead to honest conversations with people whose vibe and wavelength matched mine, I found more flow.
When we're tired, the body tells us to stop. When we're numb, even that signal doesn't get through. I'd written the goals, checkpoints, and source of truth into my loop. But when to stop was always left, by default, to the clear-headed version of me at the checkpoint. Someone with their eyes closed doesn't notice that their eyes are closed. So the stopping rule has to be set while I'm clear-headed: when I reach a checkpoint, if I can't say where this step is most likely to go wrong, I stop.
It's okay to be tired. It's okay to stop for a good meal and a good night's sleep. I just can't let myself go numb.
A real reset for myself—one that lasts, and that I can sustain.
References
Sources added for this blog edition; the essay above is translated from the Chinese original.
- Lisanne Bainbridge, “Ironies of Automation” (1983) — the automation research discussed in the essay; the account above is a paraphrase.
- Tobi Lütke on The Knowledge Project — the interview in which he discusses “slop grenades” at Shopify.
- UK Civil Aviation Authority, CAP 737, p. 41 — background on vigilance during prolonged monitoring. The half-hour discussion concerns detection performance in specific vigilance tasks, not a universal limit on human attention or a study of AI-agent supervision.
Related reading
- The Interface Is the Seed: Notes Toward an AgentOS — on organizational context, interfaces, and human attention.
- When the AI Team Becomes the Business System — on shared context, delegated work, and human judgment.