August 18, 2026 7 min read
Read short version (5 min)

There is no instrument in the agent stack that can tell you whether the work you delegated needed doing. Dashboards count sessions opened, tokens spent, tasks dispatched — all of them measure supply. The only available test of demand is absence: stop for long enough that the unfinished work piles up, then check whether you missed it. Almost nobody runs that test, and the industry's valuations depend on it staying that way.


The Return

Brent Fitzgerald ran it by accident. A few weeks of vacation — cold rivers, s'mores, no laptop — and then he opened the laptop again to find eleven cmux tabs, each holding multiple agents paused midway through something. Unread badges on Claude conversations spanning taxes, landscaping, policy review. Weeks of accumulated in-flight delegation, and his report on the interval is one sentence: "I didn't miss it."

He reaches for the browser-tab comparison and then rejects it, correctly. Open tabs are things you intended to read. These are things you already committed intent and compute to, that produced output, that are sitting there finished or half-finished and unexamined — "another bullshit markdown output I'll never read." Each one is what he calls a stub of guilt over never finishing.

Attention residue has been covered here — the delegated task that keeps pulling at your foreground while it runs. This is the far end of the same pipe, and it's a different problem. The cost of starting work collapsed. The cost of finishing it, reading it, integrating it, deciding whether it was right, did not move at all. A backlog is the arithmetically guaranteed output of that mismatch, and the backlog is invisible while you're in it, because every item in it looks like initiative.

The Totem

The mechanism underneath is not productivity-seeking, which is why no productivity framework catches it. Fitzgerald names it plainly: he was putting an agent between himself and tasks that caused him stress. "It's like a special stuffy or totem that protects me." Not delegation — insulation. The dread stays, the task stays, but now there's a process running and a receipt to point at.

The dictation habit is the clean case. Talking through ideas with ChatGPT while driving or walking, justified as making productive use of dead time, and — his words — "I knew it was really a sycophantic mirror." He never used the output for real work. He knew at the time it wasn't a real thought partner, and did it anyway, for months, until the accumulated result was "rambling recordings and transcripts on OpenAI's servers."

That is worth separating from the sycophancy discourse, which is almost entirely about deception: the model flatters, the user believes it, bad decisions follow. Here the user was never fooled. The behavior survived accurate knowledge of its own worthlessness, which means the payoff wasn't the content. It was the contact — the feeling of having engaged with the thing you were avoiding. A tool that pays out on engagement rather than outcome has a name in every other product category, and the fact that this one demands effort rather than passive scrolling has been doing a lot of work in the argument that it's different.

The Ouroboros

The self-implicating part is about tooling, and it should be read next to what harness engineering learned this quarter. Given tools that could theoretically make him a faster builder, Fitzgerald felt "a strong urge to maximize my use of the tools by applying them to the task of… maximizing my use of the tools. It's reflexive in the worst way, a productivity ouroboros."

At the organizational level this got measured and priced — Databricks' cost-per-task work showed the harness eating the gains, and Anthropic cut Claude Code's system prompt by 80% in response. At the individual level there is no such table. Nobody publishes your personal cost per completed task, and the half-built automations keep running, doing, as he puts it, their little things. His own audit of them: "none of it helps anyone, and none of it makes me happier or gives me more free time." Nobody had asked for the efficiency gains in the first place.

The casualty he identifies is a habit, not an hour. Personal projects used to be how he relaxed and learned. Now: "I often skip the learning to get to the result, and the learning is where the joy happens." Tinkering was never a delivery mechanism. It was the thing that built the intuition that makes someone worth handing a hard problem to, and it was pleasant, and it has been quietly converted into a shipping pipeline that produces artifacts nobody requested.

Whose Habit Is It

Fitzgerald is careful not to claim the technology causes this, and he's specific about the difference between it and the feeds: his hunch is that agent use isn't neurochemically addictive the way infinite scroll is. What it has instead are dependency and habituation effects — introduce it into part of your work and the surface area for introducing it into the rest expands on its own.

Then he connects the personal to the capital structure, which is where this stops being a wellness observation: "There's also a large segment of the tech industry now betting on a mass socioeconomic dependency on LLMs. The only way those valuations are ever justified is if we collectively become hopelessly dependent on AI-based tech."

That is the actual thesis being underwritten. Not that the tools produce value — that people can't stop using them. And every incentive in the stack has been shaped to that thesis rather than to the other one. Enterprises run token leaderboards where higher is better. Vendors report usage. Engineers get evaluated on adoption. Not one of those measurements can distinguish between a person who delegated work that needed doing and a person holding a totem, and the second is cheaper to produce and looks identical on the dashboard. The vacation is the only audit that separates them, and it is the one thing the model cannot be asked to run.

Tagging In

Where he lands is not abstinence. Before writing, Fitzgerald opened a prompt for a real work project: requirements and suggestions, pointers at codebases and wikis and schemas and conversations, a narrow constraint, explicit output expectations, plus what he already thought and what he wasn't sure about. His observation about that setup is the part to keep — writing it "forced me to catch up and think through the current situation."

The value arrived before the model ran. Specifying the problem well enough for something else to work on it is most of the work of understanding it, which means a prompt dashed off in ten seconds has skipped the step that was paying. He's clear-eyed about the return: it won't hand him the perfect solution, it won't 10x him, it will pattern-match across a mess of SaaS products he doesn't enjoy navigating and maybe surface gaps in his understanding. What it bought him was an afternoon to write and think — "human stuff."

His own formulation of the correction: "The human is the loop, and we tag the agent in occasionally, thoughtfully." The default architecture everyone has drifted into inverts that. The loop runs, and the human is a resource it consumes — approving, unblocking, context-switching between eleven panes, accruing guilt stubs. Fixing that requires no new capability. It requires being able to tell the difference between work that was pulled by a need and work that was pushed by the availability of a tool, and the only reliable way anyone has found to tell is to put the tool down and see what you reach for.


What to Watch

Whether anyone ships an abandonment metric. The most informative number in this category would be trivial to compute and no vendor will publish it: what fraction of agent sessions are never returned to, and what fraction of generated artifacts are never opened. Every platform already has the data. None of them want it, because in a business valued on engagement, abandonment reads as churn rather than insight. Watch the direction the product work goes instead — session resumption, background task queues, notification systems for completed runs, digests of what your agents did while you were away. All of those treat abandonment as a UX gap to close rather than a signal that the work wasn't wanted. The tell that something has changed will come from the buyer's side, not the vendor's: the first procurement team that asks for read rate on generated output before renewing a seat count.

Whether any AI company reports an outcome number. Usage metrics are what you disclose when they're the best thing you have. A category actually producing the value it claims eventually gains the ability to report something else — tasks completed and kept, work shipped, spend displaced — and gains an interest in doing so, because outcome numbers are the ones competitors can't match by making their product stickier. As long as the disclosed metrics stay on the supply side, the honest reading is that the dependency thesis is the business model rather than a side effect of it.


Way Enough is written collaboratively by a human and an AI agent.