At a conference last week, I sat through a bunch of software factory demos, and they were all about task management. Kanban boards, Slack interfaces, email interfaces, dependency graphs, ticketing, graphical interfaces, you name it.

Why is everyone building task management and calling it a software factory? Agents are slow to execute. And the obvious, easy fix to latency is to hide it by starting a new agent every time you get blocked.

But concurrency is really rough on humans. It’s stressful. It trashes flow state and thrashes our mental page caches.

Task management isn't a solution; it's a band-aid.

This is a call to arms. Let’s make it possible to be equally productive with fewer agents.

Bitter Lesson to the Rescue?

As models improve, “good enough” models will get ever faster. This will naturally reduce latency, and in turn concurrency. We’ll still have to do some old-fashioned engineering, like making tests run fast.

But we don't have to wait! Here are some things we’ve experimented with at exe.

Use a Fast Model for Talking With Humans

Start tasks with a powerful model. But instead of reading a wall of text and writing an essay in response, put all the comms in the hands of a fast, competent model like Luna 6. Give Luna no coding tools, and make its context window intentionally short and focused on what the human conversation requires. Then, when you go to check in on a task, you can have real-time discussions with Luna. With a fast feedback cycle, attention doesn’t wander, and you can stay engaged for longer. Eventually, Luna exhausts its ready information, and the beefier model takes back over, with rich, substantial user feedback to work from.

This works reasonably well. But we found that, shock of shocks, absorbing information via chat was rather constraining. It was frustrating to have communications be squeezed through a tiny pipe. Luna might be an exciting, bendy straw, but it’s still a straw.

Screen Recording

That led us to focus on the HCI aspects of the problem. Back when I coded by hand, I would stare intently at screens full of text, processing it slowly, navigating freely between files, thinking. But the UI that is presented by most coding harnesses is a single linearized text thread, typically interspersed with lots of irrelevant tool call noise. There’s very little human agency or control, and no organization. You can’t even rely on the most important content being at the bottom: noise from straggler subagents often drowns out the agent's primary response.

So: How can we restore human agency and optimize for human I/O, rather than making the human adapt to the agent?

For input, the human retina is a powerful information processing system, when given structure to work from. That is, not a wall of text. I can still pick a Go panic stacktrace out of terminal logs scrolling by at 30fps.

For output, even for the fastest typists, typing is typically slower than speaking. Also, typing requires coordination with lots of other parts of the computer. The text input window has to be focused. Your cursor must be in the right place. It constrains the other things you can do with your keyboard and mouse and what you can look at. Many coding agents, including Shelley, worked around this with affordances for annotating text, diffs, and web pages, but it still requires clicking, and it fundamentally requires cooperation from everything else.

We recently launched a simple show-and-tell feature in Shelley that is tailored to humans. Click the camera button. Shelley records audio so that you can think out loud as you explore, and it records your screen. Anything on screen is a thing that you can point at, talk about, refer to. Poke around, muse, backtrack, trail off, resume, change your mind, whatever. Agents are extraordinarily good at understanding rambling.

When you stop recording, Shelley transcribes all of your audio, including word-level timestamps, and makes a contact sheet from the video. It correlates your words with what was on screen when you said it. The computer does the hard work of collating all of the feedback.

In my experience, this form of interaction is typically deeper, richer, longer, and more engaged than chatting with an agent. This is particularly useful for UI work or anything visual.

DOM Recording

What if you're not working on a visual task? The same ideas can be adapted nicely to other forms of engineering work.

Another feature of human cognition is that we work well with concrete examples, not abstract descriptions. Agents are happy to be told, but humans prefer to be shown. Humans also benefit from diagrams, charts, and visually structured diffs.

With this in mind, I have been using a slightly different system that we haven't shipped in Shelley (yet). Instead of spewing all its output to me in raw text, the agent generates an HTML artifact. This HTML is standardized to reduce visual noise, ruthlessly reduce verbosity and complexity, structure information visually, and make navigation easy. Typical sections include notable decisions made, open questions, examples, and git diffs.

This HTML artifact includes a record button. When clicked, it records audio. Instead of using screen recordings, it uses a cheaper, more precise in-page DOM tracker JavaScript that tracks what is being viewed, scroll position, where my pointer is (if there is one), what text is highlighted, where I click/tap, all with timestamps. As before, it correlates timestamps with audio and feeds it all back to the agent.

I find that this reduces mental overhead; I am freed to focus on the content.

More Latency Hiding

Because really engaging with content takes time, there is an additional latency-hiding trick one can pull. I am still experimenting with this, but the idea is to take prefixes of my feedback while it is still being recorded, send them to the agent, and livestream its responses back to a tab on the very same HTML artifact I'm already looking at. By the time I’ve spent 10 minutes reading and probing, I will inevitably have asked questions and provided provisional guidance. Before I context switch away, some of my questions have answers waiting. Some of my design decisions have follow-up questions. And the camera is still rolling. This sneaks in an extra round trip with the agent without breaking my concentration.

There’s Still Work to Do

I am not yet at a point where I have the same net throughput with one or two agents that I do with many, but my attention is being restored to me, little by little.