Essays
Notes from the agent coding (not quite) vanguard
Everything is weirder and more capable. Models are way better at many things, unattended coding in particular. But, to everyone’s dismay, that causes them to write in a tedious style of English.
There’s plenty of room for humanity in between the tokens.
1. Models & harnesses
Frontier lab models, despite their quirks, feel like they will graduate to boring technology soon. When the servers are healthy, you can rely on Anthropic and OpenAI’s models to get work done. Cursor’s in-house models (probably other labs too, I haven’t tried Kimi, GLM, etc.) are solid enough to pinch-hit for frontier models when the inference servers are buggy.
Harnesses, when they’re relatively bug-free, are basically just as reliable. You can use the frontier labs (I’m still on Claude Code, mostly) to get most sorts of stuff done. You could substitute Cursor, Pi, or most of the others to personal taste. And, personal taste is increasingly a thing across the harnesses. Claude Code is very much for the terminal people, Cursor is increasingly for the holistic product people, the others are for the customization enthusiasts.
Occasionally one model will make a noticeable leap over the others in capability, speed, or “workhorse-ness”. Occasionally, a model gains a unique capability that is a big advantage until it’s cloned into the others. I find those momentary blips don’t matter enough to switch between models and harnesses that often.
(Despite that, I’m a novelty junky and I still oscillate between Claude Code, Cursor, and a couple of other approaches more often than I would like to quantify.)
Optimistic prediction: there will be a lull in frontier model capability race in the next 6–12 months. It will either correspond with IPO quiet periods or increasingly bad optics about the development of increasingly autonomous models. I’m optimistic about this because I think we could also use a breather in this domain.
2. Skills & prompting
Models and agents are good enough, in the late summer of 2026, to focus on customizing and leveraging skills instead of the model of the month. Bonus, skills are portable across agents, and models, sort of.
For most folks, those living in the middle of the bell-curve (definitely not the Amp folks, who seem to exist in the future), I think this means the action is now about telling the agent what the task is and how to figure out if it’s done or did a good job. Prompting and verification.
I’ve had a lot of luck lately with simplifying my instructions and moving more of the details into (Markdown) documents. In general, a lot of what I had in my CLAUDE.md ten months ago feels like over-steering now.
That said, moving the quasi-deterministic “verbs” of your interactions with coding agents to skills still makes a lot of sense. In particular, to socialize usage patterns with teammates or codify how specific tools, e.g., Xcode or CSS, are used.
Optimistic prediction: most currently-popular skills will be trained into the models. Instead of /grill-me about the new login screen or /fix the bug where the button disappears evaporate into garden-variety English.
3. The human factor
A couple of weekends ago, I was tinkering with doing Claude code via iPad and remote control. But I remembered the smarter move is to avoid tokens on evenings, and probably weekends too.😇
Lately, I’m finding that the drive for software factories and sprawling specs feels too much like the boring kind of software development work. Nudging and steering an agent, asking questions and getting answers in the form of working code, that’s more fun. It’s not as surprising and energizing and sometimes-frustrating as human collaboration. But it doesn’t feel like doing rote homework problem sets, which must count for something!
Optimistic prediction: the models whose output is most concise and least cliché could prevail over more capable but inscrutable ones.
4. Even more predictions
In the coming months:
- More hobbyist/self-employed/solo developers will utilize >1 agent harness monthly subscription. For bouncing back and forth between whichever model is in the out-performing at the time. Or if one prefers “the vibe” of one model as a collaborative partner while using the other as a coding agent. Sometimes, to get extra token usage at $40/month instead of going all the way to $100/month. More reasons to carry subscriptions to two or more models will probably come along as use-cases evolve.
- As more developers carry access to more models, skills and workflows will gain traction that use models in an “adversarial” mode. That is, model A reviews model B’s plan, model B review model A’s code, etc. Or, model A/B/C all try to find bugs/security issues in existing code in a competition. I’m not sure how this will pan out, but I’ve heard of people having good luck with this currently.
- One of the frontier model labs could go public. We get a better idea about the economics of their operation, what token subsidies actually look like in numbers, the actual economics of training and releasing a new model, and what the labs think their long-term play is currently. (Pessimistic: their filings and road-show are just “come with me, it will be good, but we’re doing it my way” like Facebook, and we don’t learn anything.)
- That developer with Pi kitted out just like they like it is the new co-worker who insists on vim (not NeoVim) or emacs with their dot files just like they like it. Same as it ever was.
It feels boring to write about agent coding so often. I try to keep it spaced out, even if it’s with modest vacation photos and ramblings. But, so much is changing and often interesting that it feels like missing out on the moment to develop one’s thoughts, in writing and publicly.
So, I regret to inform you, the discourse on AI will continue until the situation becomes less dynamic, fascinating, and rewarding to develop what I think of it by writing publicly.
RIP my vim muscle memory
I bet a lot of the grumbling about coding agents is really about losing out on the joy of moving quickly through code via a deeply learned text editor. I’m feeling it too, but I’m not so great at vim that my whole coding identity is wrapped up in it.
As goes the complaint about the decline of manual transmissions, even in sports cars, so goes the decline of emacs/vi/etc. “No connection”, “not enough focus”. Not wrong, but the toothpaste is out of the tube here.
That said, there’s a legitimate complaint to make here about software slowing down! We were almost getting to the point that raw latency was a virtue in software. And then we got uncanny, but not fast, intelligence. 🤷
It strikes me that my text editor is now, effectively, a quasi-natural language interface onto bulk and sometimes quite intelligent text operations. Well, sometimes entirely not what I want. But still, bulk and natural language.
Lately, I try Cursor every several weeks. Mostly to see if the grass is greener. But it’s not for me. Jetbrains-flavored tools are more holistic and more right for me. Regardless of the grass.
I think Cursor is an intriguing team and a good product. The best version of VS Code you could possibly use, the easiest possible onramp to coding with agents. Nonetheless, I want to be using Zed for everything. However, I think the writing is on the wall for pure, even principled, text editors. Typing in code is no longer the bottleneck. Memorizing keystroke optimizations is doubling down on a sunk cost.
Looking at the functionality of an integrated environment like RubyMine or PyCharm, I think the value of its git integration has gone up by a lot. This is essentially why I went back to PyCharm from Cursor. I can’t be bothered with simplistic, non-excellent git interfaces. (I wonder if Sublime Merge sales will start to exceed Sublime Text sales. 🤔)
VS Code, and Cursor by literal extension, seem like the keybindings were invented by 3 different people ten years apart. Which they probably were, literally. I can’t abide this.
The other thing that brought me back to PyCharm is browsing/discovering data alongside code alongside revision changes. It stands to reason that if writing code is less critical in the future, understanding the data will grow in importance. Even better, understanding your code and data in the same tool, with sensible affordances to jump between the two. This is why I’m betting on IDEs gaining ground.
Let’s not mention the part where now we’re paying for code, from our new-found editors, by the line. I’m surprised I’m not more salty about this.😬
The future of code is understanding the whole system, not constructing each part in exquisite detail. The IDE of the future probably looks more like code review alongside a debugger, than a code editor with version control and data management attached as sidecars. A bunch of Slacks and GitHubs mashed together.
Processes should serve outcomes, not the other way ‘round
In moments where process overwhelms a team’s ability to get stuff done, I’ve been fond of saying we’re not here to:
- Push Trello cards around.
- Guess if a coding task is more like a small t-shirt or a large burrito.
- Arrange our git commits in just the right order.
- Write long-lived documentation.
I’d suggest that, instead, we’re here to ship code.
Now that writing code isn’t the tricky part, I’m more convinced than ever we are both not here to operate a process and that activities like these are more important.
I never claimed my aphorisms were consistent or even logical!
Somehow, it was only within the last month that I stumbled upon a better saying. We’re not here to write code, operate processes, recite our statuses in meetings, etc.; we’re in the software development game to make outcomes. In particular, the outcomes you might expect of a software-technology company, if you’ll pardon some light hype-y jargon. Shipping code is a small part of that, but it’s in service of fixing problems, improving efficiency, solving new problems for customers, or putting an invention into the world.
I was going to include checklists in the list of things we’re not here to do. But, checklists are a pretty great way of thinking, and often lead to their intended outcomes. When they don’t lead to those outcomes, it’s typically easy to look at the items in the list and see why! “Oh, we skipped this safety check and things went sideways” or “we didn’t do the user research and, predictably, users weren’t sure what this feature is for.”
By contrast, documents, processes, and quality gates typically obscure your organization’s intended outcome. They’re too frequently a trade-off to prevent getting burnt in the same way as before. Which is probably why they correlate so frequently with failing to hit or losing track of the positive outcomes that matter.
“We’re here to learn” is a good proxy for “we’re here to generate outcomes”. If a checklist (or process or document, if you must) helps you learn more about customers/markets/your design/the world/etc. faster, it’s probably a good thing for you. If it prevents success and failure in equal probability density, it’s likely overhead.
The whole product team can use coding agents
I did a talk on coding agents, like Claude Code and Codex, advocating that everyone on product teams, not just coders, can use them to do more or better work.
Feel, fast, function, form
In that order, every time.
I make a big deal about working downhill. Get stuff done, slice it smaller, get feedback, go again, remove friction, a little faster this time.
But let me tell you, none of that matters if it doesn’t come together and feel great once it’s all assembled in front of you.
When you feel it, you know. The feature makes you smile when you use it. It fits right in, like it was always meant to be there. You want to use it again. You want to tell people about it.
This is the difference.
— Mitchell Hashimoto, You Have to Feel It
Everyone loves a little faster, even if they say they want more function or nitpick on the form. But not everyone asks for it to feel great. That’s what distinguishes the good from the great and the great from the sublime.
Feel, fast, function, form. In that order. Every time.
It has to feel great, feel quick, have the right functions, and pleasing forms. In that order of importance.
I’ll excuse anyone if they’re not a transcendent genius who can create things exhibiting those qualities in that order and every time they sit down to build. But when you can get all four of those things together, even in small quantities, then you’ve got something special.
Feels great, feels fast, gets the job done, all the right affordances and embellishments. In that order, every time. You gotta try, at least.
Finishing is a mindset
The last 10% of any creative act is the hard part. (Previously: Finishing is a skill.) You had an idea, thought it would take X days, only to find X-1 days of all these other things that have to be done. That’s functional scope creep.
Finishing, actually getting the draft or project out of your computer and out into the world, that’s a whole other list of things. Lots of work, and surprises, are lurking here. Call it delivery scope creep. Or a finishing tax. And, it’s a mental challenge, getting over the “they’re all gonna laugh at you” fear. (But not in the Adam Sandler way.)
The more steps you can remove from putting something out there, the more you can put out there. (There’s always money in the shovel stand.) Ergo, the continuous development of new systems and tools to make online publishing easier, despite the market having excited for nearly three decades.
Related, the programmers’ credo: “we do these things not because they are easy, but because we thought they were going to be easy.” So goes software, so goes any other creative endeavor.
Writing words and writing code feel similar, for me. (And, playing a musical instrument, if I go far back enough.) They are equal parts mechanical performance, exercise of taste, and act of creation. For all of them, I want to be in a flow state. I want to go as deep as I can, time permitting. I want to hold a whole world in my head. I want to work among as many details, as deeply as possible.
Finishing, whether it’s releasing software or publishing an essay, reviewing code, or editing and revising words, are a different skillset. Differently creative, but still putting something new into existence. There’s a nice symmetry there. If you can get good at editing and revising words, you can get good at editing and revising code.
Getting over the last 90% of any kind of project, whether code or words or music, is the same skill.
It’s like the last 90% of anything is just a mindset, almost a resilience of mind. If you can get good at one, the skill carries over into anything that requires finishing.
It feels like life pulls a fast one on us, at times. “What got us here, won’t get us there”. Making the software, essay, or music is different from shipping, publishing, or releasing it. But if you can get good at any one of shipping, publishing, or releasing, you are considerably further ahead on getting good at the other two.
Good enough to get going
The winning scenario for agent-assisted code, design, science, etc. is humans having more time to do creative and impactful thinking because computers/LLMs do the tedious setup, easily verified work, and gather preliminary materials that humans turn into inventions.
FWIW, I don’t think the worst scenarios are likely. The future isn’t atrophying literacy rates or people turning off their brains to tell LLMs what to do. It’s probably not Malthusian job scarcity or Keynesian leisure abundance, either.
The best outcome, IMO, is that producing almost-good-enough software, design, science, etc. is possible for more people, particularly those without specialist degrees.
You won’t have gym owners producing billion dollar SaaS companies, but they might produce software good enough to run their business without needing to contract out to a software developer.
You won’t have software developers producing the same level of design and art direction you see in major films. You might see them producing design good enough and sufficiently distinct that they can wait to bring a designer on until they’ve found their market.
You won’t have writers discovering new axioms of math and science, but you might see them correctly apply statistics and physics so that stories about finance and space battles are slightly more realistic.😉
In short: experts in topic A won’t find themselves held back by having an idea that requires expertise in topic A and topic B, where topic B is too deep for them to “just get good at”. Fewer Wozniaks will have to find their Steve Jobs, fewer Springsteens will have to find their Landaus.
It won’t exactly be you can just do stuff. But, perhaps you can get far enough along that collaborators to fill in the specialties you don’t can find you.