In a recent article, I tried to gain more precision on how the agentic shift to “judgment” carries an experientially significant reworking of a software engineer’s workflow, and how the common appeal to remaining “judgment” is too simplistic. Here, I’d like to double-down on one key aspect of this experiential shift: the lack of feedback in agentic development workflows.
Before agenetic development, there was a spectrum of ways that you could measure your success as a developer. The most immediate is task completion. You are given a task, and you write code that works to complete the task and its embedded requirements. How do know you’ve written code that works and that it meets the requirements? First, you can have a logical assurance. You can reason about the code because you wrote it and trust that it makes logical sense. Second, you can write tests, compare them against the requirements, and have a measure of assurance when the tests show green. Third, you (and perhaps others) can manually test the code. Arguably, the first kind of assurance could get obscured if you develop a mental habit of not reasoning through every angle of the AI-generated code. Assuming good habits, not too much may seem to change. But that’s on the surface level. It is on the higher-order level that the shifts emerge. You lose the innate assurance in pre-agentic workflows that you contributed something others could not produce.
Before the dawn of LLMs, you had a tacit confidence that you were doing something non-software engineers couldn’t do. You didn’t think about it. You didn’t question it. You just knew that you were writing in a programming language, untangling logical problems, and adding a dimension of quality on top of it all through not obvious refactoring and design patterns. A non-engineer would not be able to get past the first hurdle of understanding a programming language. As we absorb more and more non-engineers publishing code-driven apps, see LLMs produce something extensive from a simple prompt, and encounter how LLMs can explain concepts better than published blogs, does our value-edge weaken? This is the hard question, and it gets answered with “judgment,” but as I’ve argued before, LLMs do absorb “judgment” as well: they can review code, write code that is logically coherent, and refactor.
One might be tempted to say that their prompting is their value-edge. They have confidence that a non-engineer could not produce the quality of prompting as quickly as they did. However, not only is the path to learning prompting made easier to anyone, but the impact of your prompting can be overvalued. For example, I recently experimented with different ways of asking Claude to create a pure CSS pokeball. I used to spend hours at a coffee shop cranking out these kinds of pure CSS images, but Claude provided something quite impressive with a simple “Please write a pure CSS pokeball” prompt. When I made the end goal (i.e., final cause) explicit, it produced something more robust.
When I was a bit more vague in my writing about the end goal, I got something worse. However, it likely would have corrected itself with a few more prompts. So what did my awareness to be explicit in communication save me? Having to write a few more prompts (at most). In order to have a value-edge, I’d either need to have known to ask for a pure CSS pokemon (or other requirements) when no one asked, but this alreadys puts us beyond mere task completion into something categorical different—product engineering of sorts. Or, I’d need to know that the prompt-saving (due to my better communication) really made a difference at a certain level of scale/complexity. But such benchmarks are not obvious.
I did a similar experiment with asking Claude to produce poems with Lord of the Rings parallels (norse mythology, elvish, medieval, etc.).
What I discovered is that my most explicit prompt (Please write a poem. The essence of this poem is a elvish medieval kingdom inspired by norse myth that focuses on eucatastrophe. The rhyming scheme is accidental. The formal cause is to provide a poem that stir and uplifting sense of an adventure and final destination.) was not substantially better than my less technical prompt (Write an elvish medieval kingdom poem inspired by Norse mythology with an uplifting sense of arrival and adventure at the end, rhyme scheme doesn't matter).
“Eucatastrophe” is a specific introduced by Tolkien in his On Fairy Stories essay. My explicit mention of that may not have produced better results since the model is already “internalized” Tolkien et al’s material.
Of course, this is all subject to interpretation since these are loose tests comparing prompts with their outputs. However, so long as we do not have the interpretability models that can trace the “reasoning” of an LLM (as the LLM companies possess), then we don’t have clear knowledge as to whether our prompts actually sent an LLM in a better direction than if it simply inferred from its own “internalized knowledge.”
Bringing this back to the basic developer experience of task completion, the more obstruction of confidence in our technical knowledge as differentiating value, the more existentially destabilizing the apparatus becomes. At some point, will having only having the end product as the benchmark of our impact destabilize the whole developer experience? Or, will we remain happily naive in thinking that our prompting is uniquely impactful to generate the end product even when we don’t have an objective trace? Perhaps its existentially best to not raise these questions.





Regarding your poem prompt example, I have similarly discovered that technical precision and verbosity is of little value, especially as models improve. What is important are the key terms, which look about the same in your two prompts. LLMs fixate on terms and one malapropism could send them spinning in the wrong direction.
As a software engineer, I've always been involved in the project management aspects of the business, partly by necessity but also because I enjoy thinking about the big picture. Now that agents do most of the grunt work, I find myself more and more in the project manager role. That means tasks like your CSS pokeball example are no longer things I would consider doing myself. In an organization I would have less need of a design helper, and on my own I can accomplish more. It doesn't feel like a threat because I'm already entrenched with experience. But you're right to be concerned about humanity when there is little incentive to learn what good code and design is by working through CSS pokeballs, code katas, etc.