In a previous article, I introduced a mental model to help grasp the sectors of work emerging in AI engineering. It may be helpful to read the article first, but the gist is that there are researchers, explorers, and stitchers working acrossing the AI engineering landscape. There are institution-backed researchers primarily investigating how to interpret and evolve model architecture. There is an explosion of free-floating explorers building prototypes, small products, and general tools and workflows. Finally, there are stitchers that are taking general tools and workflows to build a particular workflow for an engineering organization and/or team within a company. While these boundaries generally map to the pre-agentic software engineering landscape, researchers are in closer contact with the day-to-day engineering. Also, explorers and stitchers have more overlap, since many companies have reacted to hype and mystery by funding exploration efforts.
Surrounding all three kinds of work is AI-trajectory news and research, such as Anthropic’s economic index, Dario Amodei’s The Adolescence of Technology (and similar essays), stream of conscious insights/updates on social media and blogs, et cetera.
From this outlook, I’ve tried to see what gaps exist in the ecosystem and what kinds of work do these gaps fall under. Of course, research is foundational. Research is what has made the development of LLM architecture possible, and it gives shape to what is possible for agentic systems to accomplish. However, explorers are foundational in making research actionable. Meaning, they can turn research into usable tools and workflows. Since stitching is scoped to particular company and/or product, it is not as general as explorers and has the least impact on the ecosystem. It’s possible that explorer and stitching work is overlapping and general tools and workflows may emerge from company/product-specific contexts. However, I think it’s best to label this as exploration that occurs alongside stitching. This narrows down the highest-impact work (relative to the entire ecosystem) to research and exploration.
Nevertheless, there is something very valuable going in stiching work. That work is often occuring in contexts where the company has decided to provide serious investment for company/product-specific AI workflows to be developed. As teams are putting together are stitching together tools and workflows, the evaluation of the architecture is much more of an art than a science at this point. I’ve heard teams talk about having built up an architecture in the past few years during the hype race, and now needing to narrow things down and focus more on cost. In a word, stitching has surfaced that tools and workflows are not the only thing that is needed. Evaluation to inform the right architecture is now becoming a priority.
Here is where I see the highest value. It may begin with research, but it’s exploration that will make it actionable. The highest value is in exploratory work that provides guidelines for the ideal architecture that can be consumer by a stitching team. It is one thing for explorers to expose tools and workflows, and it’s another thing to expose guidelines on how and when to use those tools and workflows. Since stitching teams are the consumers, and are beginning to exit the hype-prototyping phase, evaluation is the highest value work at the moment.
It’s much harder than it sounds, however, as this is a different beast than software engineering workflows. With software engineering, it was easier to establish benchmarks to compare and contrast tools and workflows. As soon as we move to the realm of the agentic, we are ultimately producing software through a model. And interfacing with models still carries a black box, for we don’t really know how much of an impact our inputs to the model are making. How do we measure that this prompt produced a better result than if we wrote something a bit different? How can we be sure that our RAG layer was essential in the end result? So on and so forth.
The ultimate objective answer to these questions would require us to be able to see the reasoning of a model. However, this is a black box in the field of interpretability that is in a separate sector of research work. We can benchmark outputs but not reasoning at this stage. To benchmark reasoning, the companies behind the models will need to pair with researches to grasp how the models work—and some progress is being made here. However, companies will also need to provide a pane of glass into the reasoning after they wrap their heads around what exactly the reasoning process looks like.
I recently read Evals for AI Engineers, which is a helpful book for leveraging evaluation tools and workflows; and, I think this is the right place to start for explorers. General evaluation tools and workflows are needed. Moreover, there needs to be work done to examine the exact trade-offs around agentic architecture. There are various ways to leverage a model. If you’re writing software, for example, you could go to Claude and talk through your approach and then write the code yourself. Or, on the other end of the spectrum, you could orchestrate a robust agent organization via the LangChain ecosystem. Guidelines are needed to determine where on the spectrum a stitching team should land, and how to dial the knobs therein. When should a company introduce a RAG layer? When should a company use an orchestration architecture? And even if an orchestration architecture is decided upon, how should the agent organization be structured? These are the sorts of answers that didn’t need as much precision in the hype-prototyping phase but will now need to be explored. The place to start, in my opinion, is enumerating all the potential workflows and the essential tradeoffs that will guide the ideal choice for a particular stitching team.

There is one final warning I must give having said all this. The agentic ecosystem lives and breathes in thinking computationally. Even in the philosophical research, a computational anthropology is often assumed when the reality may be more complex. By “thinking computationally,” I mean that we tend to wait for things to emerge through empirical tests before we commit with confidence. However, if interpretability largely remains a black box (especially to explorers and stitchers), then we risk staying in vibe architecture. Moving from vibe coding to vibe architecture only shifts the problem. But there is an alternative option, and I believe it’s the only viable option at this point in time where model reasoning is a black box. We can look for analogies of agentic organizations—other examples when tradeoffs needed to be considered to form the ideal organization. We cannot find this in the agentic discourse, for this is all new (of course). However, this isn’t the first time humans have had to evaluate how to delegate tasks through a hierarchical structure. Scipio Africanus and Napoleon would say otherwise. This is where things get interesting and impactful. But I will save that for an upcoming post.


