Brad Frost made an impact in UI design and engineering via his “Atomic Design” schema. The key insight is that if a user interface is made up of a hierarchy of elements, then a user interface should be deliberately styled at each stage in the hierarchy.
This principle also applies to software architecture. In fact, it applies to much, since all material things are compositions of form and matter. However, I will spare the lesson in metaphysics.
In software architecture, boundaries are drawn between systems. We distinguish between the software application and a CI/CD pipeline. We distinguish further between the frontend and backend. Further still, we distinguish between API endpoints. Etc.
Even as the methodology for creating software migrates to an agentic-first approach, the need for drawing boundaries remains.
Drawing boundaries is harder with LLMs.
The boundaries are inherently fuzzier. In traditional software, you implement X if you need X behavior. Even if you implemented X behavior via abstractions, there was a declarative implementation somewhere within your purview.
LLMs are behavior-constructors that don’t necessarily need your declarative instructions. LLMs dynamically evolve your thinking. This is what makes them so powerful; it is also what makes them so frustrating when they get things wrong.
Building software programmatically gives you predictability and observability, yet LLM behavior may remain a black box.
For this reason, programmatic agentic workflows have trended. You reason through an agentic system and its workflows instead of reasoning through software directly. You maintain something similar to software development, but it moves up to the agentic architectural level.
It makes sense, but it may introduce a new kind of bad abstraction. You risk creating an abstraction (e.g., agentic graph engineering) that mediates the LLM’s behavior yet accidentally weakens and/or duplicates the LLM behavior.
Thinking bottom-to-top, here are some questions that ary when architecting an agent-first software engineering methodology:
Do I leverage one model provide or leverage multiple? If I leverage multiple, do I pre-assign models by kinds of layers/kinds of tasks, or do I make “just-in-time” (JIT) determinations?
Similarly, how do I know which model and effort tier to utilize? Is this predetermined or a JIT calculation?
When do I reach for subscription-based agents versus per-use agents? Or, do I switch between the two dynamically?
What agent harness do I use, or do I create my own?
When do I create a plan before a task versus have the model dive right in?
How far out do I set the tasks and goals to be executed?
When do I preserve context internally (the agent’s memory) versus externally in documentation?
What is the optimal amount of context that I need for a given task?
Do I use a structured prompt or an unstructured one? Or, does it depend on the scenario?
How do I optimally reuse good-behavior-driving interactions with an LLM? Is it through a saving prompts, extracting skills, or creating shared packages that the LLM is pointed to? Or something else?
When do I leverage subagents, and how deep and wide should that subagent organization be? Should it be determined in advance or JIT?
All of these questions arise in an agentic methodology, and it exposes the paradox of AI engineering: it is incredibly effective when getting started, but it is very complex when aiming to optimize.
This complexity around optimizations coincides with other uncertainties:
How do I know how long a task will take?
How do I predict and justify spend?
How do I know when I even need to worry about optimization?
I won’t answer of these questions here, but I will propose an atomic design for reusing good-behavior-driving interactions with an LLM.
To begin, we need to call out two axes that determine the best fit.
One axis is how deterministic is the good-behavior-driving interaction. Prompts are the most generative, dynamic interaction; shared libraries are the most deterministic and predictable.
Prompts instruct to generate code; libraries instruct to utilize existing code. Prompts desires high variability (extension); libraries desires low variability (restriction).
Between prompts and libraries are skills and personalities.
Generative prompting can be extracted into skills. Skills do not extract code to leverage; skills extracts prompting workflows to structure the generative process.
As flagged above, as soon as we introduce abstractions to the right of prompting, we risk restricting, weakening, and/or duplicating the generative power of direct prompting.
This is where the second axis comes into play to help weigh the tradeoff: how reuseable the is knowledge.
Will the knowledge embedded in a prompt, for instance, be used again? Will it be used outside of the current task, project, and/or domain?
Evaluating along these two axes, we can make a prudential choice between a prompt and skill. If a prompting workflow is not highly variable and is used for many tasks, then a skill becomes a good choice.
Between skills and libraries are personalities. Personalities abstract judgement whereas skills abstract prompting workflows.
Personalities orchestrate which skills to leverage generate good behavior, and they specify what is needed for those skills to function well. For example, there might be a researcher “personality” that organizes the skills and knowledge needed to generate good researching outcomes.
Libraries differ from personalities in that they are completely detached from the generative. They effectively restrict an LLM not to generate the code/artifacts they contain. A downstream prompt (perhaps mediated by a skill and/or agent) instructs the LLM to merely provide the glue to that library.
When thought of this way, the atomic nature of agentic engineering methodology comes into focus.
Prompts —> Skills —> Personalities —> Libraries
When thought of with respect to the two aforementioned axes, we can construct these kinds of guidelines for a given thing/scenario:
| Thing/Scenario | Variability | Reuse | Choice |
| ------------------------------------- | ----------- | ------------- | ------------ |
| “Make this card less cramped” | High | Once | Prompt |
| Search retrieval failure analysis | Medium | Many tasks | Skill |
| Search-engineering judgment | Medium | Many projects | Disc |
| EPUB reader | None | Many projects | Library |
| Decide which search technique applies | High | Many projects | Personality |
| Calculate search metrics | None | Many projects | Library |Despite providing these guidelines, I suggest working through these complexities with empiricism rather than idealism. Bad abstractions have always been a risk. Given the blurred boundaries inherent in AI engineering, bad abstractions are easier to make and thornier. Start with prompting and evolve your methodology through experience before imposing ideas-for-abstraction before you have not yet felt the need for them.


