Anthropic recently discovered something illuminating, LLMs demonstrate planning prior to writing out individual words—shattering the notion that an LLM merely predicts and writes words one at a time:
Language models are trained to predict the next word, one word at a time. Given this, one might think the model would rely on pure improvisation. However, we find compelling evidence for a planning mechanism.
Specifically, the model often activates features corresponding to candidate end-of-next-line words prior to writing the line, and makes use of these features to decide how to compose the line.
Here, we find another instance where classical metaphysics provides the framework to understand these findings.
According to Aristotle, all action follows from an end (final cause). Even processes that appear purely mechanical or reactive are ordered toward some purpose (telos). A plant grows toward sunlight. A stone falls toward the earth. An artisan crafts toward a planned form.
When a user prompts an LLM to write a rhyming poem, the model must operate with this end in view. It cannot generate words purely sequentially—predicting each next word based only on what came before. Instead, it must coordinate its output toward the final form: a poem that both makes sense and satisfies the rhyme scheme.
Anthropic’s research reveals precisely this: the model activates representations of future rhyme words before generating the intervening text. It “looks ahead” to identify suitable end-of-line words, then constructs each line to arrive naturally at those words.
This is final causality in operation. The earlier words in each line exist for the sake of the planned end-word. The model coordinates means (word choices) toward an end (satisfying rhyme and sense constraints simultaneously).
Aristotle would recognize this pattern. When a sculptor carves a statue, he doesn’t chisel randomly and hope a form emerges. He works with the final form in mind—each strike of the chisel ordered toward the envisioned end. Similarly, the LLM doesn’t improvise randomly; it generates each word in light of where the line must arrive.
But there’s a crucial limitation: the model receives its end from the user’s prompt. It has awareness of the proximate end (produce a rhyming poem) but cannot reason to higher ends and causes—since it is not being a human—a rational animal. It doesn’t understand why rhyming matters, what makes a poem good rather than merely technically correct, or what purposes poetry serves in human life.
The model operates with functional teleology—coordinating means toward an externally-provided end—but not rational teleology—deliberating about which ends to pursue or understanding the hierarchy of goods.
This places LLMs in a fascinating middle ground between natural agents (which operate toward ends without any awareness) and rational agents (which grasp ends as ends and deliberate about them). The model has sufficient “knowledge” of its end to plan strategically, but insufficient rational grasp to understand why this end or to pursue self-determined purposes.
In short, Anthropic’s discovery of planning mechanisms in LLMs accidentally confirms what Aristotle taught: intelligent action requires final causality. Even artificial systems, when they achieve sophisticated coordination of means to ends, reveal the necessity of teleological structure.
Related Articles:
The Geometrical Universality of Thought
The Universal Language of Thought
An Aristotelian Introduction to Large Language Models




