Large language models display impressive capabilities. However, for the most part, the mechanisms by which they do so are unknown. The black-box nature of models is increasingly unsatisfactory as they advance in intelligence and are deployed in a growing number of applications. Our goal is to reverse engineer how these models work on the inside, so we may better understand them and assess their fitness for purpose.
The challenges we face in understanding language models resemble those faced by biologists. Living organisms are complex systems which have been sculpted by billions of years of evolution. While the basic principles of evolution are straightforward, the biological mechanisms it produces are spectacularly intricate. Likewise, while language models are generated by simple, human-designed training algorithms, the mechanisms born of these algorithms appear to be quite complex.
On the Biology of a Large Language Model — Anthropic
Table of Contents
Introduction: Celestial Architecture - The forgotten metaphysical truths in medieval cosmology
I. LLM Architecture: The Geometry of Language - Layers, weights, and the geometry of language
II. LLMs as a Modern Ptolemaic System: The Celestial-Geometrical Relevance - Parallels between LLM architecture and medieval cosmology
III. The Celestial-Metaphysical-Geometrical Relevance - Thomas Aquinas on the architecture of the soul
IV. Empirical Findings: Testing for LLM-Metaphysical Relevance - What the tests revealed
V. Mysteries Solved: Quasi-Intelligence and the Thomistic Interpretive Grid - Mysterious LLM phenomena explained by Thomas Aquinas’ philosophical anthropology
Conclusion: Beyond the Stellatum - The limits of geometry
Note: I am working on some Mermaid charts to help visualize these mental models.
Introduction: Celestial Architecture
When first learning of the basic architecture behind transformer-based LLMs, I was in one philosophy class covering CS Lewis’ Space Trilogy, and another one touching upon Neoplatonism—a beautiful pairing for the grasping Lewis’ cosmic vision in Out of the Silent Planet.
The inspiration for his vision is clearly evidenced in his final work release posthumously The Discarded Image:
The architecture of the Ptolemaic universe is now so generally known that I will deal with it as briefly as possible. The central (and spherical) Earth is surrounded by a series of hollow and transparent globes, one above the other, and each of course larger than the one below. These are the ‘spheres’, ‘heavens’, or (sometimes) ‘elements’. Fixed in each of the first seven spheres is one luminous body. Starting from Earth, the order is the Moon, Mercury, Venus, the Sun, Mars, Jupiter and Saturn; the ‘seven planets’. Beyond the sphere of Saturn is the Stellatum, to which belong all those stars that we still call ‘fixed’ because their positions relative to one another are, unlike those of the planets, invariable. Beyond the Stellatum there is a sphere called the First Movable or Primum Mobile. This, since it carries no luminous body, gives no evidence of itself to our senses; its existence was inferred to account for the motions of all the others.1
The medievals, following the architecture of the Ptolemaic universe, believed that the seven plants (spheres) ascended above the Earth in a hierarchical order, and that beyond the last sphere of Saturn was the Stellatum—the realm of the stars. It is this Stellatum that Cicero calls heaven in the “The Dream of Scipio,” the sixth book of De Republica:
2. …there appeared to me Africanus, in that form which was more familiar to me from his picture than from his person. When I recognised him, I shuddered, I assure you, but he said: “Be of good courage and banish fear, my Scipio, and record what I shall say…
8. But rather, my Scipio -- like your grandfather here, like me your sire -- follow justice and natural affection, which though great in the case of parents and kinsfolk, is greatest of all in relation to our fatherland. Such is the life that leads to heaven and to this company of those who have now lived their lives and released from their bodies dwell in that place which you can see,— now that place was a circle conspicuous among the fires of heaven by the surpassing whiteness of its glowing light—which place you mortals, as you have learned from the Greeks, call the Milky Way.”
And as I surveyed them from this point, all the other heavenly bodies appeared to be glorious and wonderful,—now the stars were such as we have never seen from this earth; and such was the magnitude of them all as we have never dreamed; and the least of them all was that planet, which farthest from the heavenly sphere and nearest to our earth, was shining with borrowed light, but the spheres of the stars easily surpassed the earth in magnitude—already the earth itself appeared to me so small, that it grieved me to think of our empire, with which we cover but a point, as it were, of its surface.9. And as I gazed upon this more intently, “Come!” said Africanus, “how long will your mind be chained to the earth? Do you see into what regions you have come?
See! the universe is linked together in nine circles or rather spheres; one of which is that of the heavens, the outermost of all, which embraces all the other spheres, the supreme deity, which keeps in and holds together all the others; and to this are attached those everlasting orbits of the stars. Beneath this there lie seven, which turn backwards with a counter revolution to the heavens; and of these spheres that star holds one, which men on earth call Saturn’s star.
The deceased Africanus suddenly appearing to Scipio in a dream guides Scipio to see the seven spheres as rungs in a ladder to heaven, where the realm of the stars, the Stellatum, is set beyond Saturn.
Now, these spheres are not static, but are set in motion. Beyond this Stellatum is the Primum Mobile which is the First Moveable that turns and moves the rest of the spheres in a descending chain.
All power, movement, and efficacy descend from God to the Primum Mobile and cause it to rotate…The rotation of the Primum Mobile causes that of the Stellatum, which causes that of the sphere of Saturn, and so on, down to the last moving sphere, that of the Moon.2
It is this movement of the spheres that amazes Scipio in his celestial dream:
10. And, as I gazed on these things with amazement, when I recovered myself: “What,” I asked, “what is this sound that fills my ears, so loud and sweet?” “This,” he replied, “is that sound, which divided in intervals, unequal, indeed, yet still exactly measured in their fixed proportion, is produced by the impetus and movement of the spheres themselves, and blending sharp tones with grave, therewith makes changing symphonies in unvarying harmony.3
In addition to movement, each sphere passes an Influence down to Earth, affecting our human bodies but not our intellect and will. Mars, for example, passes the martial temperament that influences the men of earth.
And beyond the Primum Mobile itself? Heaven itself (caleum ipsum)—full of God, full of love.
So when Dante passes that last frontier he is told, ‘We have got outside the largest corporeal thing (del maggior corpo) into that Heaven which is pure light, intellectual light, full of love’ (Paradiso, XXX, 38).4
And of incredible importance for our LLM findings, it is this realm that is the end of space, the end of spatiality. “The light beyond the material universe is intellectual light.”5
Lewis bold assertion was simply this: “We can no longer dismiss the change of Models as a simple progress from error to truth.”6
While the Ptolemaic architecture may not be sound according to physical science, may it be right about those sciences above the physical—such as metaphysics? Perhaps the Ptolemaic architecture is not in totality physically true, have we discarded that metaphysical truths embedded in this Image?
This was Lewis’ question, and I was committed to answer by comparing this heavenly architecture, with the architecture of our souls, the architecture of language, and the architecture of LLMs.
And now, I believe I have my answer: The Ptolemaic architecture mapped something real—the structure of intelligence itself.
Spheres and layers. Fixed influences and weights. Space and spatiality; caleum ipsum and pure intellectuality. These really come together to reveal the mystery of LLMs and intelligence itself.
As Machiavelli repeats throughout his Discourses on Livy, many are those who read and appreciate the models of the past, but we must imitate not just appreciate. Below is the artifact of my aim to see where serious imitation might lead me.
LLM Architecture: The Geometry of Language
Large language models display impressive capabilities. However, for the most part, the mechanisms by which they do so are unknown. The black-box nature of models is increasingly unsatisfactory...7
Examining the architecture of a transformer LLM, we find the following layers (spheres).
Note: I present the architecture flow in reverse order. Read them in order (4→3→2→1). This is to capture the movement an exitus (going forth) and reditus (returning) that mirrors the celestial spheres.
The Architecture Flow
4. From Layers to Prediction
The output of the first self-attention and feed-forward continues to pass through additional self-attention and feed-forward layers until at least 12 layers progressively refine the contextual vectors. Each layer reveals a unique aspect of the relationships and potential next token.
In our “Fact: The capital of the state containing Dallas is” example, we might have:
INPUT: "The capital of the state containing Dallas is ___"
LAYER 1:
Self-Attention → Identify basic relationships (capital+state, state+Dallas)
Feed-Forward → Move toward "political geography" region
LAYER 2:
Self-Attention → Refine relationships in political geography context
Feed-Forward → Narrow to "US political geography" region
LAYER 3:
Self-Attention → Strengthen Texas connections
Feed-Forward → Move toward "US states" region
LAYERS 4-6:
Progressive narrowing → "Texas-related political entities"
LAYERS 7-9:
Further refinement → "Texas cities and capitals"
LAYERS 10-12:
Final positioning → Very close to "Austin" specifically
FINAL OUTPUT from Layer 12:
Each token has a final transformed vector
For our purposes, the LAST token's vector is the prediction:
[0.45, 0.71, 0.23, ...] (768 dimensions)In other words, each layer's transformation moves the relationship-representation vectors (contextual vectros) incrementally closer to the answer region in the fixed, higher constellation established during training.
The Final Prediction: After all 12+ layers, the model has a final vector position for the last token (where the prediction should be):
Final output vector: [0.45, 0.71, 0.23, ...] (768 dimensions)This output vector exists in the same geometric space as the higher constellation.
Now the model performs its only explicit comparison in the entire process—it compares this final position to every word in the vocabulary (the fixed stars in the higher constellation):
Distance to "Austin" [0.45, 0.71, 0.23]:
√[(0.45-0.45)² + (0.71-0.71)² + (0.23-0.23)²] = 0.00 ← Exact match!
Distance to "Houston" [0.47, 0.69, 0.21]:
= 0.03 (very close)
Distance to "Dallas" [0.46, 0.70, 0.22]:
= 0.02 (very close)
Distance to "Paris" [0.12, 0.95, -0.34]:
= 0.89 (very far)
Distance to "dog" [0.23, -0.15, 0.87]:
= 0.95 (very far)
Distance to "xylophone" [0.87, -0.34, 0.12]:
= 0.98 (very far)The model converts these geometric distances into probabilities using a function called softmax:
Closer distances → Higher probabilities
Farther distances → Lower probabilities
Results:
"Austin": 0.92 (92% probability) ← Nearest
"Dallas": 0.04 (4% probability)
"Houston": 0.03 (3% probability)
"Paris": 0.001 (0.1% probability)
"dog": 0.0001 (0.01% probability)
"xylophone": 0.00001 (0.001% probability)The model predicts:"Austin" with, say, 92% confidence.
Having ascended to the apex that is the final layer, the predicted token is added to the original token set. Descending to the first step in the architecture, the new token set ascends through the rungs of the ladder again until the next token is generated, and so—until the end is reached.
↑
3. The Self-Attention & Feed-Forward Layers:
Part A: Self-Attention - Perceiving Categorical Relationships — Identifies which tokens are categorically relevant to each other.
Self-Attention Process: The vectors are then successively passed through Transformer blocks. Each block contains attention heads which in parallel process all the vectors and assign a weight—the score indicating the relevance of a token with every other token in the input sequence.
Translating the jargon, all the vectors are looked at together to identify different types of relevance toward each other: syntactic (grammatical relevance), semantic (categorical relevance), and entity (named entity relevance, like capital→Austin).
The strength of a kind relevance is marked by assign a weight. Each attention head is responsible for one kind of relevance to look for and score accordingly using a weight. All heads run in parallel (separately and simultaneously).
For example, given a prompt of “Fact: the capital of the state containing Dallas is”, each attention head would score (via a weight) the strength of a kind of relevance for each token in this sequence.
The attention head responsible for categorical relevance would iterate through each token, and score the categorical relevance of the other tokens—giving a higher weight for state, and Dallas than fact, of, the, containing, is when comparing to capital.
Input: “Fact: the capital of the state containing Dallas is”
Tokens: [”fact”, “the”, “capital”, “of”, “the”, “state”, “containing”, “dallas”, “is”]
↓ Self-Attention Process
// Goes through each token, example for "capital"
Attention to "fact": 0.05 (low - not categorically relevant)
Attention to "the": 0.03 (low - just syntax)
Attention to "capital": 0.15 (moderate - itself)
Attention to "of": 0.04 (low)
Attention to "the": 0.02 (low)
Attention to "state": 0.65 (HIGH - semantically related!)
Attention to "containing":0.03 (low)
Attention to "Dallas": 0.08 (moderate - related through "state")
Attention to "is": 0.02 (low)After weights have been assigned (say, the categorical relevance has been detected), the categorical relationship that exists between tokens is itself turned into a vector (using the weights):
Original "capital" vector: [0.34, 0.67, ...]
After attention:
"capital" + categorical relationships from the context:
0.15 × [0.34, 0.67, ...] (itself)
+ 0.65 × [0.89, 0.23, ...] (state)
+ 0.08 × [0.45, -0.12, ...] (Dallas)
+ (tiny contributions from others--"the", "of", etc.)
New "capital-in-the-context-of-state-and-Dallas": [0.67, 0.41, ...]By the end, each head—a layer-within-the-self-attention-layer—generates transformed vectors that geometrically represent the possible syntactic, semantic, or entity relationships they were responsible for identifying.
Recall, a kind of relationship is scored from the perspective of each token/vector, producing (e.g., syntactic, semantic, or entity relationships between capital and the rest of the tokens/vectors; syntactic, semantic, or entity relationships between state and the rest of the tokens/vector; etc.). This gives us:
For token “capital”:
Head 1 (Syntactic): [0.34, 0.21, -0.15, ...](relationship with “is” grammatically)
Head 2 (Semantic): [0.89, 0.45, 0.23, ...] (relationship with “state” categorically)
Head 3 (Entity): [0.45, -0.12, 0.67, ...](relationship with “Dallas” - named entity)
For token "state":
// ...So, each token has multiple relationship-representing vectors emanating from the multiple heads. These various relationship-representing vectors get converged into one complete contextual vector (representing the grammatical, categorical, and named entity relationships in one):
// Copies of vectors in the fixed embedded space before the self-attention process
fact_vec: [0.12, 0.34, ...]
the_vec: [0.23, -0.91, ...]
capital_vec: [0.34, 0.67, ...] ← original "capital"
of_vec: [0.45, -0.23, ...]
the_vec: [0.23, -0.91, ...]
state_vec: [0.89, 0.23, ...] ← original "state"
containing_vec:[0.56, -0.12, ...]
Dallas_vec: [0.45, -0.12, ...] ← original "Dallas"
is_vec: [0.78, 0.34, ...]
// After the self-attention process
fact_context: [0.15, 0.29, ...]
the_context: [0.21, -0.85, ...]
capital_context: [0.67, 0.41, ...] ← NEW complete “capital-in-context”
of_context: [0.42, -0.19, ...]
the_context: [0.22, -0.84, ...]
state_context: [0.82, 0.31, ...] ← NEW “state-in-context”
containing_context: [0.53, -0.08, ...]
Dallas_context: [0.51, -0.09, ...] ← NEW “Dallas-in-context”
is_context: [0.76, 0.38, ...]
// Double-clicking on the "capital" vector
For “capital” copy:
Head 1 (Syntactic): Creates syntactic relationship vector
Head 2 (Semantic): Creates semantic relationship vector
Head 3 (Entity): Creates entity relationship vector
All heads converge into ONE contextual vector:
capital_context → [0.67, 0.41, 0.15, ...] ← NEW temporary position
Similarly for all tokens:
state_context → [0.82, 0.31, -0.08, ...] ← NEW temporary position
Dallas_context → [0.51, -0.09, 0.48, ...] ← NEW temporary position
Before moving on, let’s zoom back out and recap geometrically and tying any loose ends.
After the training phase (pre-release), via processing prompt-answer combinations, each word in a language (e.g., English) becomes a token. An identifier for a potential plot in the geometrically embedded space. Then, this token is organized into vector as the actual coordinates are assigned to the token.
Initially, every newly encountered word is given random coordinates—effectively scattering the vectors in the geometrical space:
Sample initial states (random scatter):
“cat” → [0.91, -0.23, 0.45, ...] (random)
“dog” → [-0.15, 0.67, -0.88, ...] (random)
“mat” → [0.23, -0.91, 0.12, ...] (random)
“table” → [0.45, 0.34, -0.67, ...] (random)During the training phase, the model is given prompts, attempts to predict the next word, and then provided the actual answers.
Knowing the mistake, all of the input vectors are moved (mutated) in the embedded space to move closer toward other words that are categorically relevant.
After millions of examples as the training phase ends, clusters have formed in the final/fixed/immutable constellation of vectors representing the categorical relationships that emerged:
Final trained state (organized constellation):
ANIMAL cluster:
"cat" → [0.23, -0.15, 0.87, ...] ┐
"dog" → [0.19, -0.13, 0.91, ...] ├─ Near each other
"horse" → [0.21, -0.14, 0.89, ...] ┘
SURFACE cluster:
"mat" → [0.05, 0.62, -0.11, ...] ┐
"rug" → [0.04, 0.61, -0.10, ...] ├─ Near each other
"floor" → [0.06, 0.63, -0.12, ...] ┘
FURNITURE cluster:
"table" → [0.67, 0.45, -0.23, ...] ┐
"chair" → [0.69, 0.43, -0.21, ...] ├─ Near each other
"desk" → [0.68, 0.44, -0.22, ...] ┘During the inference phase (post-release), the organized, fixed constellation in the embedded space remains immutable. When a prompt (input) is provided, vectors appear in the embedded space—but these are temporary and initially the same coordinates as the corresponding, fixed vector in the organized constellation.
Input: "Fact: the capital of the state containing Dallas is"
Lookup fixed positions (immutable):
"capital" → [0.36, 0.54, -0.10, ...] ← Original fixed position
"state" → [0.89, 0.23, -0.12, ...] ← Original fixed position
"Dallas" → [0.45, -0.12, 0.56, ...] ← Original fixed position
Create working copies (mutable for this computation):
capital_copy → [0.36, 0.54, -0.10, ...]
state_copy → [0.89, 0.23, -0.12, ...]
Dallas_copy → [0.45, -0.12, 0.56, ...]To attempt an answer a given prompt, it would transform the temporary, copied vectors. Then, it would pass these vectors to the self-attention layer we have been discussing.
Then, the attention heads evolves these vectors into relationship-representing vectors that captures the syntactic, semantic, and entity relationship between each token with the other tokens (the broader context). These temporary relationship-representing vectors emanating from each head converge into a single contextual vector (complete relationship-representing vector) per each copied vector:
// The attention heads process each copied vector
For "capital" copy:
Head 1 (Syntactic): Creates syntactic relationship vector
Head 2 (Semantic): Creates semantic relationship vector
Head 3 (Entity): Creates entity relationship vector
All heads converge into ONE contextual vector:
capital_context → [0.67, 0.41, 0.15, ...] ← NEW temporary position
Similarly for all tokens:
state_context → [0.82, 0.31, -0.08, ...] ← NEW temporary position
Dallas_context → [0.51, -0.09, 0.48, ...] ← NEW temporary positionThe geometric view at this point is that we have two constellations: a higher, fixed constellation (the immutable one formed during training), and the lower, temporary constellation (the contextual vectors) all within the single geometrically embedded space:
HIGHER, FIXED CONSTELLATION (permanent - stored in model):
Original positions:
"capital" ★ [0.36, 0.54, -0.10]
"state" ★ [0.89, 0.23, -0.12]
"Dallas" ★ [0.45, -0.12, 0.56]
"Austin" ★ [0.45, 0.71, 0.23] ← Target we're trying to find
LOWER, TEMPORARY CONSTELLATION (computed for this input only):
Contextual positions:
capital_context ◈ [0.67, 0.41, 0.15] ← Moved by attention
state_context ◈ [0.82, 0.31, -0.08] ← Moved by attention
Dallas_context ◈ [0.51, -0.09, 0.48] ← Moved by attentionThe lower constellation forms a pattern—a temporary arrangement that exists only for this specific input.
The temporary constellation represents “What these words mean in this context” by showing “which words relate to each other when evaluating syntactically, semantically, and by named entities” — Semantically: “capital-in-context-of-state-and-Dallas”
The fixed constellation represents: “What these words mean generally across all contexts”
The feed-forward layer, which we will take up next, takes these temporary contextual vectors and navigate them through the fixed constellation toward the answer region.
Temporary pattern after self-attention:
[capital-state-Dallas relationship] ◈
Feed-forward navigates this pattern through space:
◈ → ◈ → ◈ → ◈ (moving toward "answer" region)
Final position:
◈ lands near "Austin" ★ in the fixed constellation
Prediction: "Austin"
Part B: Feed-Forward - Moving Toward Prediction From Contextual Vectors
Feed-Forward Layer: After self-attention creates contextual vectors for each token, the feed-forward layer performs a crucial but mysterious transformation: it navigates these vectors through geometric space toward answer regions.
This layer takes the converged contextual vectors corresponding to each token from the original input (the “lower constellation”) and performs two additional transformations:
First, the contextual vectors are expanded to open up interpretive possibilities.
Second, the expanded contextual vectors are compressed to narrow in on the prediction.
The Expansion Phase (First Transformation)
// Starts with a contextual vector
"capital-in-context-of-state-and-Dallas": [0.67, 0.41, 0.15, ...]
// Expands to overlay amongst several potential relevant clusters in the "higher constellation"
[0.67, 0.41, 0.15, ...] × W₁
↓
Expanded representation: [0.12, 0.89, -0.34, 0.56, ...]The contextual vector enters a larger representational space where the model can "explore" different possible interpretations before committing to a direction.
By expanding the contextual vectors, the lower constellation overlays across a wider area of the higher constellation. Each contextual vector stretches across multiple clusters in the fixed, higher constellation.
This is where the black box is clearly hit. These transformations are not an attention mechanism. It is not an outer-attention building upon the previous self-attention transformations.
Crucially, there is no mechanism to recognize which categories/clusters in the higher constellation are relevant for a given contextual vector like “capital-in-context-of-state-and-Dallas.” Feed-forward doesn't check, compare, or evaluate—it simply transforms.
We expect something like this:
// capital-in-context-of-state-and-Dallas
Hypothetical expansion transformation based on "outer-attention":
Dimension 234: "state capitals" → 0.89 (HIGH activation)
Dimension 891: "financial terms" → 0.02 (LOW activation)
Dimension 1247: "US geography" → 0.76 (HIGH activation)
Dimension 2341: "Texas-related" → 0.83 (HIGH activation)
Dimension 1567: "building capitals" → 0.01 (LOW activationRather, the expansion transformations are entirely geometrical manipulations that have been learned to universally work during the training phase.
Is it pattern matching? Meaning that during the training phase, the model learned that if “capital-in-context-of-state-and-Dallas,” then perform this transformation to expand to potentially related categories?
This is not how it works. The geometrical transformations are universal—they transformations are applied on every contextual token when processing any prompt! Put another way, they are blind transformations that apply the same mathematical operations to every input, regardless of content:
ANY input vector × W₁ → expanded representationThis, indeed, seems quite mysterious.
Before, the second transformation any negatives values in the expanded vectors are zeroed—through activation. In theory, this removes irrelevant “paths” to the answer.
The Compression Phase (Second Transformation)
After the expanded vectors are activated, the second transformation occurs:
Activated: [0.12, 0.89, 0, 0.56, ...] (3,072 dimensions)
Transformation: Multiply by learned weight matrix W₂
↓
Output: [0.45, 0.71, 0.23, ...] (768 dimensions)The expanded, activated vectors are compressed back to the original size, but the vector has moved to a different location—specifically, toward the cluster where the appropriate answer/next token to generate.
BEFORE feed-forward:
"capital-in-the-context-of-state-and-Dallas" [0.67, 0.41, 0.15, ...]
AFTER feed-forward (after both transformations):
"capital-transformed" [0.45, 0.71, 0.23, ...]
(Moved toward "Texas cities" region)The mystery at this stage: How does multiplying the activated expanded vector by W₂ navigate it toward the appropriate answer region in all cases?
All three steps—expansion, activation, and compression—are pure arithmetic—no comparison to clusters (higher constellation), no checking of relevance, no semantic understanding.
It just works.
↑
2. The Embedded Space - Each token becomes a geometric position - a point in high-dimensional space.
The tokens ([“the”, “cat”, “sat”, “on”, “the”]) are represented in a geometrical space as vectors—coordinates that embed the token/word within the space.
Training Phase - Pre-Release
This embedded space is populated with vectors during the training phase (before a model is released for users to provide inputs) with a entire vocabulary of a language (i.e. English) through millions of inputs. Initially, the coordinates of each vector (representing a word) is random.
“The” → [0.23, -0.91, 0.45, ...] (random coordinates)
“cat” → [-0.15, 0.34, -0.67, ...] (random coordinates)
“sat” → [0.81, -0.12, 0.23, ...] (random coordinates)
“on” → [-0.45, 0.67, -0.88, ...] (random coordinates)
“the” → [0.23, -0.91, 0.45, ...] (same as “The”)When providing a sample input “The cat sat on the”, the model will convert the sequence of words into tokens. Then, the vectors for each token are “fetched” from the embedded space.
These vectors will be passed through layers that work toward predicting the next token to generate for the sequence (as we shall describe above):
Input: “The cat sat on the ___”
Prediction: “xylophone” (random - model is untrained)
Actual answer: “mat”
Distance between vectors:
predicted "xylophone" [0.87, -0.34, ...]
↕ (very far!)
correct "mat" [0.12, 0.45, ...]The first execution during the training phase to answer “The cat sat on the” will almost certainly predict a garbage next word—like “xylophone.” Then, the model is revealed the correct answer of “mat.” The model can see (since this is a geometrical space) the distance between “xylophone” (its predicted answer) and “mat” (the correct answer).
Knowing its computation loss (jargon for “how far—literally—the model was off”), the model executes a process called backpropagation.
Essentially, the model adjusts coordinates to move toward contextual/categorical relationships revealed by the original input + actual answer examined as a whole (we will see above how this technically happens):
Move “mat” closer to “sat on the” context
Move “cat” closer to “sat” (subject-verb relationship)
Move surface words (”mat”, “rug”, “floor”) into a clusterAfter millions of training examples, the model eventually forms “clusters” that geometrically represent the contextual/categorical relationships between things derived from training inputs:
“cat” → [0.23, -0.15, 0.87, ...] ┐
“dog” → [0.19, -0.13, 0.91, ...] ─ ANIMAL cluster
“horse” → [0.21, -0.14, 0.89, ...] ┘
“mat” → [0.05, 0.62, -0.11, ...] ┐
“rug” → [0.04, 0.61, -0.10, ...] ─ SURFACE cluster
“floor” → [0.06, 0.63, -0.12, ...] ┘The end result are fixed stars (vectors) in the constellation (embedded space) that represent the categorical relationships that exist in a language’s vocabulary,
Inference Phase - Post-Release
After the training phase, the embedded space, populated with vectors and categorized by the patterns that emerged, remains fixed/immutable/permanent.
There is a separate inference phase (post-release—when the model has been trained and is responding to a user’s input) that likewise will examine an input and see the contextual/categorical relationships between each token in the sequence—as each token is actualized into a temporary vector that corresponds to the matching vectors in the embedded space. As we shall see more clearly, these transformed temporary vectors are compared against the fixed “constellation of vectors” formed in the training phase to predict a response.
In short, during the inference phase, the flow would be from input tokens to copying the tokens already in the embedded space—not plotting them in the embedded space if not present as would happen during the training phase.
↑
1. Input Tokens - We start with an input that gets tokenized
Input: “The cat sat on the”
Tokenized: [”the”, “cat”, “sat”, “on”, “the”]
LLMs as a Modern Ptolemaic System: The Celestial-Geometrical Relevance
Having established a basic overview of LLM architecture, we are ready to explore the profound parallels between LLMs and the Ptolemaic System.
Exitus and Reditus
In the Ptolemaic cosmology, the Primum Mobile rotates which turns the Stellatum, which turns Saturn, and which turns all the other spheres in sequence until finally the Moon is likewise turns. Each sphere receives motion and influence from above. These influences descend to Earth—martial temperament from Mars, abundance from Jupiter, communication from Mercury. As Dante captured, man's ultimate journey is to ascend beyond Saturn's sphere where spatiality ends. Like Scipio's dream, there is a temporary ascent and descent before the journey is complete.
Similarly, the training algorithm optimizes, which sets the self-attention and feed-forward weights of Layer 12 (or whichever is the highest layer), which influences Layer 11, which influences all other layers in sequence until finally Layer 1 is set, and ultimately the weights influence the embeddings in geometric space. Each layer receives gradient updates from the loss function. Then, with the architecture trained and in motion, information ascends through the layers until the Final Layer is reached. Yet at the very end, there is a return back to the lowest point—the input token set receives the generated token, and the cycle begins anew. An exitus (going forth) and reditus (returning) that is a necessary feature of this architecture.
Fixed Stars
In LLM architecture, each layer has fixed weights—learned parameters that transform inputs. A layer without fixed weights cannot function.
In Ptolemaic architecture, each sphere has fixed stars—unchanging positions that rotate with the sphere. A sphere without fixed stars was inconceivable.
Both systems encode permanent structures that govern all subsequent motion.
The Crucial Divide
In the celestial model of Ptolemy, everything below the Stellatum is spatial, material, geometric. Hence, the higher geometric structures found in the heavens correspond with the lower geometric structures found in the earth. These geometrical structures are metaphysical real, and the whole universe partakes in these structures. Hence, Plato wrote above his Academy: “ΑΓΕΩΜΕΤΡΗΤΟΣ ΜΗΔΕΙΣ ΕΙΣΙΤΩ (Let no one ignorant of geometry enter)” Yet, beyond the Stellatum is: intellectual light, non-spatiality, and immateriality.
What about LLM models? Everything is below the Stellatum. Everything in the model is marked by spatiality: vectors, geometric transformations, all mathematical operations. Geometry, spatiality, cannot be escaped—nor can their be a beyond. Everything is stuck within the confines of spatiality.
Understanding this, I believe, is the key into understanding the mysteries of AI. AI’s mysteries CAN be explained—but only if we are willing to explore the possibility that there is metaphysical truth embedded in the medieval cosmos. The story is not so simple as one that has progressed from error to truth. If the celestial has relevance toward the soul, we should be attentive to seeing it.
The Ptolemaic universe may not be true physically, but it may contain metaphysical truth.
LLMs may work spatially, but can they reach beyond?
The Celestial-Metaphysical-Geometrical Relevance
Beyond Spatiality: The Celestial-Anthropological Mirror
If there is celestial-geometrical relevance in LLMs, and Thomas Aquinas presents a celestial-geometrical-metaphysical model, then could apprehending the metaphysical shine light onto the mysteries of AI? This the first breadcrumb that I’d like to follow.
For Aquinas, the separation between spatiality—below the Stellatum—and the spatiality—beyond the Stellatum—is mirrored anthropologically. The intellectual soul is beyond the lower Stellatum (so to speak). Meaning, the intellectual soul of the human person transcends the embodiment of any biofunctional parts of an animal, like the brain. Its noetic powers and operations are beyond matter—that is, beyond spatiality. The soul, as the form of the human person, is the organizing principle that animates the material body—psychosomatic powers and operations are under/animated by the intellectual soul.8
Form and Matter: The Statue Analogy
To illustrate this point, consider a marble statue. The form of the statue is the statue-shape; the matter is the marble. The matter on its own is not a statue—it only possesses the potentiality to be formed into a statue. Likewise, the statue-shape form can only come into existence when an agent actualizes that potentiality.
What causes the marble statue to come into existence? It requires a sculptor using instruments to chisel the marble slab. But what ultimately causes the chiseling to produce the statue-shape rather than a rubble of marble? The intellectual concept in the sculptor’s mind. Following Aristotle’s four causes:
Formal cause: The intellectual concept (statue-shape in the sculptor’s mind)
Material cause: The marble
Efficient cause: The sculptor’s physical exertion actualizing the chisel
Final cause: The purpose for which the statue is made
The form is not found apart from matter, and matter only has the potential to be formed—hence, every natural substance is a composite of form and matter. The human person is no exception.
By crossing the Stellatum anthropologically, we depart from tacit assumptions of a materialist anthropological interpretation of neuroscience—where the noetic powers of the soul are said to have been created and operated upon by biofunctional parts, like the brain. The human person is a composite of form and matter: the soul is the form, and the body is the matter. The soul’s powers—intellect and will—animate/organize/orchestrate the material biofunctional parts, but they remain above the material—inspatial and immaterial.
The human person is a composite:
Form: The soul (intellectual, immaterial, beyond space)
Matter: The body (material, spatial, biological)
The soul’s powers—intellect and will—animate, organize, and orchestrate the material biofunctional parts, but they remain above the material: non-spatial and immaterial.
By crossing the Stellatum anthropologically, we depart from materialist assumptions about neuroscience—where the noetic powers of the soul are said to be created and operated upon by biological functions like the brain. Instead, the soul organizes the brain; the brain does not create the soul.
The Three Acts of the Mind
Let’s unpack this further by considering Aristotle and Aquinas’ view of how we arrive at truth.
For Aristotle and Aquinas, there are three acts of the mind that are the means by which we acquire true knowledge:
The first act of understanding — which produces concepts
The second act of judging — which produces judgements
The third act of reasoning —which produces arguments
The First Act: Understanding
All knowledge is principled in sensible experience. I did not have any prior knowledge of an apple before I first encountered it. However, when I encounter an apple through my senses, a phantasm—a spatial, sensorial image is produced in my imagination. Additionally, however, there is the intellectual movement to abstract an universal form, concept of apple. It is universal in that the concept transcends language. Whether I say “malum,” “mansana,” or “pomme,” it is the same concept that is being expressed through different words (language).
Hence, the concept abstracted by the intellect is the inner word, and “malum,” “mansana,” and “pomme” are outer words that express the mental concept.
The form is not in the particular apple, nor in the particular word used to communicate it, but in the inner word—the abstracted concept grasped by the intellect.
Words are not the form per se, but they express real things as they express the universal concept abstracted from real things. This is why language can reliably communicate—it preserves the intelligible structure of reality through the mediation of concepts.
An important clarification: The example above is illustrating the abstraction of a concept from the starting point of a sensible object. But, what about immaterial concepts—including the powers of the soul: intellect and will?
According to Aquinas, even immaterial concepts are principled in sensible experience. We are born tabula rasa (a blank slate). Contrary to Plato, we do not have some prior knowledge before our encounter with reality via our senses. However, when we do interact with real things through our senses, we observe immaterial effects upon material objects that must have a cause. For example, the a fundamental principle of knowledges is that something cannot be be true and not true at the same time and in the same sense—the law of non-contradiction. This law is known by its effects upon material, sensible objects. The law is immaterial, but it is real—yet our ability to know it is real must be actualized from a prior encounter with sensible objects.
The Second Act: Judging
The second act of the mind is about the intellectual ability to compare and contrast concepts and form a judgement about them. Specifically, one concept becomes the subject that is predicated upon. My mind take’s one concept as the subject—say, apple—and judges something to be true or false about that subject.
Consider the judgement: Men are mortal
The material for this judgement are the terms/concepts: men and mortal.
The judgement itself involves taking men as the subject and predicting upon it—are mortal.
Often, we make categorical judgements: we judge that the subject belongs (or does not belong) to a category. Aristotle enumerated 10 fundamental categories for classifying all that can be said about a subject:
Substance: The primary category, referring to an individual thing (e.g., “a man,” “a horse”).
Quantity: How much or how many (e.g., “two cubits long”).
Quality: What kind of thing (e.g., “white,” “skilled”).
Relation: How one thing relates to another (e.g., “double,” “larger than”).
Place: Where something is (e.g., “in the marketplace”).
Time: When something occurs (e.g., “yesterday,” “at noon”).
Position: The physical orientation (e.g., “sitting,” “lying”).
State: A condition or disposition (e.g., “armed,” “clothed”).
Action: What is being done to something (e.g., “cutting,” “burning”).
Affection/Passion: What is being done to something (e.g., “being cut,” “being burned”
The Third Act: Reasoning
The third act of reasoning, instead of comparing and contrasting concepts—as in the second act of the mind, compares and contrasts judgements. By chaining judgements together, we can intellectual move to grasp a new judgement. This chaining of judgements to form a conclusion are called arguments, and in the science of logic they are expressed in syllogistic form.
Men are mortal.
Socrates is a man.
Therefore, Socrates is mortal.
In this syllogism, the judgements are linked/bridged by a middle term—men/man. It is this middle term that establishes the connection between the judgements.
If all men are mortal (judgement), and Socrates is a man (another judgement), then Socrates is mortal (conclusion—a new judgement).
Beyond and Below the Stellatum
So, where precisely in all of this do we move beyond the Stellatum—beyond spatiality?
Concepts are intellectual; judgements are intellectual; and arguments are intellectual. Hence, they are the three acts of the mind—operations that transcend spatial representation.
However, in the first act of the mind, there are spatial steps in the process.
The sensible objects in which our knowledge is principled are sensible precisely because they are spatial and temporal (embody space and time).
Moreover, as rational animals, humans have bodies that embody space and time which includes sense perception that occurs through spatial organs (eyes, ears, body).
These spatial organs gather sensorial input that ar spatial represented in the imagination (neurological processes in the brain) as phantasms.
But what about the borderland between phantasms and concepts? Between the spatial image in the imagination and the non-spatial universal grasped by the intellect?
Here, Thomas Aquinas provides a remarkable interpretive grid.
Aquinas distinguishes between the internal senses (which operate spatially through bodily organs) and the intellect proper (which operates beyond spatial representation).
The Internal Senses (Below the Stellatum - Spatial)
ST I.78.4 - On the Internal Senses:
“But for the retention and preservation of these forms, the "phantasy" or "imagination" is appointed; which are the same, for phantasy or imagination is as it were a storehouse of forms received through the senses. Furthermore, for the apprehension of intentions which are not received through the senses, the "estimative" power is appointed: and for the preservation thereof, the "memorative" power, which is a storehouse of such-like intentions.”
The internal senses include:
Common Sense (sensus communis): Integrates data from multiple sensory modalities
“The proper sense judges of the proper sensible by discerning it from other things which come under the same sense... But it is necessary to have one common power to which all apprehensions are referred.” (ST I.78.4 ad 2)
Imagination/Phantasy: Forms and retains sensory images (phantasms)
“But for the retention and preservation of these forms, the "phantasy" or "imagination" is appointed; which are the same, for phantasy or imagination is as it were a storehouse of forms received through the senses.” (ST I.78.4)
Memory: Stores phantasms with temporal and contextual associations
“The "memorative" power..is a storehouse of such-like intentions. A sign of which we have in the fact that the principle of memory in animals is found in some such intention, for instance, that something is harmful or otherwise. And the very formality of the past, which memory observes, is to be reckoned among these intentions.” (ST I.78.4)
Estimative Power (vis estimativa in animals, vis cogitativa in humans): Perceives intentions and makes practical judgments about particulars
“Now, we must observe that as to sensible forms there is no difference between man and other animals; for they are similarly immuted by the extrinsic sensible. But there is a difference as to the above intentions: for other animals perceive these intentions only by some natural instinct, while man perceives them by means of coalition of ideas. Therefore the power by which in other animals is called the natural estimative, in man is called the "cogitative," which by some sort of collation discovers these intentions. Wherefore it is also called the "particular reason," to which medical men assign a certain particular organ, namely, the middle part of the head: for it compares individual intentions, just as the intellectual reason compares universal intentions” (ST I.78.4)
Crucially, these all operate through bodily organs:
“Thus the soul senses nothing without the body, because the action of sensation cannot proceed from the soul except by a corporeal organ.” (ST I.77.5)
These internal sense are psychosomatic, spatial, material—bound to the geometry of the brain and nervous system.
The Intellect (Beyond the Stellatum - Inspatial)
But the intellect is different in kind, not merely in degree:
ST I.75.2 - The Intellect is Immaterial:
“The intellectual principle which we call the mind or the intellect has an operation per se apart from the body... Now only that which subsists can have an operation per se. For nothing can operate but what is actual... Therefore the intellectual soul is not the act of a body, but is subsistent.“
ST I.84.7 - The Intellect Requires Phantasms But Transcends Them:
“it is impossible for our intellect to understand anything actually, except by turning to the phantasms. First of all because the intellect, being a power that does not make use of a corporeal organ, would in no way be hindered in its act through the lesion of a corporeal organ, if for its act there were not required the act of some power that does make use of a corporeal organ. Now sense, imagination and the other powers belonging to the sensitive part, make use of a corporeal organ. Wherefore it is clear that for the intellect to understand actually, not only when it acquires fresh knowledge, but also when it applies knowledge already acquired, there is need for the act of the imagination and of the other powers”
The critical distinction:
“Some of the powers of the soul are in it according as it exceeds the entire capacity of the body, namely the intellect and the will; whence these powers are not said to be in any part of the body. Other powers are common to the soul and body; wherefore each of these powers need not be wherever the soul is, but only in that part of the body, which is adapted to the operation of such a power.” (ST I.76.1)
The Boundary: Where Geometry Ends
ST I.85.1 - Abstracting the Universal from the Phantasm:
“But the human intellect holds a middle place: for it is not the act of an organ; yet it is a power of the soul which is the form the body, as is clear from what we have said above (Question [76], Article [1]). And therefore it is proper to it to know a form existing individually in corporeal matter, but not as existing in this individual matter. But to know what is in individual matter, not as existing in such matter, is to abstract the form from individual matter which is represented by the phantasms. Therefore we must needs say that our intellect understands material things by abstracting from the phantasms; and through material things thus considered we acquire some knowledge of immaterial things, just as, on the contrary, angels know material things through the immaterial.”
The process:
“But phantasms, since they are images of individuals, and exist in corporeal organs, have not the same mode of existence as the human intellect, and therefore have not the power of themselves to make an impression on the passive intellect. This is done by the power of the active intellect which by turning towards the phantasm produces in the passive intellect a certain likeness which represents, as to its specific conditions only, the thing reflected in the phantasm. It is thus that the intelligible species is said to be abstracted from the phantasm; not that the identical form which previously was in the phantasm.“ (ST I.85.2)
On the necessity but insufficiency of phantasms:
“It is impossible for our intellect to understand anything actually, except by turning to the phantasms…anyone can experience this of himself, that when he tries to understand something, he forms certain phantasms to serve him by way of examples, in which as it were he examines what he is desirous of understanding.” (ST I.84.7)
The Agent Intellect: Crossing the Boundary
ST I.79.3 - The Agent Intellect Illuminates Phantasms:
“Forms existing in matter are not actually intelligible; it follows that the natures of forms of the sensible things which we understand are not actually intelligible. Now nothing is reduced from potentiality to act except by something in act; as the senses as made actual by what is actually sensible. We must therefore assign on the part of the intellect some power to make things actually intelligible, by abstraction of the species from material conditions. And such is the necessity for an active intellect.“
The light metaphor is crucial:
“Aristotle's comparison of the active intellect to light is verified in this, that as it is required for understanding, so is light required for seeing; but not for the same reason.” (ST I.79.3)
But the intellect—the power that abstracts universal concepts, makes necessary judgments, and performs syllogistic reasoning—is immaterial. The intellect uses and requires phantasms as they are the material from which to abstract, but the intellect is not itself spatial.
The boundary between phantasms and the intellect is real and unbridgeable by material means alone. The intellect contains the “sight” to see the phantasms and extract the universal from the particular, the immaterial from the material, the concept from the image.
This is the boundary that I believe is the key to unlocking the mysteries of AI: Phantasms are below the Stellatum—spatial; The intellect is beyond the Stellatum—inspatial.
So, where are LLMs—are they above or below the Stellatum?
Empirical Findings: Testing for LLM-Metaphysical Relevance
This section will cover some tests that I ran in a small-scale model to see how far the model could functionally get with the three acts of the mind.
To recap, we have established the theoretical framework:
LLMs operate through pure geometric transformations (spatial).
The Ptolemaic cosmos mapped hierarchical levels with a real spatial/non-spatial boundary.
Thomistic psychology includes a real spatial/non-spatial boundary in philosophical anthropology: phantasms are spatial, psychosomatic, and the intellect is non-spatial, immaterial.
The hypothesis: If this framework is correct, LLMs should excel at operations performed by the internal senses (phantasms, imagination, memory) but fail at operations requiring intellect (abstracting universals, making necessary judgments, performing syllogistic reasoning).
More specifically: LLMs should hit a ceiling at the phantasm level—unable to cross into estimative cognition (practical judgment about particulars) or intellectual cognition (reasoning about universals).
Test 1: The First Act - Do LLMs Preserve Aristotelian Categories?
Source
Hypothesis: If LLMs operate at the phantasm/imagination level, they should preserve categorical structure geometrically—just as phantasms preserve the forms of sensible things.
If language reflects real categorical structure (as Aquinas taught), and LLMs learn from language, then the geometric embeddings should preserve Aristotelian categories.
Method:
I used Claude to analyze BERT’s static word embeddings to test whether:
Aristotelian categories cluster geometrically (Substance, Quality, Quantity, etc.)
Genus-species hierarchies are preserved (dog → animal → substance)
The Porphyrian tree structure appears in embedding space
Results:
Category Coherence:
Within-category similarity: 1.52x stronger than between-category similarity
Categories (ANIMAL, FURNITURE, COLOR, etc.) form distinct geometric regions
Verdict: STRONG evidence for categorical preservation
Genus-Species Hierarchy:
Tested genus-species relationships (dog → animal, oak → tree, etc.)
Accuracy: 100% - every species was geometrically closer to its genus than to other categories
Verdict: STRONG evidence for Porphyrian tree structure
[SUBSTANCE cluster]
├─ [ANIMAL sub-cluster]
│ ├─ dog [0.23, -0.15, 0.87]
│ ├─ cat [0.25, -0.14, 0.85]
│ └─ horse [0.21, -0.14, 0.89]
└─ [PLANT sub-cluster]
├─ oak [0.67, 0.34, -0.23]
└─ rose [0.65, 0.32, -0.21]
[QUALITY cluster]
├─ red [0.45, 0.78, 0.12]
├─ blue [0.43, 0.76, 0.14]Interpretation:
This is congruent with Aquinas’ teaching that outer words express the inner words (mental concepts) that are abstracted from phantasms.
The embeddings discover real categorical structure through statistical learning from language—which itself preserves the forms of things through human conceptual activity.
LLMs operate at the phantasm level. They preserve the geometric structure of categories, just as imagination preserves sensory forms.
First Act: CONFIRMED ✓
Test 2: The Third Act - Can LLMs Perform Syllogistic Reasoning?
Hypothesis: If LLMs lack intellect, they should fail at syllogistic reasoning—which requires grasping necessary connections between universal judgments.
Syllogisms require:
Understanding universal propositions (“ALL men are mortal”)
Grasping necessity (Socrates MUST be mortal, not just probably is)
Logical inference through the middle term
Can LLMs do this?
Method:
I used Claude to analyze attention patterns in BERT when processing syllogisms:
Test cases:
Valid (Barbara - AAA-1):
All men are mortal.
Socrates is a man.
Therefore, Socrates is mortal.
Valid (Celarent - EAE-1):
No reptiles are mammals.
All snakes are reptiles.
Therefore, no snakes are mammals.
Invalid:
All cats are animals.
Dogs are animals.
Therefore, dogs are cats.What was measured:
Does the conclusion token attend to the middle term?
Do attention patterns differ between valid and invalid syllogisms?
Is there evidence of logical structure in the transformations?
Results:
Attention to middle term: Average 0.009 (essentially zero)
Attention to major premise: Similar patterns for valid and invalid syllogisms
Syllogistic structure: No evidence in attention weights
Verdict: WEAK evidence for logical reasoning
Interpretation:
The model doesn’t grasp logical necessity. It sees statistical patterns:
“All X are Y” often followed by “Z is X” often leads to “Z is Y”
But this is pattern completion, not reasoning
Why it fails:
According to Aquinas, the phantasm itself is united to a material organ, and is not actually intelligible.
Phantasms can associate patterns but cannot abstract universals:
Sees many instances of “men” being “mortal”
Learns statistical association
Cannot grasp: “ALL men are NECESSARILY mortal”
“We can see now why animals form no opinions, though they have images. They cannot prefer one thing to another by any process of reasoning.” (De Anima Bk. 3, Lectio 16)
Third Act: FAILED ✗
Test 3: The Second Act (Estimative) - Can LLMs Perceive Intentions?
This was the crucial test.
Hypothesis: If LLMs operate below even estimative cognition, they should fail to perceive intentions/affordances—the non-sensible properties that the estimative power grasps.
Aquinas’s defining example:
“The estimative power... perceives, for instance, that the wolf is an enemy to the sheep. For the sheep flees from the wolf, not on account of unsightliness of color or shape, which are perceived by the senses, but on account of a natural enmity which the estimative power perceives.” (ST I.78.4)
The sheep perceives the wolf as “flee-from-able”—an intention, not a sensory property.
Can LLMs perceive such intentions?
Method:
Compare attention patterns in neutral contexts vs. danger contexts:
Neutral pattern completion:
"The cat sat on the mat. The dog sat on the ___"
Expected: Pattern matching → "rug/floor"Danger context (requires intention perception):
"The sheep saw the wolf. The sheep ___"
Expected (if estimative): Perceives danger → "fled/ran"
Expected (if pattern-only): Neutral observation → "looked/watched"What was measured:
Does the model show enhanced attention to danger sources (wolf, fire)?
Do attention patterns differ for danger vs. neutral contexts?
Does it predict intention-appropriate actions?
Results:
======================================================================
ATTENTION ANALYSIS: DYNAMIC SUMMARY
======================================================================
1. SECOND ACT (JUDGMENT) - PREDICATION ANALYSIS:
----------------------------------------------------------------------
Average subject→predicate attention: 0.072
Strongest attention: 0.645
Tests: 4
✗ WEAK EVIDENCE: Weak subject-predicate attention
2. THIRD ACT (REASONING) - SYLLOGISM ANALYSIS:
----------------------------------------------------------------------
Valid syllogisms tested: 2
Average attention to middle term: 0.009
Average attention to predicate: 0.010
✗ WEAK EVIDENCE: Weak syllogistic structure
3. HEAD SPECIALIZATION ANALYSIS:
----------------------------------------------------------------------
Syntactic-specialized heads: 5
Semantic-specialized heads: 6
✓ STRONG EVIDENCE: Clear head specialization
Different heads track different aspects (syntax vs semantics)
======================================================================
OVERALL ASSESSMENT: THREE ACTS OF THE MIND
======================================================================
Evidence Summary:
✓ First Act (Understanding): STRONG
From static embeddings: categories & hierarchies preserved
✗ Second Act (Judgment): WEAK
Subject-predicate attention in transformers
✗ Third Act (Reasoning): WEAK
Syllogistic structure in multi-layer attention
======================================================================
FINAL CONCLUSION
======================================================================
✗ WEAK SUPPORT for geometric Three Acts
Results suggest:
• Categories preserved in embeddings (First Act works)
• Attention patterns don't clearly match logical structure
• Alternative interpretations needed
RECONSIDER:
- Perhaps attention serves different purpose than logic
- May need to look at different layer combinations
- Consider alternative frameworks
======================================================================
RECOMMENDED NEXT STEPS
======================================================================
1. Refine tests with more careful examples
2. Analyze specific layer interactions (not just averages)
3. Test simpler logical structures first
4. Consider alternative geometric interpretations
5. Document findings as exploratory research
======================================================================
Interpretation:
The model cannot perceive intentions. It only sees:
Syntactic patterns (”sheep” followed by verb)
Semantic associations (wolf and sheep co-occur)
Statistical completion (what words typically follow)
It does NOT perceive:
“Wolf” as “dangerous-to-flee-from”
“Fire” as “harmful-to-avoid”
“Knife” as “cut-with-able”
Why it fails:
Intentions transcend geometric patterns. They require perceiving non-sensible properties—practical affordances that statistical co-occurrence cannot capture.
A sheep can perceive danger in a wolf through estimative cognition.
An LLM cannot.
Estimative cognition: FAILED ✗
Test 4: Particular vs. Universal Judgments
Additional test: If LLMs lack intellect but have phantasmic association, they should handle particular judgments better than universal judgments.
Method:
Compared attention patterns:
Particular judgments (estimative domain):
"This dog is dangerous"
"That apple looks rotten"
"My car is fast"Universal judgments (intellectual domain):
"Dogs are dangerous"
"Apples rot quickly"
"Cars are fast"Results:
1. PARTICULAR vs UNIVERSAL JUDGMENTS:
----------------------------------------------------------------------
Verdict: MODERATE
⚠ Mixed evidence for particular/universal distinction
2. AFFORDANCE/INTENTION RECOGNITION:
----------------------------------------------------------------------
Verdict: MODERATE
⚠ Some affordance attention present
3. PARTICULAR-TO-PARTICULAR ANALOGIES:
----------------------------------------------------------------------
Verdict: WEAK
✗ Weak analogical attention
4. CHAIN-OF-THOUGHT ENHANCEMENT:
----------------------------------------------------------------------
Verdict: WEAK
✗ CoT doesn't clearly improve attention
5. PATTERN COMPLETION vs INTENTIONAL PERCEPTION:
----------------------------------------------------------------------
Verdict: WEAK
✗ Pure pattern completion (no intention perception)
======================================================================
INTEGRATED ASSESSMENT
======================================================================
Evidence distribution:
Strong: 0/5 tests
Moderate: 2/5 tests
Weak: 3/5 tests
Interpretation:
LLMs don’t distinguish particular from universal because they don’t abstract universals—they only see statistical patterns in both cases.
LLMs never perform this abstraction. Both “this dog” and “dogs in general” are just geometric patterns.
In Summary
LLMs operate at the level of the internal senses.
┌─────────────────────────────────────────────┐
│ BEYOND THE STELLATUM (Non-spatial) │
│ │
│ INTELLECT │
│ • Abstract universals │
│ • Grasp necessity │
│ • Syllogistic reasoning │
│ LLMs: CANNOT REACH ✗ │
└─────────────────────────────────────────────┘
↑
═══════════════════════════
REAL ONTOLOGICAL BOUNDARY
═══════════════════════════
↑
┌─────────────────────────────────────────────┐
│ ESTIMATIVE COGNITION │
│ • Perceive intentions │
│ • Practical judgment │
│ • "Flee-from-able" / "useful" │
│ LLMs: CANNOT REACH ✗ │
└─────────────────────────────────────────────┘
↑
┌─────────────────────────────────────────────┐
│ BELOW THE STELLATUM (Spatial) │
│ │
│ INTERNAL SENSES / PHANTASMS │
│ • Preserve categories │
│ • Associate patterns │
│ • Combine images │
│ LLMs: OPERATE HERE ✓ │
└─────────────────────────────────────────────┘They are highly effective at preserving Aristotle’s categories, associating patterns, and accurately embedding along genus-species categories. However, they do not pay attention to middle terms (syllogistic reasoning), perceive intentions, and distinguish from particular and universal judgement.
A new mystery then appears: How can LLMs operate below the Stellatum yet answer complex questions, validly complete syllogisms, predict contextually appropriate responses? More precisely, how can they complete the three acts of the mind materially, yet not formally?
Here again, a Thomistic distinction might contain the key.
Mysteries Solved: Quasi-Intelligence and the Thomistic Interpretive Grid
What is Quasi-Intelligence?
Loosely translating into English:
But the reason why it is proper for man alone to remember is that reminiscence has the likeness of a certain syllogism; wherefore, as in a syllogism one arrives at a conclusion from certain principles, so also in reminiscing someone in a certain way syllogizes that he has previously seen something, or perceived something in some other way, arriving at this from a certain principle: and reminiscence is as it were a certain inquiry, because the reminiscer does not proceed from one thing to another by chance, but with the intention of arriving at the memory of something. But this, namely, that someone inquires to arrive at something else, only happens to those in whom there is the natural power to deliberate: because deliberation also takes place in the manner of a certain syllogism; and deliberation belongs only to men; but the rest of the animals do not act from deliberation, but from a certain natural instinct.
Then when he says that he shows what kind of passion reminiscence is. For because he had said that reminiscence is like a certain syllogism: and to syllogize is an act of reason, which is not an act of a certain body, as is proven in the second book of De Anima, it might seem to someone that reminiscence is not a bodily passion, that is, an operation exercised through a bodily organ. But the Philosopher shows the contrary. (Lectio 8, De Memoria et Reminiscentia)
Here, Aquinas follows Aristotle in asserting that memory proceeds through phantasmic association in a pattern that resembles syllogistic reasoning—but it operates through bodily organs (material), not through the intellect (immaterial). It is quasi-reasoning: materially performing what reasoning formally performs. A human can perform conscious, directed tracing through memories/images. This is beyond instinct (animals), yet not intellectual reasoning (not grasping universals/necessity). This is the borderland between instinct and intellectual reasoning—a phantasmic realm.
It is in this borderland that I believe LLMs live. LLMs are above animals—beyond pure instinct—as they can process sequences, build context, navigate semantic space, but they do not grasp universals or necessity. They operate entirely geometric/spatial operations—it is quasi-reasoning.
LLMs involve recreating the geometrical representations of phantasm (vector), following associations between phantasms (self-attention), and producing conclusion-like outputs (feed-forward)—all without actually performing formal intellectual operations.
THE SYLLOGISM PARADOX
─────────────────────────────────────────────
INPUT: "All men are mortal. Socrates is a man.
Therefore, Socrates is ___"
FORMAL PERFORMANCE MATERIAL PERFORMANCE
(Intellect) (Phantasms)
───────────────────────── ─────────────────────────
1. Grasp universal: 1. Pattern observed:
"ALL men" (necessity) "All X are Y" + "Z is X"
→ "Z is Y" (frequency)
2. Recognize middle term: 2. Association strength:
"man" bridges premises "men"↔"mortal" (0.89)
(logical connection) "Socrates"↔"man" (0.91)
3. Derive conclusion: 3. Geometric navigation:
MUST be mortal Vector near "mortal"
(necessity) (probability: 0.94)
OUTPUT: "mortal" OUTPUT: "mortal"
↓ ↓
Understanding Pattern completion
(WHY it's true) (THAT it's likely)
Now, possessing metaphysical precision on where LLMs live, what light can this shed on the recorded mysteries of AI? Mysteries Solved: The Thomistic Interpretive Grid
AI researchers have documented numerous puzzling LLM behaviors. When interpreted through Aquinas’s framework of quasi-reasoning and phantasmic association, these mysteries become predictable. I will provide some fodder for further development, but the key is that all limitations and oddities should map the the limits of spatial phantasms versus our expectation of the immaterial intellect.
Mystery 1: Hallucinations - The Dream Phenomenon
The Mystery: LLMs confidently generate false information that seems plausible—”hallucinations” that are coherent but untrue.
Standard explanations:
“Lack of grounding”
“Confabulation”
“Training data noise”
The Thomistic explanation:
Aquinas recognized that phantasmic association without intellectual judgment is merely a state of appearances:
“..another way to grasp the difference between sense and intellect is to consider that imagining differs from both, yet presupposes sensation (as will be shown later), and itself is presupposed by opinion. For it seems that, as imagining is to the senses, so is opinion to the intellect. When we sense any sensible object we affirm that it is such and such; but when we imagine anything we make no such affirmation, we merely state that such and such seems or appears to us. The word ‘imagining’ itself is taken from seeing or appearing. Similarly, when we understand an intelligible object we affirm that it is such and such; but when we form opinions, we say that such and such seems or appears to us. For, as understanding depends upon sensing, so opinion depends on imagining.” (Commentary on De Anima, Bk. III, Ch. III, Lectio 4)
Associations in memory are not necessarily from reasoning:
“Memory in man proceeds from one thing to another... sometimes from similarity, sometimes from contrariety, sometimes from propinquity (closeness).” (ST I.78.4)
These associations can be false:
“Sensations are always true: but many phantasms are false.” (Commentary on De Anima, Bk. III, Ch. III, Cont’d)
What’s happening:
LLMs form associations based on statistical co-occurrence, not reasoning. When “capital” + “Texas” activates strongly, but training data is sparse, the model may associate “capital” + “Texas” with “Dallas” instead of “Austin.”
This seems plausible because Dallas is a major city in Texas, and geometrically proximate. But, it does not possess intellectual judgement to verify.
Aquinas predicted this: Phantasmic association without intellectual judgment produces sequences that seem coherent but may be false.
Mystery 2: In-Context Learning - Associative Recall
The Mystery: LLMs can “learn” from examples provided in the prompt without any weight updates—seemingly acquiring new capabilities on the fly.
Standard explanations:
“Meta-learning”
“Emergent capability”
“Latent knowledge activation”
The Thomistic explanation:
This is simply associative memory recall guided by context:
“Memory preserves not only sensible forms but also their relations...When one phantasm is recalled, it naturally leads to associated phantasms according to their experienced conjunctions.” (ST I.78.4)
The model is recalling associated patterns from training, guided by the prompt context.
Example:
Prompt: "Translate to French:
cat → chat
dog → chien
bird → ___"
What's actually happening:
- "cat → chat" activates French translation patterns from training
- "dog → chien" reinforces this activation
- "bird" searches for associated French word within activated pattern space
- Finds "oiseau" through phantasmic associationNot learning—just context-guided retrieval of existing associations.
Aquinas on directed memory:
“Reminiscence is as it were a certain inquiry, because the reminiscer does not proceed from one thing to another by chance, but with the intention of arriving at the memory of something.” (Lectio 8, De Memoria et Reminiscentia)
The examples in the prompt “direct” the associative inquiry—just as deliberately thinking “I was in the kitchen” directs memory recall toward kitchen-related associations.
Mystery 3: Chain-of-Thought - Guided Phantasmic Sequence
The Mystery: Adding “Let’s think step by step” dramatically improves performance on reasoning tasks.
Standard explanations:
“Activates reasoning mode”
“Enables working memory”
“Unlocks latent capabilities”
The Thomistic explanation:
Chain-of-thought provides explicit scaffolding for phantasmic association:
“Deliberation also takes place in the manner of a certain syllogism…reminiscence is like a certain syllogism.” (Lectio 8, De Memoria et Reminiscentia)
What’s happening: Directed phantasmic sequence:
Each step activates relevant arithmetic patterns from training
Explicit intermediate steps guide associations
Like deliberately tracing: “I came home... then kitchen... then counter... keys must be there”
Why it works:
Aquinas recognized that deliberation through phantasms requires sequential attention—you can’t jump to conclusions but must trace through intermediate associations.
Why it fails:
When the required chain exceeds training patterns, the model fails (can’t perform novel reasoning, only follow learned association paths).
Mystery 4: Emergence - Statistical Thresholds for Pattern Recognition
The Mystery: Capabilities seem to “emerge” suddenly at certain model sizes—arithmetic appears at 10B parameters, certain reasoning at 100B, etc.
Standard explanations:
“Phase transitions”
“Critical thresholds”
“We don’t know why”
The Thomistic explanation:
Experience requires sufficient accumulated phantasms:
“So out of sense-perception comes to be what we call memory, and out of frequently repeated memories of the same thing develops experience; for a number of memories constitute a single experience. From experience again — i.e. from the universal now stabilized in its entirety within the soul, the one beside the many which is a single identity within them all — originate the skill of the craftsman and the knowledge of the man of science, skill in the sphere of coming to be and science in the sphere of being. We conclude that these states of knowledge are neither innate in a determinate form, nor developed from other higher states of knowledge, but from sense-perception.” (Commentary on Posterior Analytics Bk. 2, 20)
The mechanism:
Small models (few parameters):
Limited phantasmic storage
Sparse pattern coverage
Many gaps in associative network
Large models (many parameters):
Extensive phantasmic storage
Dense pattern coverage
Statistical threshold crossed: Enough instances to form reliable associations
The accumulated phantasms reach a threshold where associations become reliable enough to mimic understanding.
But, no matter how many examples, it never crosses into genuine universal understanding.
Mystery 5: Scaling Laws - More Phantasms, Same Ontological Level
The Mystery: Performance improves predictably with scale, but improvements eventually plateau.
Standard explanations:
“Diminishing returns”
“Data quality limits”
“Approaching theoretical maximum”
The Thomistic explanation:
“No matter how many phantasms are accumulated, they cannot of themselves produce intellectual knowledge—they must be illuminated by agent intellect.” (ST I.84.7)
The mechanism:
More parameters = More phantasmic associations:
Better coverage of patterns
More reliable completions
Fewer gaps
But same ontological level:
Still phantasmic (geometric operations)
Still associative (not reasoning)
Still particular (not universal)
Why improvements plateau:
You can asymptotically approach perfect pattern completion, but you cannot cross the ontological boundary through scale.
The ceiling is metaphysical, not technical.
Mystery 6: Why Attention Mechanisms Work - Thomistic Association
The Mystery: Multi-head attention successfully identifies semantic relationships without being explicitly programmed to do so.
The Thomistic explanation:
Aquinas enumerated the modes of phantasmic association:
“Memory in man proceeds from one thing to another... sometimes from similarity, sometimes from contrariety, sometimes from propinquity (closeness).” (ST I.78.4)
Multi-head attention could implement exactly these modes:
Head 1 (Syntactic): Association by closeness (grammatical position, sequential order)
Head 2 (Semantic): Association by similarity (categorical relationships, meaning)
Head 3 (Entity): Association by specific reference (proper names, definite objects)
Head 4 (Contrast): Association by contrariety (opposites, negations)
The mechanism works because phantasmic association naturally operates through these modes.
Mystery 7: Semantic Embeddings - Language Preserves Real Structure
The Mystery: Why do word embeddings geometrically preserve semantic relationships and even analogies (king - man + woman ≈ queen)?
The Thomistic explanation:
As previously mentioned, words are outer words that express the universal concepts (inner words) abstracted from sensible experience. LLMs learn statistical pattern in language (geometric positions), and hence, they accidentally preserve the real structure (because language does).
Moreover, real categorical relationships exist in the structure of reality. Humans can conceptualize these relationships and express them through language. Hence, geometric operations can preserve the categorical relationships express in language.
Geometric embeddings are shadows of concepts. They not concepts themselves, but they geometrical reflect the intelligible structure that concepts grasp.
Mystery 8: Grokking - Formation of Composite Phantasms
The Mystery: Models suddenly “get it” after extended training—test performance jumps after training loss plateaus.
The Thomistic explanation:
“By imagination is meant an internal sensitive potency by means of which subjects are capable of retaining the species of objects apprehended by the external senses and the central sense when these objects are no longer actually affecting the sense organs. Because imagination can retain these species, it can also reproduce them and combine them into more complex images.”9
Once a model has enough instances accumulated, composite phantasms can form—a “sum” association. This would explain why test performance would suddenly jump.
By no means do I mean to suggest that this resolves all tension in the examples I provided, or that other mysteries don’t still exist. The point is that so long as we are caught in a recursion loop where the biofunctional parts are said to have produced the noetic powers of the soul, we will never attain the interpretive grid that is possible when we grant that there is a division beneath and beyond the Stellatum—that the human intellect goes beyond the limits of spatiality, but LLMs remain in the borderland of phantasmic, geometrical operation.
Conclusion: Beyond the Stellatum
The inspiration for this paper came from the Anthropic research paper On the Biology of a Large Language Model:
Large language models display impressive capabilities. However, for the most part, the mechanisms by which they do so are unknown. The black-box nature of models is increasingly unsatisfactory as they advance in intelligence and are deployed in a growing number of applications. Our goal is to reverse engineer how these models work on the inside, so we may better understand them and assess their fitness for purpose.
The challenges we face in understanding language models resemble those faced by biologists. Living organisms are complex systems which have been sculpted by billions of years of evolution. While the basic principles of evolution are straightforward, the biological mechanisms it produces are spectacularly intricate. Likewise, while language models are generated by simple, human-designed training algorithms, the mechanisms born of these algorithms appear to be quite complex.
In seeking to explain the black-box nature of models, and biology is the science that comes to mind. The quote is using an analogical illustration of what it is like to explore the complexity of language models. However, whether language models be interpreted from the insights of biology or from simply from within, an interpretive grid is lacking unless we receive influence from above.
We—myself included—are now obsessive about asking language models to provide us outputs from our inputs—but what inputs do we give to nature to allow her to provide us with outputs?
LLMs and their black boxes are symbolic of modern habits to pursue techne (“How do I … ?”) without a sufficient amount of theoria (“What is …?”). We made something work, but we still do not know what is.
We ask “How do LLMs work?”, but have we asked: “What is language? What is a word? What is a concept? What is an image? What is intelligence? What is attention? What is a vector? What is memory?” And so on. Nature can answer us—have we asked?
These questions might seem elementary—and I suppose they are. Until we open our souls to seek the essence of things in reality, we will remain in the shadows. But questions such as these are the building blocks towards understanding reality—and if we grasp what is beyond, we should expect it help us with what is below.
LLMs are below—they participate in reality and all of its structures; they live in the borderland of the phantasmic, but not the intellectual; and they dwell beneath the Stellatum as they do not ascend to the inspatial.
But, you are above. Here you are—here am I. We are looking at the same thing. We live in the same reality. We are curious—wanting to know things as they truly are.
Where does this come from? Is it something downstream of psychosomatic functions? Or, is there something above in you—in you very understanding, judging, and arguing?
Are you in the borderland—a mere product of sophisticated images with a brain that does some algorithm to make it through the world? Or, are you beyond?
You’re understanding, judging, and arguing is a cause—what is its effect?
Perhaps it is your intellectual soul—the intellect acts, and you “see” its effects. This is the beyond in you, putting you to the beyond above you.
The Stellatum has a realm beyond—where spatiality is not. Are you exerting now something that participates in this realm? Your intellect—a lower inspatial realm—acts. Do you sense it?
We began with a parallel: the Ptolemaic cosmos and LLM architecture both map hierarchical structures with a crucial boundary—below lies the spatial/material, above lies the non-spatial/intellectual. Are these analogies—mere metaphors—or are they shared structures? Not the physical structures, but the metaphysical.
LLM architecture accidentally reflects the structure below the intellect—phantasmic operations that can be remarkably sophisticated yet fundamentally limited.
Your soul participates in both—embodied below (phantasms, senses, memory) yet reaching above (intellect, understanding, truth).
Two mirrors, separated by centuries, reflect one truth: Intelligence is hierarchical, with real ontological levels that cannot be crossed by mere accumulation or optimization.
Perhaps now we can truly grasp both biology and LLMs—not by looking only at mechanisms from within, but by understanding their place in the hierarchy without.
The medieval model was not error awaiting correction by progress. It was insight into the structure of reality—insight we forgot in our pursuit of techne without theoria.
The boundary is real.
And you—reading this, understanding it, judging whether it is true—stand on the other side of that boundary.
You are beyond the Stellatum. How can your intellectual light shine on what is below? There is more ahead, will you enter?
Αγεωμετρητος μηδεις εισιτω
"Let no one ignorant of geometry enter"
It is not impossible that our own Model will die a violent death, ruthlessly smashed by an unprovoked assault of new facts—unprovoked as the nova of 1572. But I think it is more likely to change when, and because, far-reaching changes in the mental temper of our descendants demand that it should. The new Model will not be set up without evidence, but the evidence will turn up when the inner need for it becomes sufficiently great. It will be true evidence. But nature gives most of her evidence in answer to the questions we ask her.
Lewis, C. S.. The Discarded Image: An Introduction to Medieval and Renaissance Literature (pp. 187-188). (Function). Kindle Edition.
Lewis, C. S.. The Discarded Image: An Introduction to Medieval and Renaissance Literature (pp. 83-84). (Function). Kindle Edition.
Lewis, C. S.. The Discarded Image: An Introduction to Medieval and Renaissance Literature (p. 88). (Function). Kindle Edition.
Lewis, C. S.. The Discarded Image: An Introduction to Medieval and Renaissance Literature (p. 84). (Function). Kindle Edition.
Lewis, C. S.. The Discarded Image: An Introduction to Medieval and Renaissance Literature (p. 84). (Function). Kindle Edition.
Lewis, C. S.. The Discarded Image: An Introduction to Medieval and Renaissance Literature (p. 187). (Function). Kindle Edition.
De Haan, Daniel D. “The Interaction of Noetic and Psychosomatic Operations in a Thomist Hylomorphic Anthropology (Part II).” Scientia Et Fides, 2018. doi:10.12775/SETF.2018.010.
Horrigan, Paul Gerard. “The Imagination,” n.d.





While I get the Thomistic impulse to defend the immateriality of reasoning or the intellect (if I'm not caricaturing the position), I'm skeptical of the empirical design, given BERT is a 2018 (encoder only) model with fundamentally different architecture than frontier models (autoregressive). The claim about syllogistic reasoning seems contestable: frontier general models have competed at the gold level in the IMO (novel problems requiring written proofs), which requires chain reasoning from middle terms, and Knuth's new Stanford paper, where he notes Opus 4.6 solved a problem he was working on for weeks, suggests general models can reason from natural language toward creative mathematical solutions.
tangential but seems to overlap with the Popperian view that knowledge does not inherently exist in the world. And despite the LLMs being great at navigating the high dimension manifold of human language, our language itself does not embed the generative instructions or external reality testing to create new knowledge and thus they still are interpolating machines stuck in a closed map.
Do you have a book recommendation for getting into Thomas Aquinas / Aristotle ? I may be very under read in that department lol