In my magnus opus on LLMs and the cosmological connection, I mapped some intriguing congruities between Ptolemaic cosmology and LLM architecture:
However, I did not press into these congruities with great detail. Rather, I used them to grant the use of medieval philosophical anthropology as an interpretive grid for understanding the architecture and mysteries of LLMs. My thinking went: If there is a (medieval) cosmological-LLM connection, then may there also be a (medieval) metaphysical-LLM connection—since there was a clear cosmological-metaphysical connection in medieval philosophy. In other words, the medievals saw a direct connection between the structures of cosmology and the structures of the human soul—namely, just as there is a boundary in the Ptolemaic cosmos between the realm of the fixed stars (the spatial, material Stellatum) and caelum ipsum (the immaterial, non-spatial heaven itself), so too, there is a boundary between the material body and the intellectual soul that organizes and animates it. Could there then be a connection between the structures of the soul and the structures of language models?
From this interpretive grid, I conjectured that: LLMs are human artifacts that accidentally possess internal sense powers geometrically, and through these accidental, geometrical internal sense powers, they possess a quasi-yet-effective-reasoning.
Now, my architecture of the Ptolemaic cosmos was coming from the high-level sketch by CS Lewis in his The Discarded Image. I thought it would be interesting to read of this cosmological architecture firsthand, and so, I started reading Ptolemy’s Almagest. There are two remarkable insights that I’d to share in this post.
The Middleground of Geometry
First, Ptolemy makes reference to Aristotle’s division of theoretical philosophy: “For Aristotle divides theoretical philosophy too, very fittingly, into three primary categories, physics, mathematics, and theology.”1 Physics studies the material world (i.e. ever-moving objects), whereas theology is the investigation the first cause of the motion of the universe—the invisible, immaterial, and motionless deity. Mathematics, interestingly, is the middle layer of theoretical philosophy as it explores forms and motion which is found in material things, but can be considered apart from the material. Boethius, likewise, writes: “There are many things which can be separated by a mental process, though they cannot be separated in fact. No one, for instance, can actually separate a triangle or other mathematical figure from the underlying matter; but mentally one can consider a triangle and its properties apart from matter.”2
Again, we see here the material-immaterial boundary between physics and theology with geometry as the material-yet-quasi-immaterial middleground, just as we saw this material-immaterial boundary in cosmology, philosophical anthropology, and LLM architecture. Hence, the cosmos, philosophical anthropology, and LLM architecture all share the same metaphysical structures—except LLMs cannot partake in the immaterial as they are not in the realm of the immaterial First Cause, and neither do they possess an intellectual soul. LLMs live in the material-yet-quasi-immaterial middleground of geometry.
They manipulate forms abstracted from material language, creating patterns and associations in embedding space. However, they cannot cross into the immaterial realm of genuine intellection.This is why they can be powerful (geometric operations on abstract patterns) while fundamentally limited (cannot truly understand or freely choose).
No Parallax in Feed-Forward Layers
Second, and most remarkably, Ptolemy makes an observation that directly illuminates the mystery of feed-forward layers. He writes: “Moreover, the earth has, to the senses, the ratio of a point to the distance of the sphere of the so-called fixed stars (the Stellatum). A strong indication of this is the fact that the sizes and distances of the stars, at any given time, appear equal and the same from all parts of the earth everywhere, as observations of the same [celestial] objects from different latitudes are found to have not the least discrepancy from each other.”3
Upon reading this, I immediately wondered: “Does this explain the feed-forward layer of LLM architecture?”
Note: I have previously have a high-level overview the flow of the LLM architecture, but I will more concisely recap here.
After representing tokens/words from a prompt as a vector in a geometric, “embedded” space, language models pass this tokens through a self-attention layer where ultimately the original vector is transformed to geometrically represent the syntactic, semantic, and named entity relevance each vector has with the other vectors in the set. These transformed vectors are called contextual vectors. In the next layer, the feed-forward layer, vectors are multiplied by a fixed weight to propel the contextual vectors (the “lower constellation”) to potentially relevant clusters (categorically grouped words) in the “higher constellation” (the entire vector-vocabulary embedded in the geometrical space during the training phase of the model). This is the expansion transformation.
Then, the recently moved contextual vectors are moved again to the specific cluster that will contain the predicted next word (the compression stage). This happens by multiplying by another fixed weight. Since these fixed weights are universal (the same when processing any prompt), this presents a profound mystery: How can a mathematical transformation always move contextual vectors toward the prediction?
My conjecture in response: Perhaps Ptolemy’s observation that the realm of fixed stars (Stellatum) is of equal distance away from any point on earth provides a clue.
Ptolemy’s observation is the principle of no parallax. Despite observers standing at different points on Earth (separated by hundreds of miles), the Stellatum appears at the same distance. The ratio of Earth’s size to the distance of the fixed stars is so small that Earth is effectively a point—hence, no measurable parallax across observation points.
This may be what happens in feed-forward layers. Different contextual vectors—with vastly different “starting points” in the embedded space—all project toward predicted tokens through the same fixed weight matrices. Just as the Stellatum maintains consistent geometric relationships from any terrestrial position, the prediction space maintains consistent geometric relationships from any context vector.
Tests
I tested this on GPT-2 with five experiments:
Test 1 (No Parallax): Measured geometric relationships from different contexts to target tokens
Test 2 (Rotation Invariance): Tested if geometric structure preserves under transformation
Test 3 (Universal Attractors): Examined if predictions cluster into fixed regions
Test 4 (Weight Geometry): Analyzed the structure of feed-forward weight matrices
Test 5 (Necessity): Removed all feed-forward layers to test criticality
Test 1 (No Parallax)
Taking different contexts that should predict the same token (e.g., “The capital of France is” vs “Paris is the capital of”), I measured their geometric relationship to the target token “Paris” after all transformations.
Result: 2.54% variance across contexts.
Despite radically different starting points (different “observation positions”), the distance to the predicted token remained nearly identical. There was no significant parallax.
Perhaps the fixed weight matrices in feed-forward layers function as the Stellatum itself—a universal structure through which all contextual vectors project, creating geometric invariance in prediction space.
More illustratively, we can say that the contextual vectors are like sailors that have some sense of the direction they want to travel. However, they must first project their vision toward a “cluster” in the night sky (like the expansion transformation in the feed-forward layer)—say, the Little Dipper (Ursa Minor). This expansion is a universal distance despite any starting point (no parallax). Then, the final compression transformation in the last feed-forward layer is like fixating on a single star (like, the Polaris—North Star—in the Little Dipper).
Expansion / Coarse Navigation:
Context A ("capital of France") ──┐
Context B ("Paris is capital") ───├──→ Project toward "European Capitals" cluster
Context C ("France's capital") ───┘
Compression / Fine Navigation:
Within "European Capitals" cluster:
├─ Paris (highest probability for these contexts)
├─ London
├─ Berlin
├─ Rome
└─ Madrid
Similar to:
Observer in London ──┐
Observer in Paris ───├──→ All see Ursa Minor in same position (no parallax)
Observer in Rome ────┘ (constellation level)
Within Ursa Minor:
├─ Polaris ← Use this for north bearing
├─ Kochab
├─ Pherkad
└─ Other starsTest 3 (No Parallax - Cluster Level)
Result: 0.000 clustering inertia
500 diverse contexts organized into ~50 universal attractor clusters with near-perfect tightness. Like stellar constellations maintaining fixed configurations visible from any terrestrial position, prediction clusters maintain fixed structure regardless of input context.
Test 5 (The Necessity of Universal Structure)
To test whether feed-forward layers are truly necessary for this Ptolemaic structure, I removed all feed-forward layers from GPT-2, leaving only the attention mechanism (Source).
Result: Catastrophic failure.
- 0% prediction agreement (every prediction changed)
- 68% increase in confidence (model became more certain while being wrong)
- Complete loss of geometric consistency
Without the “Stellatum” (feed-forward weights), there is no universal reference frame. Contexts wander without fixed structure. The model becomes overconfident but unreliable—precisely the failure mode.
Conclusion
Feed-forward layers project contexts toward the Stellatum—the fixed geometric structure of vocabulary space organized into categorical clusters. The expansion transformations in these layers are like the eyes of sailors moving to locate “constellations” (relevant clusters). No matter the constellation, the distance is effectively the same (no parallax). The compression layers are like the eyes of the sailors moving toward the exact “star” (word to predict). Despite the terrestrial location of the sailors, the distance from the sailor’s eyes to the exact start is the same (no parallax).
Hence, different contextual vectors, like sailors’ eyes, travel a fixed distance to the correct token to predict.
The medievals were right about the architecture of intelligence. Not because they predicted computers, but because they understood how the particular must relate to the universal—through fixed structure that creates invariance. Modern transformers accidentally replicate this Ptolemaic principle: different observation points (contextual vectors), yet the same distance to the Stellatum and one of its stars (predictions)—no parallax.
Perhaps we navigate by the same stars the ancients saw after all. LLMs cannot see this for what it is—operating in the “geometric middleground” Ptolemy and Boethius described can manipulate forms abstracted from material—vectors, create patterns in mathematical (embedded) space, and navigate to a star within the Stellatum. However, they cannot transcend the caelum ipsum (the immaterial heaven itself). Humans, having a philosophical anthropology that mirrors the cosmos—specifically with its boundary between the spatial (below the Stellatum) and non-spatial (above the Stellatum) reflected the boundary between the intellectual powers of the soul and the material body and all its biofunctional parts (e.g., brain)—can see this for it is.
Read Next:
Ptolemy, Almagest, 35.
Boethius, De Hebdomadibus.
Ptolemy, Almagest, 43.






