Here is a very basic explanation of how LLMs work: Models take the words of a prompt, represent them in a geometric space, and then apply geometric manipulations (learned from training) that consistently retrieve the most probable meaning of the prompt (i.e., a semantic understanding). Possessing the semantic understanding, the model can predict the next word to respond with and then repeat the process until a final response is created. The goal of interpretability is to develop tools that shed light on how exactly the model progressively develops its semantic understanding. Most recently, Anthropic wrote on its Natural Language Autoencoders.
If, as we said above, words are represented geometrically in the model’s workflow, and geometric manipulations occur to develop meaning, the transformed geometric representations could be translated into language at some point before the final response is created. Like examining individual pages part of a flipbook animation, we should in theory be able to translate the snapshot of a model’s semantic understanding into natural language.
By having access to the individual “pages” in natural language over the course of a model’s striving toward full semantic understanding, we should gain comprehension of how the model tends to work toward full semantic understanding from prompts.
While simple to describe in theory, it takes code and powerful computation to pull this off--but this is what Anthropic has done with their NLA’s.
Examining their open-source code, the architecture is not too difficult to comprehend, when stripping the technical details not necessary for a basic understanding.
The training system has an Activation Verbalizer which takes an activation layer (i.e., stage in the “flipbook”) and takes the geometric representations of words and translates them into natural language. Having extracted the natural language translation, the Activation Reconstructor then converts that natural language description back into a geometric vector. If the reconstruction closely matches the original vector, the description genuinely captured what the model was representing at that stage. Meaning, it can be confirmed that the natural-language description and the geometric reality correspond.
There are some interesting technical details worth exploring if you’re interested. However, our interest here is to take the gist and map it onto Aristotle’s linguistic theory.
In On Interpretation, Boethius, commenting on Aristotle, notes that interpretation refers to the significant vocal sound that signifies something by itself. Hence, all parts of speech with exception to conjunctions and prepositions are signs.1 Aquinas summarizes, “a name signifies the substance of a thing and then that the verb signifies action or passion proceeding from a thing.”2 The simple point is that interpretations signify existing things.
Hence, Aristotle implies that our speech maps to real things (technically, also imaginative transformations of real things), and that real things may be internalized in humans well enough to articulate them through language.
The natural question that emerges is how humans internalize real things so as to be able to articulate them. Aristotle answers this in On the Soul. Very briefly, Aristotle posits that all knowledge is principled in sensory experience.3 It is through contact with sensible objects that we eventually attain the concept of them, and those concepts become the foundation from which language emerges. After contact with the five senses, the disparate sensory inputs must be unified, then our imaginative power creates an image that represents the basic elements (e.g., color, shape, number) of that sensory input.4 Moreover, as the Aristotelian-Thomistic interpreters develop more explicitly, the cogitative power creates an image of the perception embedded in the unified sensory input, and this perception is comprised of aspectual, actional, affectional percepts.5 Simplifying the jargon, individual things and their features, potential actions, and the emotional impact of the thing are all represented via an image.
What is quite incredible is that flash forward to the present, and Peter Gärdenfors has presented a conceptual spaces framework that posits that the cognitive phenomena may be best modeled geometrically rather than merely symbolically (classical logic) or associatively (statistical probabilities) since geometric representations allow for similarities and differences between things to be “visualized” and measured.6 I’ve drafted an extensive paper connecting Gärdenfors’ framework with that of Aristotle for those interested in a deeper dive on this subject.
But Gärdenfors’ framework, for all its explanatory power, only accounts for the cogitative layer, meaning the geometric categorization of perceptual images. Aristotle goes further. Fleshing it out, the perceptual images of things generated by the cogitative power does create a categorization of things that allow for basic inferences, as Gärdenfors articulates. However, there is an additional step “above,” for there is an immaterial intellect that is able to strip the material conditions of the perceptual image to possess an universal, absolute conception of a thing (i.e., a concept). From this concept, things can be compared and contrasted necessarily, logically rather than from material geometric comparisons studied by cognitive science. There is discussion as to whether the signs in Aristotle’s interpretations map to the perceptual images or the immaterial concepts. The answer is likely both.7
In either case, the fascinating connection with LLM architecture is that geometric representations of words, and transformations upon them, are the means to developing semantic understanding. This parallels Aristotle’s theory that real things take on a geometric representation as an image in the brain and that words, in some manner, are derived from these geometric representations. Hence, while Natural Language Autoencoders are a breakthrough in interpretability tooling, they are not a wholly original idea. Aristotle’s theory, begun long ago, would predict such a translation to be possible. While I’ve enjoyed commenting on the Aristotelian parallels as Anthropic releases new research, my interest and hope is that Aristotle’s framework could be applied to inform interpretability tooling (among other things at the center of AI research) instead of waiting until it is accidentally discovered. In either case, I’m very much enjoying the ride.
Aquinas, Commentary on On Interpretation, Introduction.
Ibid., Bk. 1, Lesson 4.
Aristotle, On the Soul, Bk. 2, Ch. 6.
Ibid., Bk. 3, Ch. 2.
De Haan, Perception and the Vis Cogitativa, pp. 416, 429.
See Gärdenfors, Conceptual Spaces as a Framework for Knowledge Representation (2004); Reasoning with Concepts: A Unifying Framework (2023).
De Haan, The Interaction of Noetic and Psychosomatic Operations in a Thomist Hylomorphic Anthropology, p. 20.





