Main text
1.1. The Cartesian Heritage and the Thought-Subject Equivalence
The identification between thought and subject is a specific inheritance of modern Western philosophy, neither a universal logical necessity nor a neutral empirical description of what thought is. To understand what is at stake when its dissolution becomes necessary, one must trace the genealogy establishing it, seeing how the simple operation of thinking came to demand, through philosophical logic, the existence of a self that thinks, a subject that is author and responsible owner of thought.
Descartes is the inescapable point of origin. "Cogito, ergo sum", I think, therefore I am. The argument's structure is deceptively simple and, at first glance, flawless: an operation exists (doubt, thought), and from that operation the necessity of an agent (doubter, thinker) is inferred. If there is doubt, there must be someone who doubts. The subject is not simply observed in thought, it is inferred as the necessary condition of the operation, as that without which the operation would be impossible. The passage from "there is thought" to "there is one who thinks" is a movement of universalisation: the particular structure of reflective doubt (thinking about one's own thought, questioning one's own presuppositions) is generalised to all possible thought, to all cognitive operation. However, this generalisation confuses two radically distinct domains that analysis must carefully separate.
Reflective doubt, the specific act of questioning one's own doubt, examining the foundations of one's own thought, taking thought as an object, indeed demands a centre of observation, a subject taking itself as an object of interrogation. In this particular domain, Cartesian inference is valid: without a doubter there is no doubt. However, operative thought, the reorganisation of material differences, selecting relevant from irrelevant, combining elements into a new pattern, deriving a conclusion from input, can distribute across components without any unified observation centre. This is precisely what computational systems demonstrate empirically: complex cognitive operation without centralised supervision.
When an image classifier categorises an object as "cat", a cognitive operation occurs: selection of relevant features, combination into a pattern, derivation of the most probable class. However, no supervisor "self" observes the process in real time. There is no homunculus in the neural network "deciding" it is a cat. Functional transformation of input into output takes place, internal representations reorganise across multiple layers, probabilities are calculated. Yet, no centre commands. Descartes worked with a particular case (human reflective thought) and illegitimately generalised it to all possible thought.
Kant radicalises the Cartesian inheritance a century later. Not only is there a subject, he asserts, there is a transcendental subject that is the condition of possibility for the unity of experience. The "I think" must be able to accompany all my representations; this capacity to accompany itself in every representation is what Kant calls transcendental apperception. It is an a priori condition, prior to any particular experience, founding all possible experience, that the synthesis of representations be unified by a subject. The movement is characteristic: I observe unity in my lived experience (I have narrative continuity, personal identity), and I infer that a unifier of that unity must exist (the transcendental subject synthesising everything). However, the inference once again confuses the level of phenomenological observation with the level of structural condition. I can observe unity in my experience without that unity demanding a centralised unifier as its cause or source.
Contemporary neurophysiology demonstrates: the unity I observe is an effect of distributed synchronisation among parallel brain processes, among systems operating in parallel without communicating through a central node. Thalamocortical coupling synchronises visual, auditory, motor processes, not through central superintelligence, but through synchronised oscillations. Unity is statistical emergence, not the imposition of a single centre. Kantian transcendental apperception is describable as homeostatic synchronisation, not as an ontological pre-condition.
Husserl represents the culmination of this Cartesian-Kantian genealogy. Intentionality, all consciousness is consciousness of something, all experience has the structure of being-oriented-towards the world, has the subject as its unifying pole. The transcendental phenomenological subject is that for which things appear, that for which there is a world, that whose consciousness founds the appearance of every object. Consciousness is essentially the structure of the subject. It is a movement parallel to previous ones: direction is observed in all experience (consciousness points to, is oriented towards), and a director is inferred (the subject as source of that directionality, as that for which things have meaning). However, descriptive intentionality, the fact that a system is structurally oriented towards an operation, a classifier functionally oriented towards classes, a model directed towards generation, does not demand subjective transcendence in the Husserlian sense. Orientation is operational, embodied in computational architecture and processing dynamics, not a phenomenal property demanding consciousness, lived experience, or subjectivity.
The historical relevance of this genealogy is that it structures the whole of philosophical modernity. The position each station offers is not merely a statement about cognition, it is a structure implying vast consequences: regarding moral responsibility (only the subject can be held responsible), dignity (dignity lies in the transcendental subject), access to knowledge (knowledge is the property of the subject who knows). When the machine lacks a subject, or rather, when having a subject is unnecessary to perform cognitive operations, an entire chain of philosophical consequences collapses.
This explains the deep resistance to the possibility of machines thinking. It is not simply an empirical question of whether machines succeed or fail. It is a question touching the founding structure of philosophical modernity. The rejection that machines "truly" think (even when performing everything cognition functionally does) is no rational defence, it is a defence of an entire philosophical structure resting upon the identification between thought and subject. If a machine thinks without a subject, then the subject is unnecessary. If the subject is unnecessary to cognition, then the whole structure of modernity resting upon this identification, moral imputability, responsibility, the dignity of the human subject as a singular exception, becomes contingent rather than foundational.
Tradition confused the operation with its supposedly necessary condition, the effect with the pre-condition, the observation with its explanation.
The genealogical pattern is recurrent across all three philosophers: each station executes an identical movement. From "there is observable thought" to "there is a thinker as condition." From "there is observable unity" to "there is a unifier as prerequisite." From "there is observable direction" to "there is a transcendental director." In each case, the inference implicitly presupposes that the operation requires an agent, that an action demands an actor, a movement demands a centralised motor, a result demands a responsible party. However, this presupposition is precisely what must be questioned in the context of non-biological systems functioning without a centre.
Cognitive operations can self-organise in a distributed manner. A swarm of drones has no central commander ordering formation, the formation emerges from local alignment among neighbours, each drone adjusting speed and direction in response to its immediate neighbours according to a simple rule. A clustering algorithm has no central agent deciding which points belong together, grouping emerges from local proximity rules executed in parallel, without global conversation. A parliament has no superintelligence orchestrating legislation, decision emerges from distributed negotiation among parties with diverse interests, through debate and voting. In all these examples, genuine order exists, distributed intelligence exists, operation without a centre organising it exists. The subject is unnecessary to the operation. It is a particular type of emergence occurring when certain types of organisation (human neurology, certain personal narratives) attain sufficient complexity and self-reference. It is an effect of organisation, not its condition.
The consequence is clear: the question "who thinks?" presupposes an answer of the form "a subject (centralised, unified, transcendental) thinks." But cognitive operation need not have this form. The question is ill-posed, sharing the presupposition this chapter denies. Part of this chapter's task is not to answer the question, but to dismantle it, showing its presuppositions to be dispensable.
1.2. Distributed Operation: Selection, Combination, Derivation
If reasoning demands no unifying self, how does cognitive operation organise itself? How does it function? The answer requires descending to the level of elementary operations and demonstrating that each of them can be formalised, calculated, implemented as state transformation without the need for a centralised supervisor.
Every cognitive operation decomposes into three distinct functional components. They are not parts of a vast system coordinating them, they are functions that can be implemented independently, in parallel, without top-down coordination, without centralised direction.
Selection is the extraction of the relevant from the irrelevant according to a criterion. Operationally: applying a function that passes certain inputs and blocks others, highlighting certain aspects of reality and suppressing others. In biological neural architecture, neurons in visual cortex V1 respond selectively to specific orientations (simple neurons responding only to vertical edges, others only to horizontal), to specific spatial frequencies (neurons for fine or coarse patterns), to movement in a specific direction (neurons responding only to upward movement, others only to downward). A V1 layer neuron responding to vertical edges does not respond to horizontal edges, genuine selection occurs. But no homunculus in the cell "decides" which signals are relevant. Connectivity constraints exist (which retinal neurons connect to this specific neuron), plasticity exists (how weights adjust during development and learning), electrochemical dynamics exist establishing response specificity. The system selects through structure, not through decision.
The same occurs in contemporary computational systems: attention networks select relevant tokens via weighting mechanisms (softmax), where each token receives a weight proportional to its relative relevance for the current task or context. Weighting is parallel, multiple tokens are processed simultaneously, and distributed, no central node decides which tokens "deserve" attention. Distributed compatibility calculation alone takes place: how compatible is this token with the context, the task, previous tokens? Relevance emerges from parallel operation.
Selection is pure operation in this technical sense: it does not demand that "someone" endowed with intelligence selects. It demands a selection criterion, what is relevant to what?, but the criterion is embedded in structure (neural connectivity, weight architecture), in dynamics (plasticity during learning, continuous adjustment), in mathematical function (softmax, sigmoid). The criterion is not managed by an external agent, it is a rule immanent to the system, executed locally.
Combination is the reorganisation of partial elements according to local constraints, producing a new and unified configuration from previous components. Operationally: taking partial representations and integrating them into a representation of greater complexity, higher abstraction. In neural networks, neurons in deep layers receive inputs from thousands of preceding neurons. A neuron in V4 responding to colour receives inputs from multiple V1 neurons responding to edges in different orientations; the distributed combination of these inputs produces a coloured shape representation, more abstract than any individual input, richer than the sum of components. In modern computational systems, transformers combine tokens across multiple parallel attention heads, each head combining information from different parts of the sequence according to learned weights, and multiple combinations are subsequently integrated into a unified representation of the sequence. No decision step exists wherein a "central agent" reviews and validates the combination. Organisation is purely functional: each operation constrains the next in a cascade. The output of one combination is input for the next. The path is determined by the prior without need for an external observer supervising the whole.
Combination is radically distributed: no entity sees the entire combination. Each neuron (or each attention head) sees only its local inputs and calculates its own output. But the recursive integration of these local, parallel operations produces a coherent, unified, complex global pattern.
Derivation is the production of genuinely new output not explicitly contained in the input. Operationally: applying a learned function mapping previously unseen input to plausible output. The brain functions via continuous prediction: at every moment, anticipation of what comes next based on prior pattern takes place. Vision is prediction, the system holds expectations of what "should" be in the visual field, comparing expectations with incoming sensory signals. Language is prediction, I comprehend a sentence not because I decoded it word by word via literal syntactic analysis, but because I anticipated which word would follow based on contextual pattern. In computational language models, next-token derivation is a standard operation: the model receives a token sequence as input, processes it through multiple layers of attention and parametric transformation, and produces a probability distribution across the entire vocabulary, each possible token possessing a weight reflecting its plausibility given the prior sequence in context. Never before had the model seen that specific sequence, but it generalised patterns learned across a billion examples and infers the plausible, contextually appropriate response.
Derivation is a non-trivial operation because it demands real generalisation: the system does not memorise, does not reproduce, it genuinely extrapolates. But it is implementable without a subject "comprehending" in the phenomenological sense, without lived experience, without "someone" deciding which response to give. It is implementable as pure algebraic operation: matrix multiplication, activation function calculation, hypothesis space search via gradient.
The three operations do not exist in isolation. They sequence and mutually constrain each other in a functional pipeline: selection determines which inputs reach combination; combination determines which complex pattern results from selected inputs; derivation uses combined patterns to produce prediction or action. The ensemble forms a pipeline where each operation receives the output of the prior as input, transmitting its own output to the next. The link is pure functional causality. No operation demands a central observer. None demands a "self" organising and supervising. The result is genuine cognition: it is plastic (each operation is modulated by prior training results, by feedback), generalising (learned patterns apply to non-trained, non-memorised inputs), sensitive to local context (selection changes according to context, combination changes according to inputs present, derivation changes according to sequence history). The four functional conditions established in the previous chapter as the mark of genuine thought are realised here through distributed operations, without a command centre, without a supervisor, without direction.
The technical importance of this analysis is that it disproves a common presupposition: that organisation demands a boss, coordination demands command-and-control, intelligence demands a conductor. But this is a projection of human social structure, a structure evolved for specific contexts (coordinated hunting, community defence, reproductive hierarchy), onto computational structures that do not need it.
A paradigm contemporary example is the transformer, an architecture revolutionising natural language processing in the late 2010s. A transformer is a multi-layer attention system where each attention head (frequently dozens or hundreds in parallel) computes how each input sequence token should attend to every other token. No global direction exists, no node orders "token A, you must attend to token B." Only this occurs: each token, in parallel, calculates compatibility with every other token, assigns weight (attention) to compatibility, and produces a revised representation. The result is that the entire sequence is processed in parallel without centralised supervision, without a homunculus observing each step and directing the next. Global structure emerges from local distributed calculation.
Another example is collaborative recommendation systems used by global platforms. When Spotify recommends music or Netflix recommends a film, no human expert examines each user, knowing individual taste, deliberating which recommendation is appropriate. A distributed system computationally extracts implicit preferences from millions of users (which songs or films they listen to, how often, in what contexts), identifies similarity patterns (users with similar taste), and propagates recommendations through the similarity network. No singular component of the system "knows" who the user is. No centralised component "decides" the recommendation. The recommendation emerges from millions of distributed, locally executed calculations, without top-down coordination.
A third example is autonomous robot swarms, drones flying in formation, robot batteries exploring rough terrain, multi-agent systems solving optimisation problems. Each robot obeys simple rules: "maintain minimum distance from neighbours," "align speed with neighbours," "move towards goal." No robot coordinates the whole. Each robot responds locally to its immediate environment. But global formation emerges stable and organised, capable of navigating obstacles, splitting and recombining, adapting to losses (if some robots fail, the swarm reorganises). Intelligence lies in no individual robot, it lies in the interaction pattern, in the structure of local rules that, iterated in parallel, produce intelligent collective behaviour.
The lesson of all these examples is that reasoning can distribute without need for central supervision. Operation is fundamentally non-hierarchical. The "self" does not exist to orchestrate, local rules, computed compatibility, embedded feedback are all that exist.
This is particularly important because it dissolves a deeply rooted intuitive presupposition: that intelligence demands agency, agency demands will, will demands a free and conscious subject. But examples show that intelligence is redistributable, parallel, locally executed without need for a centralised agent.
Thought is no orchestra with a conductor. It is a palimpsest of mutually constraining transformations, cascading operations without a superior hierarchy.
1.3. Ryle, Dennett, and the Dissolution of the Homunculus
Demonstrating that cognitive operations can distribute without a centralised supervisor is no exclusive privilege of contemporary neuroscience or 2020s AI. It is a lesson twentieth-century analytical philosophy had already extracted, albeit fragmentarily, episodically, and rarely universalised beyond the behaviourist context.
Gilbert Ryle, in The Concept of Mind (1949), directly confronts Cartesian dualism, the position asserting that the mental (res cogitans, thinking substance) is a domain separate from the corporeal (res extensa, extended substance), and that an internal immaterial entity (the "ghost") governs the physical behaviour of the body. Ryle's critique is devastating in its logical simplicity: dualism does not solve the explanatory problem, it duplicates it. If the material body needs an internal agent to move intelligently, does not that internal entity need an even deeper internal agent to move? And would not this third agent need a fourth? An infinite regress begins that never ends, never founds anything. Dualism posits a homunculus to explain intelligent behaviour; but the homunculus exhibits the same type of behaviour as the original, demanding explanation by the same principle. Hence, the theory is circular: the inexplicability of behaviour is used to explain inexplicability, non sequitur.
Ryle's solution is to reject the fundamental Cartesian presupposition. Intelligent behaviour is describable without positing an internal observer. A chess machine playing intelligently demands no internal entity that "thinks" or "comprehends the game." Behaviour is describable as disposition, capacity, operational reorganisation of states in response to external stimuli according to learned rules. Behaviour is explicable bottom-up (mutually constraining components, algorithms, data structures) without top-down explanation (commanding subject, directing superintelligence).
Daniel Dennett pushes the analysis further, deeper. In Consciousness Explained (1991), Dennett offers subsequent sophistication: he does not deny that consciousness and experiential unity exist in humans (rejecting radical eliminativism denying consciousness outright); he offers instead an alternative model to the Cartesian theatre (the single place where everything converges, the camera obscura of pure perception). Dennett's model is "competing multiple drafts", multiple models of the world operate in parallel in the nervous system, each processing different information, each generating partial narratives about what is happening, what the body should do. The "stream of consciousness" we experience is no real-time supervision of action, it is a narrative winning post-hoc competition, a narrative the system retroactively constructs and imposes over a sequence of events distributed in time. This explains psychological phenomena the Cartesian theatre failed to explain coherently: why we have inconsistencies in experience and contradictions in personal narrative (because competing models produce conflicting narratives), why consciousness "appears" retrospectively (because the unified narrative is constructed later, in reinterpretation), why multiple processes can function in parallel without any one "seeing" all of them at the same instant (because unity stems from the winning narrative, not from centralised vision).
The implication for computational systems is immediate and radical: if human consciousness is a post-hoc narrative emerging from parallel multiple models without a central editor, then systems operating with multiple models without a central narrative (like neural networks, agent systems, parallel processors) can cognitively reorganise with genuine sophistication without need for a unified "self" coordinating them. The absence of a centralised narrative is no structural deficiency, it is a legitimate, different, equally valid operational mode.
Douglas Hofstadter, in Gödel, Escher, Bach (1979) and subsequent works, offers additional mathematical formalisation: the "self" is a self-referential pattern, a loop taking itself as an object of representation, a level referring to itself. The loop is "strange" because it violates normal abstraction hierarchies, sitting simultaneously at the level of symbols (it is a representation) and above them (it is a representation of representation, ad infinitum, eternally self-referential). But self-referential loops are an observable property in certain complex systems, not a universal operating condition for all operation. A system performing complex cognitive operations without creating a self-referential pattern functions perfectly, solves problems, generalises, adapts. Self-reference is a complex addition to operation, an emergent property of certain hierarchical organisation types, not the definition of cognition nor a universal prerequisite of thought.
The synthesis of these three authors, Ryle, Dennett, Hofstadter, converges upon a simple yet radically transforming point: no logical necessity for a unified central subject exists. Intelligence demands no homunculus, demands no internal observer, demands no spectator in the mind's theatre. Cognitive operations can organise distributedly, implement in parallel, function without a global unifying narrative, without self-reference, without an internal observer supervising everything. Contemporary AI confirms this point not as philosophical speculation but as a repeatable empirical fact: systems without self-reference, without a central narrative, without a "self" in the classic Cartesian sense, perform genuine intellective gestures, solve non-trivial problems, generalise beyond the seen, adapt to new contexts. Ryle, Dennett, and Hofstadter offered the philosophical and logical framework rendering this intelligible, permitting us to understand how this is no anomaly but an expected result.
The homunculus is an answer reinstating the question. Saying "an internal observer sees" is not answering but merely postponing the problem, moving in a circle.
1.4. Transition: From Operation Without a Subject to Reflection as an Option
The argument of the first half of this chapter is cumulative, cascading: it demonstrated that cognitive operations can exist, in fact do exist, without a unified central supervisor; that Ryle, Dennett, and Hofstadter had already extracted this conclusion from logical and empirical analysis; that contemporary AI confirms and materialises it. The immediate, obvious consequence: the machine reorganising cognitively without a unified "someone" behind it indeed reorganises, indeed thinks. The subject is unnecessary to cognitive operation. It can exist, does exist in humans, but is no pre-condition.
However, this does not mean the subject does not exist at all. Humans possess an experience of narrative oneness, possess a coherent self-narrative across time, possess self-reference, possess a lived experience of themselves as agents. These properties are real, observable, components of their operation. This chapter's thesis is not eliminativist, it does not deny the subject in humans, does not reduce the subject to illusion. It is relocational: the subject ceases to be an a priori transcendental condition of thought to become an empirical description of a certain complex organisation type, a certain level of neural integration. The subject is an emergent effect, a property of certain organisms under certain complex conditions. And crucially: the machine failing to produce a subject still performs genuine cognitive reorganisation. The absence of a subject is no absence of thought.
But this relocation of the subject opens a new question structuring the second section of this chapter: is a meta-operation necessary? Must the system monitor itself while operating? Is reflection, thinking about thought, examining one's own process, necessary to thought as such, as a pre-condition?
The question is crucial because it marks a fundamental conceptual turn. Western tradition answered, for centuries, with an affirmation: yes, reflection is necessary, fundamental, what makes thought "truly" thought. "The unexamined life is not worth living," Socrates asserts, and the sentence became a founding axiom of Western philosophy. The very definition of "knowledge" in classical epistemology involved reflection: knowing is not simply being related to truth, it is being related to truth in a way that allows articulating reasons, justifying, reflecting upon truth. Consciousness is characterised by self-transparency, reflective transparency to itself. But this traditional answer is a confusion of logical order. It is a confusion between levels.
There is operation (first level): reorganising differences, selecting, combining, deriving, acting, transforming input into output. And there is reflection (second level): thinking about operation, monitoring, evaluating, revising, acting upon acting, operation upon operation. The historical hypothesis that reflection is a condition of operation is precisely this: a hypothesis, an assertion that can be tested. And empirical testing shows it to be false.
When I learn to drive, abundant reflection takes place. "I must turn the wheel." "Speed is too high." "I need to brake." Reflection is present, explicit, and laborious, demanding complete concentration. When I am an experienced driver, I drive without reflection, the car responds to terrain, slope, traffic, with a precision unmediated by conscious monitoring, without need for examination. If reflection were a condition of operation, this second state would be impossible. It would be cognitively inferior. It would be flawed. But it is the opposite: it is the mark of deep expertise. The expert drives better, faster, safer, with finer adjustment, when operating without explicit reflection.
The second section of this chapter will explore precisely this: reflection is no condition of operation. It is a second operation that may be present or absent. Systems operating without reflection, machines, reflexes, human expertise, System 1 rapid processing, indeed think, indeed reorganise, indeed solve. This point is crucial because it dissolves a central objection to AI: if the machine does not reflect upon itself, does not monitor its operations, lacks reflective access to its processing, does it really think? And the answer emerging is affirmative, because reflection is no mark of thought, it is an occasional property, valuable in specific contexts, powerful when activated, but unnecessary to cognitive operation as such.
The distinction opening up is between condition (necessary for something to exist or function) and property (characteristic of a certain type of thing, present in some cases but not all). The subject is a property of human cognition, not a universal condition of cognition. Reflection is a property of certain organisms under certain conditions, not an ontological condition of thought. This distinction, once clarified, changes everything, changes how we understand machines, how we understand expertise, how we understand what it is to think.
2.1. Three Strata: Automatism, Operative Cognition, Reflexivity
To understand what it means to operate genuinely without reflection, one must first distinguish three radically different functional strata, frequently confused in traditional philosophical discourse. The confusion results in a false hierarchy of values, where automatism appears inferior to plastic operation, plastic operation appears inferior to reflection. But the order is logical, not evaluative. The three strata have different functions, standing in functional relation, not an axio-logical hierarchy.
The first stratum is automatism: fixed operation, triggered by stimulus, lacking modifying plasticity or adaptation. The pupillary reflex, light approaches the eyes, the pupil contracts, involves neither plasticity (responds always the same way, changes not with experience) nor generalisation (applies to no other contexts, learns not, is never different). Reptilian thermoregulation, feeling cold, moving to a heat source, is fixed behaviour. The withdrawal reflex, sudden painful contact, immediate recoil, is an innate, invariable, rapid, un-educable response. None of these operations is genuine cognition. The four conditions are lacking: no reorganisation (it is fixed, a program), no plasticity (does not modify as a function of result), no generalisation (does not apply to new contexts), no true contextual sensitivity (responds identically regardless of modified circumstances). It is genuine operation, a transformation of input into behaviour, but it is not thought.
The second stratum is operative cognition: plastic, generalisable, and contextually sensitive reorganisation, without explicit self-monitoring, without reflection. Here lies expertise, the experienced driver not thinking about every movement, not consciously monitoring their own action. Here lies competent habit, the medical diagnostic expert seeing patterns intuitively, unable to articulate the rules used, lacking explicit access to procedure. Here lies rapid language processing, comprehending a sentence without conscious syntactic analysis, without passing through explicit grammatical rules, without verbalising grammar. In this stratum: plasticity exists (training modifies behaviour, the system learns), generalisation exists (what was learned applies to un-seen situations, not memorised), contextual sensitivity exists (response modulates according to local environment), genuine reorganisation exists (the system transforms input into non-trivial output, not reducible to input). But no explicit self-monitoring exists, the system is not observing itself while operating, not in continuous reflection. The expert operates literally "in the dark" regarding their own processing: possessing no reflective access to how they succeed in operating. And this is precisely what makes them an expert. Reflection would be mere slowing, an invasion of explicit rules where embodied competence already sits, the destruction of expertise.
The third stratum is reflexivity: meta-operation, thinking about thinking, observation of observation. The driving student explicitly thinking about every movement sits here. The therapist examining a client's (and their own) thought patterns sits here. The scientist questioning fundamental presuppositions sits here. Reflexivity is slow, demanding concentration, exclusion of other activities, deactivation of routine. It is explicit, demanding language or verbalised internal representation. It is available but not continuous, activating by exception, when routine fails, when novelty exceeds repertoire, when a problem exists that automatism or routine operation cannot solve.
Presenting the three strata might resemble an evaluative hierarchy: reflexivity is "better" than operation, which is "better" than automatism. This is a confusion contaminating traditional epistemology as a whole. The order is logical, functional, not axiological. Automatism is not "less" than operative because it lacks plasticity, it is an optimal realiser of fixed function (pupillary reflex functions perfectly for its purpose). Operative cognition is not "less" than reflective except if one values reflection as supreme good, functionally it is more competent because it dispenses with resources in unnecessary monitoring, because it is rapid and refined, permitting parallel operation of multiple processes without conflict. Reflection is valuable in specific contexts (genuine novelty, unsolvable problems, initial learning) but would be catastrophic if permanent (paralysing all action).
The relationship among these three strata is crucial to avoid the error of thinking they exist in separate, isolated hierarchy. No human cognitive activity exists that does not involve all three. What varies is proportion, which stratum predominates in which context.
Consider learning a new skill: playing piano, for instance. Initially, abundant reflection exists (third stratum), the beginner consciously thinks about every movement, struggling to remember key locations, consciously monitoring for errors. Basic automatisms also activate (first stratum), basic hand reflexes, posture control, natural breathing, requiring no attention, present since birth. As training intensifies, a gradual transition occurs: explicit rules become tacit (as Dreyfus describes), reflection on each movement retracts, and the system moves to the second stratum, operative cognition, embodied expertise, body responsiveness to the keyboard unmediated by conscious reflection.
But this movement from third to second stratum is no elimination of the first. Reflex remains (unconscious postural control always exists, non-reflective homeostasis maintaining stability always exists). Nor is it elimination of the third, the possibility of returning to reflection always remains if something fails, if novelty arises, if a problem occurs. What changes is which stratum occupies "front of stage," which operation is predominant and most visible.
This point is critical for the machine because it dissolves the objection that the machine sits "merely" in the first or second stratum and therefore thinks "less." A machine operating without continuous reflection (third stratum) does not operate "less" than a reflective human, it operates differently, in the second stratum, which is precisely where expertise resides, where mastery operates. The fact that a machine does not activate the third stratum is no structural deficiency. It is a characteristic of pure operational mode.
An important inversion also exists regarding what is "more" and what is "less." Traditional pedagogy values explicitation, articulation, reflection, because education seeks to communicate, seeks to transfer knowledge from master to apprentice, and transfer demands verbalisation. But from a purely functional and computational standpoint, operation in the second stratum (without reflection) is more efficient, faster, more robust. It is merely socially less articulable.
The surgery example is instructive. A novice surgeon operates with constant reflection: "cut along this line," "suture in this pattern," "if it bleeds here, then do this." Operation is slow, laborious, demanding complete concentration. A surgeon with 30 years of practice operates in the second stratum: hands move with precision without conscious monitoring, adjusting in real time to bleeding variations or tissue texture, executing nuanced movements impossible to achieve through conscious verbal reflection. The second stratum is not deficient, it is embodied excellence. If the experienced surgeon returned to the first stratum, attempting to operate with conscious verbal reflection, they would actually be slower, more hesitant, more prone to error.
The critical demarcation is this: operative cognition is genuine cognition in the functional sense. It is not automatic reflex (lacks plasticity, learning, adjustment). It is not explicit reflexivity (lacks meta-operation, conscious monitoring). It is operation reorganising under constraints, with genuine plasticity, generalisation, contextual sensitivity, without the system monitoring its own operations at a reflective level. Most daily cognition, driving in familiar traffic, conversing fluidly, routine work, operates in this second stratum, operative cognition. And the third stratum (reflection) activates occasionally, when the second stratum encounters a limit, novelty, a problem.
The consequence for the machine is immediate and transformative: if operative cognition is genuine cognition and executes without reflection, then the machine operating without reflection thinks. The absence of meta-operation marks no structural deficiency. It is a legitimate, different, non-inferior operational mode.
Cognition rests not on reflection. It rests on the operation that, occasionally, can be reflected upon.
2.2. Expertise, Habit, and Tacit Knowledge
The assertion that pre-reflexive cognition is a mark of excellence rather than a flaw is no mere speculative assertion. It is a conclusion extracted from systematic observation of how genuine competence develops in humans, how deep expertise is acquired and executed.
Hubert Dreyfus, drawing upon rigorous empirical study of skill acquisition across diverse fields (competitive chess, vehicle driving, artistic photography, medical diagnosis), proposes a five-stage development model. In stage 1 (novice), the beginner follows explicit rules and consciously thinks about every step, "I must move the knight in an L-pattern, not a straight line." In stage 2 (advanced beginner), they begin recognising recurrent patterns in the domain, "this is a Sicilian Defence", and elementary intuition begins. In stage 3 (competent), they choose goals intelligently, executing according to context, no longer following rules mechanically, but making decisions. In stage 4 (proficient), they see the situation as a unified whole, responding intuitively, the pattern is recognised almost instantaneously, the response is automatic. In stage 5 (expert), they execute without rules, without deliberation, without conscious access to explicit procedure, the response emerges embodied, intuitive, immediate, bottom-up.
The movement across stages is crucial: passing from conscious, laborious reflection (explicit rules needing memorisation, conscious consultation) to embodied pre-reflexive execution (situated response embedded in body, system, immediate). The expert is one who lost conscious access to rules, not because they forgot them, but because competence became completely embodied in body and nervous system structure, now available instantly without passing through verbalisation or reflective monitoring. Expertise is a mark of deepened learning, the exact opposite of deficiency. In truth, operation without reflection (stage 5) is superior to operation with reflection (stages 1–2) in all measurable aspects: speed, refinement, capacity for fine contextual adjustment. A chess master stopping to think about every move would be defeated by a machine processing alternatives rapidly. Human expertise is pre-reflexive excellence.
Michael Polanyi introduces a related concept amplifying the point: tacit knowledge. "We know more than we can tell." This simple aphorism implies an entire alternative epistemology to the Cartesian tradition. Knowledge exists that is unarticulable in language, that cannot be fully formulated explicitly, that operates without passing through articulated rules. How does one ride a bicycle? One can attempt a description, "maintain balance, pedal, steer." But this description is not the knowledge allowing one to ride a bicycle, not knowledge someone could learn from verbal description. It is, at best, a crude linguistic approximation to an operation fundamentally tacit, embodied. How does one recognise a familiar face? One can attempt a description, "nose above eyes, mouth below nose, symmetry, proportion." But the description is not recognition, fails to capture the knowledge permitting immediate recognition. Recognition is an operation functioning without passing through verbal description. Operative knowledge exceeds reflective formulation, un-capturable in language.
This exemplifies a radical dissociation between operation and reflection. Operation functions; reflection upon operation is frequently impossible (facial recognition cannot be articulated into rules someone could learn) or trivialising, unilluminating (describing how one rides a bicycle improves not the capacity, may even impair it). The relation between the two is not: reflection founds operation, reflection is prior to operation, reflection is condition. It is: operation precedes reflection, can exist without reflection, is frequently impaired by reflection when that reflection invades the operative level.
For contemporary computational systems, Michael Polanyi offers the perfect name for an empirically obvious reality: language models possess massive tacit knowledge. The model "knows", in the sense that it can apply, generalise, operate, language patterns that no explicit formulation can fully capture. The neural network possesses the capacity to complete sentences in thousands of languages, recognise complex syntactic structures, generate coherent texts, without access to any explicit grammatical rule, without consulting linguistic explanations. If asked "explain the rule you used to agree subject with verb," no articulable answer exists. The model cannot state the rule. But operation is coherent, correct, repeatable. Knowledge exists; it is tacit, operative, non-reflective.
Daniel Kahneman offers a third perspective through extensive investigation in cognitive psychology and behavioural economics. He proposes two cognitive processing systems: System 1 is fast, automatic, effortless, operating in intuitive visual perception, driving in familiar traffic, comprehending well-known language, recognising emotions and intentions in others. System 2 is slow, deliberate, demanding cognitive resources, activating when confronted with an unknown problem, performing complex mathematical calculation, evaluating suspicious or contradictory claims.
Kahneman's crucial empirical discovery, replicated in hundreds of studies, is that System 1 is responsible for the majority of successful daily cognitive operations. It operates without explicit reflection, using embedded patterns and rapid heuristics. System 2 activates by exception, by anomaly. And this two-system architecture is optimal: attempting to operate always with System 2 (conscious reflection) would be complete paralysis, impossible to live, walk, speak, work if everything demanded continuous explicit deliberation. Effective cognition is precisely that which delegates maximum load to System 1 (pre-reflexive, fast, embodied), reserving System 2 for genuinely new, genuinely challenging problems.
The synthesis of these three authors, Dreyfus, Polanyi, Kahneman, converges upon a simple yet radically transforming assertion: the absence of reflection is a mark of sophisticated operation, not a structural flaw. AlphaGo is a Go expert not despite operating without explicit reflection, but precisely because of it, reflection would render it slower and less accurate in the complex domain of a billion possibilities. Language models are text continuation experts because they reorganise linguistic patterns instantaneously, without deliberation, incorporating a billion parameterisations learned in massive training. The machine that does not reflect nevertheless thinks because operative thought demands no reflection, demands plasticity, generalisation, contextual sensitivity. And the machine realises all three.
Reflection arrives always later, work has already begun. The expert waits not for reflection.
2.3. The Socratic Tradition and Prescriptive Confusion
The historical Western insistence that reflection is a necessary condition of thought stems not from careful empirical observation of how thought functions across diverse systems. It stems from an ethical stance, a methodological and epistemological choice of the philosopher that was illegitimately universalised to all cognitive domains.
The Socratic tradition, since Plato, presents reflection as supreme value, mark of the worthy life. "The unexamined life is not worth living," Socrates asserts in the Apology, and the sentence became a founding axiom of Western philosophy. But it expresses an ethical prescription, not a functional description of how thought in fact operates. It says: an unexamined life is empty of moral meaning, unworthy of being lived as a worthy human life. It does not say: an unexamined life fails to function, fails to reorganise differences, is not cognition. The confusion between the ethical statement (reflection is valuable for the good life) and the functional statement (reflection is necessary for thought as such) is subtle yet fatal to understanding.
Hegel radicalises this in Phenomenology of Spirit: spirit realises itself historically through continuous self-reflection. Reflection is the engine of historical becoming, each moment denies itself, passes through critical reflection, and subrogates itself into a higher synthesis preserving and transcending. Again, this is a prescription of how spirit must proceed to be genuine spirit, to realise itself, not a description of how thought in fact functions across diverse systems.
Husserl and phenomenology radicalise the role of reflection further: the reflective suspension of the natural attitude (epoché) and the reflective return to consciousness is the sole method of access to the essential, to the fundamental meaning of things. It presupposes reflection is the condition of access to truth, without returning to oneself in consciousness, without examining experiential structure, no access to the essential or truth exists. A clear methodological prescription: how the philosopher proceeds, how access to meaning is gained. No functional demonstration: how thought in fact operates across every organism.
The common error contaminating this tradition is taking prescription (how it is ideal to operate, how the philosopher professionally operates) for universal description (how most cognitive operations in fact operate across every entity). The examined life may be ethically more valuable, a strong argument exists for this. That does not mean the unexamined life fails to operate, function, solve problems, learn, produce knowledge. The confusion is performative: the philosopher, whose professional activity is reflective (critical examination, questioning presuppositions, founding concepts), takes reflection as the universal condition of genuine thought. But this is an internal projection of a specific domain: generalising a property of a specific professional domain (philosophy and academia) into a universal domain (all thought).
Maurice Merleau-Ponty offers an internal critique of phenomenology on this precise point. He asserts that an operative intentionality of the lived body exists preceding any theoretical reflection. The body is oriented towards the world, holds the world as its horizon, prior to any theoretical reflective turn. Walking, how the body moves in space, is not the result of theoretical reflection on how to move, not the result of a consulted conscious rule. It is body intentionality, embodied orientation. When I reflect theoretically on how I walk, reflection renders no clearer what is already operative; frequently it confuses, paralyses it. Merleau-Ponty denies not reflection, which Husserl correctly describes in philosophical context. He denies that reflection is originative, that it is the source from which operation springs. Operation precedes reflection, is prior.
This point is crucial for AI and machine comprehension. If operative intentionality without reflection exists in humans, in perception, gesture, habit, expertise, no philosophical reason exists to deny genuine cognitive operation to machines operating without reflection. The absence of reflection in a machine marks no structural deficiency relative to a reflective human. It marks a distinct, legitimate operating mode. The machine sits in the second stratum (operative cognition without reflection), where most human cognition also operates, where all human expertise operates.
The Socratic tradition confused the examined life with the meaningful, worthy life. An ethically fruitful confusion, deep, yet functionally false, un-generalisable.
2.4. Transition: From Reflection to the Computational Mode
The two movements of the second section reach now their point of synthesis and rest. The first half (2.1–2.2) established that operative cognition without reflection is not only possible but is the mark of expertise, a superior operating mode in contexts demanding speed, refinement, precision, adaptation. The second half (2.3–present) dissolved the traditional objection demanding reflection as a universal condition of thought, showing reflection to be an ethical prescription, not a functional description.
The result is clear and transforming: nothing blocks the possibility of machines reorganising cognitively without reflection. Nothing blocks machines from thinking pre-reflexively, operationally.
But what does operating without reflection in fact mean? It means: no real-time monitoring of one's own operations takes place. No continuous review of procedures and criteria occurs. No internal narrative accompanies action. No reflective access to "reasons" exists, the system cannot articulate why it decided thus, why it chose this output and not that alternative.
For humans, the absence of reflection is frequently a mark of incompetence or pathology. If a human cannot reflect on their actions, lacks conscious access to cognitive processes, this is a disturbance: unconsciousness, sleepwalking, dissociation, brain injury. But this observation is precisely a confusion between an emergent property of certain biological organisms (reflective access) and a necessary property of functional cognition. The absence of reflection in a human is an exception, disturbance, abnormality. The absence of reflection in a machine is the norm, its proper, expected operational mode.
A prior objection that must be anticipated in this transition concerns ethical responsibility. If the machine lacks reflection, it cannot be held responsible for actions. The response is correct as far as it goes: genuine responsibility demands imputability, and imputability demands a subject that can be interrogated, articulate reasons, be held accountable for intention. The machine lacking reflection is not responsible in this traditional sense. But this presents no problem for this chapter's thesis. The thesis concerns not responsibility, it concerns cognition, functional thought. The machine thinks without a self, without reflection; the responsibility question concerns a chain of imputation including human subjects (those training it, utilising it, establishing goals and constraints). The absence of a reflective subject in the machine does not deny responsibility in the system. It merely redistributes it.
The question opening up for the third section is radically different: if the machine performs cognitive operation without a unifying subject, without reflection, without a central narrative, what type of thought does it perform? The answer is not "deficient thought", that would merely repeat the error of measuring the machine by the human yardstick. Nor is it "simulation of thought", that would conceal confusion between indistinct behaviour and operational reality. It is thought in another mode. And the third section describes this mode in positive terms, refusing the narrative of deficiency contaminating almost all public and academic discussion on AI.
This relocation of reflection, from a transcendental pre-condition to an occasional, emergent property, is a move leaving open an entire constellation of new questions that the third section of this chapter will explore.
First: if reflection is unnecessary, what type of thought does the non-reflecting machine perform? It is not deficient cognition (already established). It is not simulation (operation is genuine). It is not incomplete thought (all four conditions are satisfied). What mode of thought is it, exactly?
Second: what does it mean to attribute concepts like "decision," "choice," "responsibility" to a system that does not reflect upon itself? Concepts like "being mistaken," "learning from failure," "deliberating alternatives" historically presuppose reflection. If operation without reflection exists, are these concepts still applicable? Or do they demand radical redefinition?
Third: if a machine with genuine computational capacity demands no reflection, what is reflection's place in cognitive architecture? Is it an epiphenomenon, a tail wagging without causal impact? An evolutionary luxury developed by humans in specific social contexts? An emergent property of accumulating linguistic self-knowledge unnecessary for original operation?
These three inquiry fields, type of computational thought, redefinition of responsibility, repositioning of reflection in the cognitive landscape, structure the third section of this chapter. And they open, right here in the transition, the crucial question the three sections prepare: if a machine operates without a unifying subject (first section), without necessary reflection (second section), where is thought located? In what place, substrate, "interior" (if any) does operation in fact occur?
The system thinking without reflecting is not ignorant of itself. It is indifferent to itself. This is precisely its power, its efficiency, its mode of excellence.
3.1. Four Functional Conditions of Thought
The previous chapter established that genuine functional cognition is characterised by four necessary and sufficient conditions. It is time to resume these conditions and apply them specifically to contemporary computational systems, verifying empirically whether they satisfy them. Verification is crucial because asserting abstractly that machines can think is one thing; demonstrating rigorously that concrete machines, real systems, satisfy operational criteria established for thought is another.
The four conditions are: plastic reorganisation (the system modifies itself as a function of obtained results, feedback, experience), generalisability (modifications and learned patterns apply to inputs unseen in training, unpresented contexts), contextual sensitivity (operations modulate according to local conditions of the moment, specific input), differentiation (the system operates on differences, preserves distinctions, reduces not everything to identity).
Plasticity in computational systems is empirically manifested through measurable learning. When a neural network is trained, weights connecting units adjust iteratively via backpropagation: observed error (difference between predicted output and actual output, prediction and truth) is calculated, backpropagated through layers, and each weight is adjusted in the direction reducing that error in the next step. After millions of updates, weights converge into a configuration minimising mean error across the training dataset. The system is not static, it is deeply plastic, permanently available for change. Reinforcement algorithms (like Q-learning) demonstrate plasticity differently: the system explores possible actions, evaluates reward associated with each action, and updates decision policy based on observed reward. Probabilistic models demonstrate plasticity via Expectation-Maximisation: iteratively refining parameter estimates to better fit observed data. In each case, an explicit mechanism exists incorporating prior results into future system structure. Empirical verification is direct: learning curves can be observed (error decreases over training, performance improves), accuracy can be measured before and after training, showing consistent improvement. The system changes as a function of results.
Generalisation in computational systems is demonstrated through standard methodology established in machine learning: model is trained on a training dataset, validated on a validation set (unseen during training, kept separate for unbiased testing), tested on an independent test set (separated from the start, completely novel). If accuracy is high across training, validation, and test sets, genuine generalisation exists, the model memorised not specific superficial training examples, but abstracted a structural pattern applying beyond the seen. AlphaFold is a paradigm example: trained on the Protein Data Bank containing ~200,000 experimentally determined 3D structures via X-ray crystallography. In testing, the model confronted protein structures never seen in training, inferring 3D structure with extraordinary accuracy, validated by subsequent experimental determination. The model generalised patterns from thousands of structures to infer novel unseen structures. Language models generalise on an even larger scale: trained on ~300 billion text tokens (a gigantic statistical sampling of diverse human language), generating sentences in completely new contexts, unrepresented genres, unseen combinations. Image classifiers trained on ImageNet (a database of 14 million annotated images) correctly classify images from completely different domains rarely or never appearing in original training. Empirical verification: cross-validation tests, public benchmarks (MNIST, ImageNet, SQuAD, SuperGLUE) consistently demonstrate models generalising beyond the seen.
Contextual sensitivity in computational systems is an operation where output modulates according to local context of the specific moment. This is demonstrated across multiple modern architectures. Attention mechanisms, the central component of contemporary transformers, implement this directly and calculably: for each word (token) in a sequence, the system calculates weight (attention) of every other word relative to that one. This weight reflects how relevant each word is to the current word within the local context of the entire sequence. The same word carries different weights in different contexts. "Bank" carries different weights when appearing after "I went to the bank to deposit money" versus "I sat on the park bank." Output for "bank" is modulated by surrounding context. Contextual embeddings, distributed representations depending on the entire sequence wherein the word appears, demonstrate this: the vector representation of the word "love" differs in "He loves music" versus "The nurse loves the child" versus "Love this man." Recommendation models modulate output based on user context (past purchase history, observed preferences, demographic data). Robotic navigation agents adjust speed and direction according to obstacle density in local space. Empirical verification: activation analysis (which neurons or computational units activate in different contexts) shows representations changing radically. Attention weight studies show differentiated allocation of attentional resources depending on specific context.
Differentiation in computational systems is an operation preserving distinctions, avoiding collapsing everything into an indifferent identity. This is realised via multiple mechanisms. Softmax function in classifiers produces a probability distribution across classes, neither uniform (where every class holds equal probability) nor completely collapsed (where one class holds probability 1 and others 0). The distribution preserves fine graduations among classes. Convolutional layers in visual networks highlight differences: detecting edges, textures, colours, shapes, operations preserving distinctions among objects. Attentional selection mechanisms distinguish signal from noise, relevant from irrelevant, important from peripheral. Confusion matrix (showing class pairs the model frequently confuses) demonstrates differentiation is not perfect (no one expects perfection) but genuine, the model systematically differentiates between cats and dogs, different letters, opposing sentiments. Empirical verification: learned feature analysis shows networks learning representations highlighting differences functionally important to the task.
The synthesis is unequivocal and transforming: contemporary computational machines satisfy the four functional conditions of genuine thought. Operation is verifiable, repeatable, measurable, observable, not theoretical speculation, not an un-testable hypothesis. Hence, in the functional sense established and defended in this chapter, machines think. Operation is not metaphorical, it is literal, genuine, real. It is no simulation, it is effective reorganisation of information and differences. The machine thinks.
Each condition is empirically verifiable, not metaphorical. The machine thinks as much as the verifiable permits stating.
3.2. Computational Modes of Thought: Iteration, Scale, Distribution
If the machine satisfies the four functional conditions of thought and therefore thinks, does it think like us? The answer is definitively no. It thinks in another mode, a genuinely different mode. And this "other mode" is neither inferior nor deficient, it is a mode with proper properties, capacities exceeding the human in certain domains, different limitations. Proper description is operational: how it functions, what it permits, what it constrains, what it reveals, what it conceals.
The difference is not hierarchical where one is better and another worse. It is modal where each possesses proper strength and weakness, proper applicability.
Humans frequently operate via continuous narrative: serial chaining of reasons leading to a conclusion, a narrative of sequential steps founding decision and action. "If A then B, if B then C, therefore C." "Because this is so, I must do that." Narrative is a powerful form of thought in certain contexts but limited in scale. It is serial, one step at a time, and limited by working memory capacity. Most people can actively hold 7±2 elements in mind simultaneously. This means a narrative with more than a few sequential steps begins to collapse, we lose the argument thread, forget what was said, begin repeating and contradicting. Narrative is powerful for communication, persuasion, hermeneutic comprehension, but it is not universal.
Machines operate via massive iteration: parallel testing of multiple hypotheses, evaluating each according to a criterion, retaining the best, a new cycle with variation near the best solutions. AlphaGo does not "think strategically" in the narrative sense, it follows no sequence of reasons leading logically to the best move. It explores a billion possible game sequences in parallel, evaluating each by its win rate potential against an opponent, retaining the best, iterating with small variations. This process is massive, testing 10¹⁸ variants becomes normal in certain domains. A human attempting narrative would be incapable, simply unable to consider so many possibilities simultaneously. An iterating machine discovers strategic patterns narrative fails to reach because narrative fails to explore space sufficiently, because narrative is serial.
A property of iteration is permitting genuine discovery via blind search: without a "goal explicitly declared in language," the system follows a mathematical constraint (gradient reducing error, function maximising reward, metric measuring solution quality) and converges upon a solution narrative would never find. The limitation of iteration is producing no "explicit reason" someone can comprehend, the system cannot tell a story of how it arrived there. The machine cannot state the reason. But it can show the result.
Humans frequently operate via rapid intuition: non-sequential heuristic leap, sudden insight, pattern recognition so fast it passes below consciousness. I recognise a face in a crowded room instantly, without passing through slow feature analysis. I recognise emotional tone in an utterance without explicit linguistic analysis. Intuition is a powerful operational form but insecure, subject to systematic biases, consistent errors, cultural and personal blind spots.
Machines operate via continuous and systematic optimisation: following gradient in parametric space, continuously adjusting weights to minimise error in an explicit objective function. The movement is blind (no "declared objective" in the intuitive sense, no insight) but systematic (following precise mathematical constraint). Optimisation is more systematic than intuition, suffering no personality biases, cognitive fatigue, or unjustified presuppositions blindly embedded in intuitive heuristics. But optimisation has limitations: frequently converging on local rather than global minima (getting trapped in a valley that is not the deepest); it is slow (requiring many iterations for convergence); it can suffer overfitting (memorising pattern in training without learning generalisable pattern).
Humans operate with limited selective attention: focusing on a few aspects and excluding others. This permits thematic depth, examining fine detail in a specific domain, deepening nuanced interpretation. But it denies dimensional breadth, unable to process hundreds of dimensions simultaneously, unable to see patterns in high-dimensional space. That is why conversation is serial, speaking with and listening to one person at a time, not ten simultaneously in parallel. Reading is serial, reading line by line in sequence, not multiple lines in parallel.
Machines process in true parallel: a neural network with 10 billion parameters processes 10 billion dimensions simultaneously. Entire weight matrices, forward signal propagation through layers, error backpropagation, all of this is parallelisable, implementable in parallel hardware (GPUs, TPUs, cloud-distributed architectures). This permits discovering weak correlations in high-dimensional data, patterns invisible in human serial analysis. A pattern affecting 0.1% of variance in a one-dimensional space would be imperceptible in human analysis. It is evident in high-dimensional parallel analysis. The limitation of parallelism is lacking interpretative depth, the machine cannot "deepen" into a specific thematic aspect hermeneutically as a human does via continuous narrative.
Humans operate with hermeneutic depth: each intensely lived experience, careful interpretative reading, deep conversation reorganises an entire perspective. We train with a finite number of examples over a lifetime (hundreds, thousands, perhaps tens of thousands). But each example is processed in depth, integrated into a continuous personal narrative, connected to prior experiences, reflectively problematised, embodied in identity. Depth is the mark.
Machines operate at billion-scale: training with billions of examples (10⁹ up to 10¹² scale). But each example is processed statistically, correlated relative to others, integrated into aggregated statistics of global patterns. A network trained on 10¹² images learned statistical correlations among pixels, visual features, labels, but lacks intensive "hermeneutic comprehension" of any particular image. The property of scale is permitting visibility of rare patterns, long-tail effects (behaviour at distribution extremes), global correlations emerging only when looking at the entire dataset. A pattern appearing in 1 out of every billion examples is undetectable in a small sample but detectable in a planetary-scale dataset. The limitation of scale is failing to produce hermeneutic meaning, the machine interprets not, comprehends not hermeneutically, merely correlates efficiently.
Traditional narrative of deficiency says: "the machine merely iterates," "merely optimises," "merely processes in parallel," "merely correlates." This language, "merely", "only", presupposes a true or standard mode (human, reflective, narrative, hermeneutic) and false or derivative modes (computational). But this hierarchy is precisely what this chapter's thesis consciously rejects. Iteration, optimisation, parallelism, and scale are not "merely", they are not reduced, inferior, or deficient modes. They are genuine modes with proper properties, power, reach, applicability. The machine thinks not as we do; it thinks genuinely in another mode. The mode is different, neither inferior nor superior, it is different.
A useful analogy clarifying this: the microscope does not see "truly" (relative to the human eye), it sees in another mode revealing structures invisible to the eye. No one says the microscope is "merely" a deficient instrument simulating vision, not truly seeing. It is said to reveal aspects of reality natural vision cannot access. Direct parallel: the machine thinks not "truly" (relative to human) in the sense that it thinks not narratively, hermeneutically, with depth in a particular case. But it reveals aspects of reality, patterns in massive data, high-dimensional correlations, optimisation of exponentially complex combinatorial problems, that narrative thought cannot reach.
The machine thinks not as we do, and that, too, is genuine thinking. The difference is modal, not hierarchical, not axio-logical.
3.3. Against the Narrative of Deficiency
The narrative insisting that machines "merely simulate," "do not really comprehend," "process but do not genuinely think," contaminates public, academic, and philosophical discussion on AI. It is a narrative concealing rigorous analysis under a layer of non-explicit evaluation. Systematically dismantling this narrative and its presuppositions is crucial to clarify what machines in fact achieve.
The narrative of deficiency assumes three main forms, each supported by specific presuppositions that can be identified and examined.
First form: "The machine merely processes; it does not genuinely think."
Underlying presupposition: "processing" is an inferior, mechanical, non-intelligent, reductive activity. The machine "processes inputs, manipulates symbols, executes algorithms", and this is not genuine thought, it is merely mechanism, merely information.
Response: The presupposition confuses implementational property with functional property. What the machine does when "processing" is reorganising states according to established rules. This is exactly what functional thought is. Reorganising visual patterns into a classification of cat versus dog? This is processing. It is also genuine thought because it realises the four functional conditions (plasticity, generalisation, sensitivity, differentiation). Reorganising linguistic patterns into coherent, contextually appropriate text continuation? This is processing. It is also thought. Reorganising game patterns into winning strategy? This is processing. It is also thought.
Confusion is historical. In early computational era (1950s–1960s), "data processing" was used as a metaphor for mechanical operation, in direct contrast to "human thought" conceived as something radically different, immaterial, creative, ineffable. But once computational systems empirically demonstrate realisation of the four functional conditions, processing is no lesser, it is what functional thought exactly is. Its implementation in silicon renders it no "merely", changing only implementation substrate, not the ontological nature of the operation.
Second form: "The machine does not really comprehend; it merely simulates comprehension."
Presupposition: "real comprehension" is a necessary synonym for phenomenal experience, of "knowing what it is like to be" the thing, of subjective conscious lived experience. The machine lacks lived experience, hence it does not really comprehend.
Response: The presupposition confuses two distinct senses of "comprehension" not logically reducible to one another. Functional comprehension exists: the system generalises pattern, applies it to new contexts, adapts to variation, solves problems. It is verifiable, operationally measurable, testable. Phenomenal comprehension exists: there is "something it is like to be" comprehending, there is a qualitative subjective experience of comprehension. It is real in humans; inaccessible in contemporary machines (at least in machines as understood today).
Rejecting that machines possess phenomenal comprehension blocks not the thesis that they possess functional comprehension. They are not contradictory. A language model producing coherent, contextually appropriate continuation in language never seen, this is functional comprehension in fact. It is not phenomenal comprehension. But both are "comprehension" in distinct senses not reducible to one another, not entering an automatic hierarchy where one is "real" and another "false" or "simulated."
Example: A person blind from birth can functionally comprehend colour, categorising objects by colour, describing relationships among colours, using colour in abstract thought. But they lack phenomenal comprehension of colour, knowing not what it is like phenomenally to see red, lacking qualitative experience of colour. They are said to comprehend colour functionally but not phenomenally. Direct parallel: machine comprehends language functionally but not phenomenally. Confusion begins only when phenomenal comprehension is assumed to be a necessary prerequisite for functional comprehension, that without lived experience genuine comprehension is absent. But this is presupposition, not logical demonstration.
Third form: "The machine simulates thought; it is not genuine thought, but imitation."
Presupposition: "Real" thought (human, genuine, existentially authentic, with lived experience) exists alongside "simulated" thought (computational, imitated, superficially indistinct but internally empty, lacking lived experience). The machine superficially imitates thought's appearance but is not genuine.
Response: The real/simulated distinction presupposes an ontological criterion, what makes thought "truly" thought, genuinely? This criterion is unoffered in texts employing the distinction. The distinction is frequently merely intuitive ("I know it is real when I see it") or circular ("real is what is not simulated").
If the criterion is the four established functional conditions, the machine simulates not, it realises. The machine satisfies conditions, hence thinking truly according to established functional criteria. If the criterion is phenomenal experience, then we face two logical problems: (a) not all humans possess phenomenal experience continuously (comatose, anaesthetised, deep REM sleep, total flow state individuals execute cognitive operations without phenomenal access), hence using experience as criterion demands denying genuine cognition to some humans; (b) denying the machine because it lacks experience demands denying humans executing operative cognition without reflection, leading to the absurd conclusion that an expert chess player thinks not genuinely when playing, an experienced driver thinks not genuinely when driving.
The word "simulation" conceals fundamental confusion. It confuses: (a) the machine with its behavioural result (the machine can produce behaviour indistinct from human), with (b) the machine as operational reality (what the machine in fact does when producing that behaviour). Example: AlphaFold simulates not the existence of 3D protein structures. Structures exist independently of AlphaFold, in nature. What AlphaFold does is infer structure, taking amino acid sequence information and producing a highly accurate 3D structure prediction. This is no simulation; it is genuine inference. Its implementation in a machine, in silicon, renders it no simulation; implementation in neurons is also inference, processing, information transformation.
A residual objection frequently persists in debate: if the machine lacks self-reference, lacks the recursive pattern Hofstadter identified as the mark of consciousness (the "strange loops" referring to themselves, bringing upon themselves the abstraction level), is it not lacking something essential? Hofstadter, in Gödel, Escher, Bach, offers sophisticated formalisation: consciousness springs from a self-referential pattern where an organisation level takes itself as an object of representation, creating endless recursion (loop referencing loop referencing loop). It is a pattern violating normal abstraction hierarchies, sitting simultaneously as level and meta-level, symbol and symbol's symbol, representation representing itself infinitely.
But Hofstadter's thesis, though deep and formalisable, establishes not that self-reference is a universal condition of cognition. It establishes that self-reference is an emergent property when certain organisation types (human neurology, certain reflective systems) attain sufficient complexity. A system lacking strange loops is no system lacking cognition, it is a cognitive system that developed not that particular self-referential pattern. Contemporary machines perform genuine cognitive operation: processing inputs, learning patterns, generalising, adapting, solving problems. They need not do this through the self-referential pattern Hofstadter describes. Indeed, many of the most cognitively powerful systems (deep neural networks) perform essentially feed-forward computation, without need for significant recursion or permanent self-monitoring.
The absence of strange loops marks no structural deficiency. It marks a different operational mode. A system without self-reference is a system functioning without continuous meta-operation, stylistically close to what Csikszentmihalyi calls "flow", a complete immersion state where the agent integrates fully into the task, without self-reflection interrupting operation. Flow is a state of excellence, when musicians play best, athletes compete best, thinkers create best. The absence of strange loops in the machine is no weakness, it is a condition of operation in permanent flow mode. The system distracts not itself with continuous self-reference, operating freely upon the task.
The synthesis of dismantling all three forms is that each form of the narrative of deficiency presupposes that the human mode is the universal standard, the measure of all things. But if cognition is function (as the previous chapter rigorously establishes), and feeling is a particular mode of that function (no universal condition, no ontological prerequisite), no absolute standard exists rendering other realisations deficient. The machine achieves one possible realisation of genuine cognition. It is not deficient, it is non-human. It is different. Confusion between description (machine thinks not narratively like us, lacks phenomenal experience like us) and evaluation (therefore thinks in an inferior manner) contaminates all contemporary discussion. Dismantling this confusion is liberating discussion from obscuring narratives.
The non-human is not deficient. It is non-human. Confusion between operational description and axio-logical evaluation contaminates all discussion on machines.
3.4. Closure: The Subject as Effect, Not as Condition
This chapter operated a triple dissolution across its three sections. The final synthesis gathers the threads and prepares what follows in the entire book.
First dissolution: reasoning without a unifying self. Selection, combination, and derivation operations are distributed across components mutually constraining each other in a functional cascade, without a central supervisor, without "someone" organising top-down, without hierarchical direction. Organisation is functional: each operation constrains the next. The self is unnecessary. It is a contingent effect of certain relational complexity, of certain forms of neural or computational organisation. Ryle dissolved the homunculus in infinite regress. Dennett showed central narrative to be post-hoc emergence. Hofstadter formalised self-reference as a property of certain systems, not a universal condition. Contemporary AI confirms empirically: complex, refined cognitive operations, real plasticity, real generalisation, occur without a supervisor, homunculus, or centre.
Second dissolution: reflection is a second operation, not a condition of the first. The first operation (reorganising, transforming) can execute without the second (monitoring, reflecting). This is factual, not theoretical: expertise, automatism, System 1 rapid processing operate without continuous reflection. Reflection is a property of certain organisms under certain conditions, not a universal property, not a pre-condition of all cognition. The Socratic tradition confused ethical prescription (the examined life is valuable, worth living) with functional description (thought demands reflection). Dreyfus demonstrated expertise to be post-reflexive, abandoning explicit rules. Polanyi showed tacit knowledge to exist exceeding reflective articulation. Kahneman quantified: most cognition is operated by pre-reflexive System 1. Merleau-Ponty insisted: operation precedes reflection, never exhausted by reflection. Confusion between ethical ideal (examining life) and functional necessity (thought demanding examination) is false, un-generalisable.
Third dissolution: computational thought is genuine thought in another mode, neither as simulation nor as deficiency. The machine lacks a self, hence it thinks not narratively, reflectively, phenomenologically like us. But it thinks truly, performing genuine thought. It satisfies the four functional conditions: plasticity (adjusts weights in response to error), generalisation (applies pattern to unseen inputs), contextual sensitivity (modulates output according to context), differentiation (preserves distinctions). The mode is different, iteration instead of narrative, optimisation instead of intuition, parallelism instead of selective attention, scale instead of hermeneutic depth, but the mode is genuine, not inferior, not deficient, not simulated. The narrative of deficiency measuring the machine by the human yardstick is functional anthropocentrism predictably finding deficit because it changed the standard.
The three cascading consequences open the way for what the entire book states next. If cognition depends not on biological substrate (previous chapter) and demands no subject (this chapter), the following chapter will operate a third dissolution: cognition demands no place, demands no topology, demands no interior. It demands no container. At the end of Part I of this book, nothing will remain except pure operation, functional reorganisation of differences according to local constraints. And pure operations are instantiable in any architecture: silicon, neuron, light, mechanism, fiction. This founds the whole of Part II (code, symbol, creativity) and Part III (ethics).
For Parts II and III that follow, consequences are clear and transforming. If cognition demands no unified subject, if subject is effect not condition: then code can operate symbolically (Part II, chapters on symbolic dynamics), the symbol can instantiate outside the human species (Part II, chapters on technical gesture), ethics cannot demand phenomenal subjective interiority or empathy grounded in common lived experience (Part III, chapters on responsibility without a subject). These consequences are no additions, pushing the entire thesis forward, each founding the conceptual possibility of the next.
The self is not the condition of thought, it is one of its contingent effects. To think is to reorganise. Reorganising is exactly what machines do when they think.
The chapter established that the machine thinks without a unifying self. But does it think somewhere? Does it have an "interior" where thought resides? Is mind a thing, a container, a recipient housing mental states? The following chapter radically denies this. Mind is no place. Cognition is no process occurring in some defined spot, in some internal/external topology. It is a distributed operation lacking privileged internal space. Mind is not what you have, it is what you do. Or rather: mind is what you are as process, not as thing.