reductio.EXPLANATION / COGNITION / AI
ESSAY 0016 SEPTEMBER 2026

Understanding
Modern AI.

From connectionism to language models,
evolutionary search, and systems that can act

30 MIN READ CONNECTIONISM · EVOLUTION · LANGUAGE MODELS
Contents

The most useful way to understand contemporary AI is to separate how a model acquires capabilities, how it computes an answer, and how a larger system turns answers into successful activity. These are different mechanisms, operating on different timescales. Much confusion comes from treating them as one thing called “the AI.”

Three ideas point toward a coherent account. A small number of developments can explain much of the transition from 1990s connectionism to useful AI. Evolutionary selection really has been implemented, although it is usually not what ordinary artificial neurons are doing. And the surrounding software—the harness—can materially change what a model accomplishes. Hofstadter supplies a further question: does a system preserve the right relationships when a problem changes its appearance?

My central assessment is that modern AI combines learned representations with increasingly effective ways to spend computation, consult external information, and select or repair candidate solutions. That is a stronger explanation than either “the machines now think like us” or “they merely autocomplete.” The five-part account below is my synthesis, not an agreed historical canon or a measured allocation of credit.

The report concentrates on language models and agents used for intellectual and practical work. It is not an exhaustive history of robotics, computer vision, scientific AI, or generative media. Research and engineering reports are identified as such; older experiments establish mechanisms and historical findings, not performance ceilings for September 2026 models. Public descriptions do not reveal every frontier training recipe.

Three places capability comes from
LearningData + objective + optimization
Trained modelReusable representations and behaviors
Working on a taskModel + current context + computation
Candidate action or answerA proposal, not yet a verified result

Training changes parameters. Working and testing can improve an outcome while leaving those parameters fixed.

1. The five developments that explain most of the change

Assume the starting point is already familiar: interconnected units, distributed representations, adjustable weights, and learning from experience. From that starting point, these are the five developments I would retain.

Development Bottleneck it addressed What it made possible
1. Training deep networks reliably at useful scale Networks could represent complicated functions but were difficult or expensive to train well. Learning rich internal representations from large datasets rather than relying mainly on human-designed features.
2. Broad self-supervised pretraining and transfer Every task appeared to require its own examples, labels, and model. One model acquiring reusable capabilities from ordinary data and adapting to tasks described in its context.
3. Attention and the transformer Sequential processing made long relationships and large-scale training difficult. Flexible exchange of information among token positions and efficient parallel training.
4. Post-training for instructions, preferences, and reasoning A capable predictor was not automatically a useful assistant or persistent problem solver. More effective instruction following, tool behavior, and extended problem-solving strategies.
5. Inference-time computation and an external working environment A single answer had limited opportunities to investigate, test, revise, or preserve progress. Systems that search, use tools, evaluate results, and continue work across multiple steps or sessions.

These developments overlap chronologically. Scale, data quality, and engineering run through all five. If we counted distinct inventions, five would be far too few. If we counted explanatory transitions, five is defensible.

1.1 Deep learning became trainable and economical

The central connectionist idea survived: useful representations can emerge through adjusting connections rather than being specified in advance. Rumelhart, Hinton, and Williams's 1986 paper showed how error signals could shape hidden representations using backpropagation. It was a pivotal exposition and demonstration, not the first appearance of every underlying mathematical idea. Original paper.

It helps to separate three things that are often collapsed:

  • The objective: what counts as a good output.
  • Backpropagation: calculating how a small change in each parameter would affect the objective.
  • The optimizer: using those calculations to update parameters.

Reinforcement learning is a way of learning from rewards for behavior. It is not another name for backpropagation. A reinforcement-learning system can use backpropagation to update a neural network; a network can also use backpropagation without reinforcement learning.

The later breakthrough was an accumulation of practical advances: hardware able to perform enormous amounts of matrix arithmetic, large datasets, useful architectures, and methods that kept deep training workable. AlexNet's 2012 ImageNet result made the advantage of large, GPU-trained convolutional networks conspicuous. Residual networks subsequently demonstrated how shortcut connections could make very deep models easier to optimize. AlexNet paper; residual learning paper.

A residual connection lets a layer modify a running representation while preserving a direct path for existing information. That makes refinement easier to learn than requiring each layer to reconstruct everything. This is an example of a seemingly modest architectural decision having a large cumulative effect.

Calling hardware and optimization “mere engineering” misses their causal role. An algorithm that is theoretically adequate but cannot be trained within a feasible budget is not a usable route to intelligence. The engineering changed which hypotheses could actually be tested and which models could exist.

1.2 Prediction became a source of broad competence

The most important correction to “basic neural nets with reinforcement learning” is that much of an LLM's broad capability is acquired through self-supervised pretraining.

For an autoregressive language model, a training example can be made by taking existing text, exposing an initial portion, and asking the model to predict the next token. Tokens are pieces of text, not necessarily whole words. The actual continuation supplies a target automatically. One document therefore provides many training signals without someone separately labeling its grammar, argument structure, factual content, and style.

Why might this produce more than superficial word associations? Consider my illustrative sentence: “The package was too large for the drawer, so we replaced the …” Predicting a plausible continuation can benefit from representing which object is larger, what containment requires, and why someone might replace one object rather than another. Across varied examples, reusable abstractions can improve prediction.

This is an explanatory argument, not proof that every model learns the intended causal structure. Shortcuts, memorization, and genuine abstraction can coexist. The objective rewards successful prediction; it does not separately demand a philosophically satisfactory theory of the world.

The decisive empirical development was broad transfer. GPT-3 demonstrated substantial performance on many tasks described through instructions and examples in the prompt, without updating the model's weights for each task. That is in-context learning: the model's behavior adapts to the current material even though its stored parameters remain fixed during that evaluation. GPT-3 paper.

Think of pretraining as acquiring a repertoire of ways to interpret and transform information. The prompt supplies a situation in which to deploy that repertoire. This analogy is useful as long as we do not confuse the temporary situation with permanent learning.

Scale mattered, but parameter count alone was an incomplete story. The Chinchilla experiments found that, under their training-budget assumptions, the balance between model size and amount of training data was crucial. A smaller model trained on more data could outperform much larger models. Their particular scaling prescription is not an eternal rule for every deployment budget; it demonstrated the importance of allocating computation intelligently. Chinchilla paper.

For your purposes, this explains why the same assistant can discuss philosophy, draft prose, interpret a table, and write software. Those activities need not be separate handcrafted modules. It also explains why apparent familiarity with a legal citation is not evidence that the system has retrieved or checked the actual source.

1.3 The transformer changed how information moves

An old recurrent language model processes a sequence through a changing internal state. Information from an early word must survive that sequential path to matter later. Attention provides a different route: a position can draw information directly from other available positions, with weights determined by their current representations.

The 2017 transformer organized sequence processing around attention rather than recurrence or convolution. An especially consequential advantage was parallel training across positions. Causal language generation still ordinarily produces successive tokens in sequence; parallel training does not mean an entire answer is generated simultaneously. Transformer paper.

Here is the essential mechanism, in plain language. Each position constructs numerical descriptions of what information it is seeking and what information it can supply. Their compatibility determines how information is combined. Layers repeat these operations, interleaved with transformations that compute new features. The words “seeking” and “supply” describe numerical roles, not intentions.

For an invented example, compare “The witness contradicted the report because it was inaccurate” with “The witness contradicted the report because she had been there.” Useful processing must connect different parts of the sentence depending on the relationship being expressed. Attention provides an adaptable information-routing mechanism; it does not by itself guarantee the correct interpretation.

This remains recognizably connectionist. Its representations are distributed numerical patterns, transformed by learned parameters. A mathematical analysis of simplified transformers shows how attention and other components can implement identifiable information-processing circuits. It also shows why “one neuron equals one concept” is generally a poor starting assumption. Transformer-circuits framework.

The transformer is therefore a substantial architectural innovation within the connectionist tradition. It did not require discovering a little symbolic reasoner hidden inside every network.

1.4 Post-training turned prediction into useful behavior

A model trained to continue text has learned from dialogues, stories, essays, errors, arguments, and many other patterns. It is not automatically committed to treating your message as a request to fulfill accurately. Additional training shapes which behaviors it tends to produce.

Supervised fine-tuning teaches through examples of desired responses. Preference-based methods use comparisons or evaluations of outputs; reinforcement learning can increase the probability of behaviors that receive better rewards. These are distinct methods that can be combined. InstructGPT is an early influential demonstration: on the authors' evaluation distribution, human raters preferred outputs from a much smaller instruction-trained model over a far larger base model. That finding concerned preference and assistance, not superiority on every cognitive task. InstructGPT paper.

Reasoning-oriented training adds another dimension. A system can generate a solution, receive a reward based on whether the answer is correct, and update its future behavior. Where correctness is checkable, it becomes possible to favor strategies that reliably reach successful outcomes rather than merely resemble polished explanations.

DeepSeek-R1 provides a publicly documented example. Its R1-Zero experiment applied reinforcement learning to an already pretrained base model without preliminary supervised fine-tuning. “Zero” did not mean starting with an empty network and learning all language and knowledge from reward. The released R1 used a more elaborate combination of supervised stages, reinforcement learning, and selected training outputs. The researchers also transferred capabilities into smaller models through distillation. R1 paper.

The conceptual advance is the combination of a powerful pretrained starting point, usable feedback, and enough training to favor sustained problem solving. Reward can reinforce checking, backtracking, and longer solution attempts. This does not mean every apparent act of reflection is reliable, or that reward training invents all capabilities from nothing.

1.5 Systems learned to spend effort after receiving the task

There are two fundamentally different places to spend computation: before deployment, changing the model; and after a request arrives, working on that particular request.

Search itself is much older than LLMs. AlphaGo's 2016 paper already combined learned policy and value networks with tree search. Learned judgment helped allocate exploration, and exploration made the whole system stronger than a bare prediction of the next move. The present transition is the extension of such combinations into broad language-based work and tool environments, not the invention of search after 2020. AlphaGo paper.

Additional inference-time computation can mean longer intermediate reasoning, multiple candidate answers, search over possible steps, or repeated revisions with a verifier. A study by Snell and colleagues showed that how this computation is allocated can matter as much as simply increasing it, with the useful strategy depending on task difficulty and the system's components. It does not establish that unlimited reflection can compensate for any weakness in a model. Test-time compute study.

External tools add something qualitatively different from another paragraph of thought. A calculator supplies an actual calculation; a database supplies retrieved records; a code runner supplies observed execution results. Retrieval-augmented generation is one influential implementation of combining a generator with external information. ReAct demonstrated the value of interleaving reasoning and actions so that observations could change subsequent decisions. RAG paper; ReAct paper.

The harness is the software that manages this activity: supplying instructions and context, executing tool requests, returning results, preserving state, handling interruptions, and deciding when to continue or stop. An agent is the resulting system when model outputs help determine its next actions. A particular implementation may be a simple loop or a more elaborate arrangement.

My assessment is that this fifth development explains a large part of the gap between an impressive chatbot answer and useful delegated work. It does not establish that the harness is more important than the model in all situations. A competent research environment cannot fully compensate for a model that consistently misreads its evidence.

2. What “training alchemy” actually consists of

The phrase captures something real: knowing the architecture and a headline training objective does not tell you how to reproduce a capable model. The recipe includes data mixtures, filtering, duplication control, sequence lengths, optimization settings, curricula, synthetic examples, reward design, and the order in which stages are performed.

Those choices affect what the system encounters and what behavior is reinforced. To use an analogy from education, “we trained people by having them read and answer questions” leaves out almost everything needed to reproduce the outcome. Which readings? Which questions? In what sequence? Who grades them? What errors receive correction?

Not every detail is mysterious. Many mechanisms are well understood locally, and experiments can establish which change improves a defined evaluation. What remains incomplete is a comprehensive theory predicting the final mix of abilities, failure modes, and transfer from the entire recipe.

A concrete public example is DeepSeek-V3: its report describes mixture-of-experts organization, attention changes, lower-precision training, load management, and other system choices. These details affect the amount of capability obtainable from a computation budget. A mixture of experts activates selected parameter groups for a token, rather than using all parameters equally on every step. It does not imply a committee of independent miniature people. DeepSeek-V3 technical report.

Three distinctions help evaluate claims about recipes:

Claim What to ask
“This training method improved the model.” On which tasks, compared with what baseline, and with how much additional data and computation?
“This smaller model matches a larger one.” On which distribution, with what inference budget, and did it learn from a stronger teacher?
“This agent beats the same model in chat.” Did it receive more evidence, better tools, more attempts, or an external evaluator?

None of those advantages is illegitimate. The point is to identify the mechanism so you can reproduce the benefit and recognize where it will disappear.

3. Dawkins: persistence, made operational

Your formulation has a direct textual connection to Dawkins. Early in The Selfish Gene, he places biological evolution within a broader discussion of “survival of the stable.” The passage also expressly recognizes that stability alone is insufficient to explain the origin of highly complex organisms. Dawkins, authorized book preview, printed pp. 15–17.

I would sharpen your proposition this way: in a specified environment, patterns that are retained or regenerated more successfully become more prevalent; when variants reproduce with inheritance, differential success can accumulate adaptations across generations. The second clause is doing additional explanatory work.

There are three distinguishable cases:

  1. Persistence: an object remains because it resists destruction.
  2. Repeated production: a pattern remains common because a process continually produces it.
  3. Evolution through heredity: variants leave descendants that retain relevant differences, allowing those differences to affect future representation.

A long-lived rock illustrates the first. A repeatedly generated wave pattern illustrates the second. A lineage of replicators illustrates the third. The third can accumulate a history of adaptations in a way the first two, by themselves, cannot.

The important quantity need not be an individual molecule's lifespan. It can be the continuation of a lineage or information pattern through many short-lived instances. Dawkins's gene-centered account emphasizes genes as units of hereditary information and organisms as vehicles through which replication occurs. This is not a claim that genes consciously want anything or that cooperation is incompatible with selection. Oxford University Press's account of the book.

A small thought experiment

Suppose there are initially 100 instances of each of two heritable patterns. In an intentionally simplified, unlimited-resource environment, each time step has independently specified survival and production rates:

Pattern Existing instances surviving per step New instances per starting instance Expected population multiplier
A: durable, non-reproducing 99% 0 0.99
B: fragile, reproducing 50% 0.70 1.20

After ten steps, A has about 90 instances and B about 619. B's individual instances are less durable, but its lineage expands. These are illustrative expectations, not biological measurements. The calculation assumes faithful inheritance, constant rates, and no resource constraint; changing those conditions can reverse the result.

The explanation is not merely “the things that survived survived.” We specified rates independently and derived a prediction. For adaptation, we must additionally explain variation, inheritance, the relationship between traits and these rates, and the environmental conditions sustaining the process.

Explore the persistence example

Keep A’s survival at 99% and B’s at 50%. Change B’s reproduction rate or the elapsed steps. Both start with 100 instances. These are expected counts under fixed rates, faithful inheritance, and unlimited resources; the model does not simulate mutation.

A: durable
B: reproducing

3.1 Has this been implemented in AI? Yes, quite literally

Evolutionary computation represents candidate solutions, generates variants, measures performance, and preferentially retains or reproduces useful candidates. Genetic algorithms belong to an established research tradition, including John Holland's work; they should not be described simply as a later application of Dawkins's 1976 book. Holland's foundational book.

Neuroevolution applies this approach to neural networks. NEAT evolves both connection weights and network topology. Its design includes mechanisms to align inherited structures and protect novel forms long enough to develop. Here, a representation of a network really does serve as something like a genome, with evaluated descendants. That is a much closer evolutionary implementation than saying a highly active neuron has “won.” NEAT paper.

Evolution strategies can optimize network behavior from scores obtained by perturbing parameters and evaluating the resulting candidates. A 2017 study demonstrated large-scale applications to control tasks. Such methods can estimate useful parameter updates from performance scores without backpropagating through the policy network. Their existence also shows that “reinforcement-learning task” and “evolutionary optimization method” are not mutually exclusive descriptions of a project. Evolution-strategies paper.

Population-based training combines ordinary neural-network learning with an outer selection process. Multiple training runs proceed; less successful runs can copy stronger ones and alter training settings. Local learning and population selection can therefore coexist. Population-based training.

AlphaEvolve is the most illuminating bridge to present-day language-model agents. Its LLMs propose changes to programs, automatic evaluators score the resulting programs, and a database retains candidates that inform later proposals. The evolved entities are programs, not the individual neurons of the proposing LLM. The authors report applications to mathematical constructions and computational infrastructure. The method is especially suited to tasks with machine-executable evaluation. AlphaEvolve paper.

This combines a learned generator with an evolutionary search process. The generator contributes informed variation; the evaluator supplies selection pressure; retained candidates preserve useful history. The model need not independently invent a survival drive for this arrangement to work.

3.2 Are ordinary neurons or model nodes selfish replicators?

Usually, that is the wrong level of description. In an ordinary transformer training run, numerical parameters are adjusted within an architecture. A neuron does not ordinarily reproduce into a lineage of competing descendants simply because it activates strongly. Attention allocates influence among representations; it is not automatically a reproduction mechanism.

A feature may become more useful to prediction, its associated circuitry may strengthen, and some parameters may become dispensable. That licenses an optimization or selection analogy. It does not establish literal Darwinian evolution among the nodes.

There is another difficulty: “the node” may not be the relevant unit of representation. Work on superposition demonstrates in toy networks how more features than available dimensions can be represented through overlapping patterns. This is part of the reason researchers investigate features and circuits rather than assuming every concept occupies its own neuron. A feature's distributed implementation complicates any proposed mapping from concept to replicator. Superposition research.

For biological brains, there are explicit selectionist and Darwinian hypotheses. Edelman's neural Darwinism concerns selection among neuronal groups. Fernando, Szathmáry, and Husbands distinguish selectionist accounts from stronger proposals involving actual replication of neural information. They also explore mathematical relationships between learning rules and evolutionary dynamics. Their neuronal-replicator proposals are research hypotheses, not established descriptions of how human thinking or mainstream LLM training works. Selectionist and evolutionary approaches to brain function.

The useful test is specific: what varies, what is copied, what is inherited, and what determines differential retention or reproduction? If those questions cannot be answered, “evolution” may be an evocative metaphor rather than an implementation description.

3.3 The strongest practical consequence: inspect the environment

Your emphasis on “in a particular environment” is essential. Consider my hypothetical AI workflow that generates ten legal arguments and retains the one receiving the highest score. If the evaluator mainly rewards fluency and confidence, selection will favor those properties. If it checks source support, factual consistency, and the strongest counterargument, selection pressure changes.

The evolutionary lesson is not that retained outputs must be true. It is that outputs become adapted to whatever actually determines retention. A successful test suite, persuasive prose, a pleased reviewer, and a correct real-world result are different selection criteria.

There is a particularly important failure mode: if the generator can also redefine what counts as success, an apparent improvement can reflect an easier test. In any serious delegated workflow, preserve the requirement being tested independently of the proposed answer. That is my practical inference from the selection framework, rather than a claim that a particular legal product implements evolutionary search.

4. Hofstadter: influence, implementation, and an unfinished challenge

It is useful to distinguish Hofstadter's work on analogy and fluid concepts from his work on self-reference and consciousness. They raise related but different questions.

4.1 Analogy was a computational research program

Hofstadter and Melanie Mitchell's Copycat was an implemented model of analogy-making. It explored how a situation could be represented in different ways and how one representation could guide a corresponding transformation elsewhere. The program's letter-string microworld was deliberately restricted; it was not a general conversational assistant. Hofstadter and Mitchell's Copycat chapter.

Its architecture included a workspace holding developing structures, a network of concepts, small processes called codelets that proposed and tested structures, and a temperature-like control of randomness. Local proposals and global context influenced each other. Mitchell explicitly presented it as a complex adaptive system, drawing comparisons with distributed biological systems. Mitchell's account of Copycat.

The compelling idea is that intelligent analogy requires choosing what the objects and relations are, not merely matching objects whose descriptions are already fixed. In my legal example, two disputes may involve completely different technologies but share the same information asymmetry and commercial pressure. Conversely, two disputes involving the same technology may have very different strategic structures.

This is why “find me similar cases” can be underspecified even before a search begins. Similar according to what relationship, at what level of abstraction, and for which purpose? Hofstadter's work addresses that representational flexibility directly.

4.2 Did the transformer adopt his architecture?

The defensible answer is not in the direct architectural sense. The transformer paper and the mainstream training accounts discussed here describe attention, learned representations, prediction, and optimization. They do not present their systems as implementations of Copycat's codelet-and-slipnet architecture.

This is a scoped conclusion from the documented designs, not proof that Hofstadter influenced no researcher. Intellectual influence is harder to establish than a shared mechanism. Similar vocabulary—attention, context, emergence, competing possibilities—does not establish historical descent.

There are interesting functional parallels: context changes what relationships are salient; partial interpretations affect later processing; alternative possibilities can compete. The modern implementation usually learns much of that behavior in numerical parameters, while Copycat explicitly organized a small world around mechanisms for building and revising interpretations.

An LLM can also describe a Hofstadter-style analogy because his work is discussed in the culture on which models may train. That behavioral fact, by itself, would establish neither architectural adoption nor reliable transfer to unfamiliar problems.

4.3 Was he wrong about analogy being central?

There is no need to force a winner. Contemporary LLMs achieved broad useful behavior without engineers first implementing Hofstadter's proposed architecture. That is evidence against treating that architecture as an established prerequisite for useful AI. It does not establish that abstraction and analogy are peripheral to intelligence.

There is also a real empirical dispute about how robust LLM analogy-making is. Webb and colleagues reported strong performance on several analogy tasks. Lewis and Mitchell then used counterfactual variants to probe generality and found weaknesses in the models they tested. Both results can be true: substantial competence on one distribution and fragility when the structure or conventions change. Positive results; counterfactual tests.

Those studies concern particular models and tasks. They cannot settle what every current model can do. Their lasting contribution is a method: test whether an apparent principle survives meaningful changes to the problem.

For your use, I would turn that into four questions. Does the answer survive changing irrelevant names and surface details? Does it change when the decisive fact changes? Can the system identify which relationship supports its analogy? Can it recognize when that relationship fails? These questions test usable understanding more directly than asking the model whether it “really understands.”

4.4 Strange loops are not a synonym for an agent loop

Hofstadter's interests explicitly include consciousness and self-reference as well as analogy. Indiana University profile.

A program that reads its own output, edits its code, or maintains a description of its current task has a form of computational self-reference. That does not, by itself, establish the richer account of selfhood associated with I Am a Strange Loop. Nor does a fluent first-person explanation reveal whether there is subjective experience.

My assessment is that current systems make self-description and self-monitoring experimentally tractable while leaving the consciousness question unresolved. Neither the existence of a feedback loop nor the use of the word “I” settles it. For effective delegation, the nearer question is whether self-monitoring predicts and prevents errors under independent testing.

5. Returning to McClelland, the Churchlands, and Dennett

McClelland: the core connectionist wager largely survived

The broad wager—that cognition can depend on learned distributed representations rather than a complete hand-authored inventory of symbolic rules—has a strong engineering vindication. That conclusion concerns artificial systems' capabilities, not a demonstration that transformers reproduce the brain's learning algorithms.

An especially relevant older idea is complementary learning systems. McClelland, McNaughton, and O'Reilly explained why rapid storage of individual experiences and gradual integration into general knowledge can require different systems. Their account addressed the interaction between hippocampal and neocortical learning. 1995 paper.

There is a useful structural analogy to pretrained parameters plus a store of recent episodes or retrieved documents. The parameters support general abilities; external records preserve particular events without requiring immediate retraining. That is an analogy, not a claim that every retrieval system descended from this paper or that a document database functions biologically like a hippocampus. Ordinary retrieval also lacks the full consolidation process proposed in complementary learning systems.

The Churchlands: representational geometry became practically important

Paul Churchland argued for bringing neuroscience and computational accounts of the mind to bear on philosophical questions about knowledge and representation. Publisher's account of A Neurocomputational Perspective.

Patricia Churchland and Terrence Sejnowski's The Computational Brain addressed how groups of neurons support perception, decisions, and action, bringing abstract models into contact with neurobiological evidence. This matters for the comparison: explaining a capable artificial system and explaining the biological brain are related projects, but evidence for one does not automatically establish the other. The 1992 book.

The modern resonance is clear: systems can encode important distinctions in distributed numerical organization without storing an explicit sentence for every distinction. Their ability to use language does not imply that their internal representations are simply miniature written propositions.

But successful neural models do not establish eliminative materialism, settle the status of beliefs, or demonstrate a complete theory of human cognition. There is also a productive tension with the earlier ambition to explain intelligence through the brain: much recent AI progress came from computationally effective methods rather than close biological fidelity. The connectionist engineering program can succeed while stronger philosophical and neuroscientific claims remain open.

Dennett: competence can precede comprehension

Dennett repeatedly connected Darwin and Turing through the possibility that sophisticated competence can be built from processes that do not themselves understand what they are doing. Dennett's discussion in a recorded interview.

That perspective helps avoid two errors. First, simple local operations do not establish that the organized system lacks sophisticated abilities. Second, impressive behavior does not automatically establish the entire range of human comprehension, self-knowledge, or experience.

Applied to current AI, this suggests assessing explanations at the right level. A model can implement useful reasoning without individual nodes reasoning. A research agent can pursue a task through coordination among model calls, files, and tools without every component representing the whole project. Whether intentional language is helpful depends partly on how well it predicts behavior; it should not replace inspection of the mechanism when reliability matters.

My synthesis of these intellectual threads is that competence can be distributed across components and timescales. That makes the boundaries of “the thinker” less obvious than the chat interface suggests, while leaving empirical questions about robustness very much intact.

6. The harness is part of the working intelligence

Imagine the same capable model in three environments. In the first, it receives a question and must answer immediately. In the second, it can inspect relevant documents. In the third, it can inspect documents, maintain a structured account of findings, test factual assertions, compare alternative arguments, and revise the result.

Those are different problem-solving systems even if the model's weights are identical. The differences arise from evidence access, working memory, available actions, and feedback.

Anthropic distinguishes predetermined workflows from agents whose models decide more of the action sequence. Its engineering guidance emphasizes using the simplest arrangement that works. This is a useful corrective to treating more agents or more orchestration as automatic progress. Building effective agents.

Two later reports make the harness contribution concrete. A November 2025 report described improving continuity through explicit requirements, progress records, incremental work, and end-to-end tests. A March 2026 report described planner, generator, and evaluator roles, then reducing parts of the scaffolding as a stronger model made them less necessary. These are developer reports about particular application-building experiments, not universal effect-size estimates. Long-running agents, 2025; harness design, 2026.

The conclusion is more useful than “the harness matters more than the model”: the right harness depends on the task and on what the current model can already do reliably. A scaffold can compensate for a weakness, remain valuable at the edge of capability, or become unnecessary overhead.

What actually changes when an AI learns, remembers, or improves?

Layer What changes? Does it necessarily persist?
Training or fine-tuning Parameters shaping future behavior Yes, in the saved trained model.
Current conversation The context and temporary computational state Only while retained or supplied again.
Saved memory or retrieval External records later inserted into context Only if stored and successfully retrieved.
Tool use Information available to the model, or the external environment Depends on the tool and resulting artifact.
Candidate search Which proposed solution is retained The selected result can persist without changing model weights.
Harness revision Instructions, tools, orchestration, or evaluation procedure Yes, if the revised configuration is retained.

This table is an operational taxonomy, not a statement about every product implementation. Research such as Titans explicitly explores neural memory updated at test time, so the boundary between training and inference can be engineered differently. It would be a mistake to infer that an ordinary assistant is updating its core weights in real time merely because it remembers a correction within a conversation. Titans paper.

The practical question is always: where was the improvement stored, and how will it affect the next task? A corrected answer, a saved preference, a better model, and a better workflow are four different achievements.

7. Understanding, reasoning, and reliability are separate questions

“Next-token prediction” describes a training objective and a generation process. It does not, by itself, establish the complexity of the internal computation learned to perform that task. Conversely, plausible reasoning text does not establish that the computation was correct or that the explanation faithfully describes how the answer was produced.

For practical purposes, split understanding into observable abilities: applying a relation in a new setting; distinguishing a relevant change from an irrelevant one; identifying missing evidence; tracking consequences; and correcting an error when confronted with a decisive observation. These abilities can be tested individually without requiring a prior solution to consciousness.

Intermediate reasoning can help because earlier generated material becomes available for later computation. Chain-of-thought prompting demonstrated benefits on several reasoning benchmarks. But producing a longer explanation is not equivalent to performing a sound proof. Chain-of-thought study.

Self-correction is also conditional. One study found failures and regressions when models were simply asked to correct reasoning without external feedback. Another found intrinsic correction under different prompting and sampling conditions. The sensible conclusion is not that self-correction never works or always works; it is that we must distinguish tested correction procedures from generic exhortations to reconsider. Huang and colleagues; Liu and colleagues.

Similarly, a larger context window is capacity, not a guarantee of complete use. “Lost in the Middle” documented sensitivity to where relevant information appeared in the contexts of the models tested. It is historical evidence for a failure mode, not a current score for every long-context model. The operational response is to test retrieval and evidence coverage, rather than equating successful upload with successful comprehension. Long-context study.

The most consequential practical separation is therefore proposal versus verification. A generator can be imaginative and useful even when some proposals are wrong. A verifier can reject those errors if it has the relevant evidence and a dependable test. If both share the same blind spot, repetition can merely stabilize the error.

8. Engagement at the right level

The question to ask after an AI failure is not just “Which model should I use?” It is “Which part of the system was missing the capability, information, or feedback needed to succeed?”

Observed problem Likely place to investigate first A targeted intervention
The system gives the wrong factual premise. Evidence access and retrieval Supply or retrieve the operative source; require identification of the supporting passage.
It has the right material but misses the decisive relationship. Task framing and reasoning State the decision at issue; compare competing interpretations and their consequences.
It produces an elegant but unsupported account. Objective and evaluation Make support and contradiction checks part of completion, independently of style.
It loses earlier decisions during a long project. State preservation and context management Maintain a short authoritative record of adopted decisions, unresolved questions, and supporting files.
It repeats an ineffective action. Feedback and stopping rules Define what evidence would justify a retry, a different approach, or a stop.
It performs well on the example but fails on variants. Generalization Change surface details and then decisive assumptions separately.
It still fails after good evidence, clear framing, and checks. Model capability or task difficulty Try a stronger model, a narrower subproblem, or direct expert work.

This is my proposed diagnostic method, not a vendor's benchmark result.

For litigation work, I would make the unit of delegation a bounded decision or artifact with inspectable support. For example, asking for the strongest argument is useful, but asking for the strongest argument under specified facts, together with the evidence that could defeat it, gives the system a better-defined task. The aim is to create the conditions for useful reasoning, not to accumulate ritual prompt language.

An effective task specification tells the system the practical objective, which materials control, what it can do to gather additional evidence, what counts as completion, and which unresolved issues must remain visible. That is closer to designing a working environment than finding a magic sentence.

Evaluation deserves the same specificity. Anthropic's January 2026 account distinguishes evaluation components and emphasizes inspecting actual outcomes and traces of agent behavior. The general lesson is that a convincing final message is an insufficient measure of whether the external task succeeded. Agent evaluation guidance.

For an overnight project, autonomy is easiest when the system can recognize progress and failure without repeatedly asking you to supply the success criterion. The constraint is often not the number of hours available. It is whether each additional hour produces new evidence or merely more text.

9. What the five-part account leaves out

Several omitted developments are substantial. Convolution, recurrent memory architectures, optimization algorithms, normalization, tokenization, efficient attention, mixture-of-experts routing, distillation, and low-precision computation all matter. Some introduced new capabilities; others made powerful methods feasible. Their grouping into five transitions should not erase their inventors or technical significance.

Across AI as a whole, game-playing search, scientific structure prediction, generative image methods, and physical-world learning would require additional histories. This report deliberately explains the path to the language-based systems you can engage with as intellectual and practical collaborators.

There are also open questions the evidence here does not settle: how robustly learned abstractions transfer far outside familiar distributions; which reasoning improvements genuinely expand a model's solution repertoire rather than improve access to it; how to learn continually without destructive interference; how to evaluate outputs where the ground truth is contested or delayed; and what, if anything, current systems establish about consciousness.

My overall assessment is that there was no single missing philosophical insight that turned old networks into today's assistants. There was a sequence of changes that made representation learning scale, made training data abundant, made computation better organized, shaped behavior toward useful objectives, and connected generation to an environment capable of supplying corrective feedback.

Dawkins helps us ask what is retained and why. Hofstadter helps us ask whether the same relationship survives a change of representation. Connectionism helps explain how useful internal organization can be learned. The modern systems perspective asks what combination of model, memory, tools, and tests makes a whole task succeed. Together, those questions support a more precise form of engagement than treating the chat window as either a person or a passive reference book.

10. A short reading route

The report's links support the adjacent claims. For a focused return to the underlying work, I would read these in order:

  1. Rumelhart, Hinton, and Williams (1986): reconnect with learned hidden representations and credit assignment. Paper.
  2. Vaswani and colleagues (2017): read the introduction and architecture diagram before the mathematical details. The key question is how attention changes information flow and parallel training. Paper.
  3. Brown and colleagues (2020), then Hoffmann and colleagues (2022): separate broad transfer from the economics of model/data scale. GPT-3; Chinchilla.
  4. DeepSeek-R1 (2025): follow the training stages carefully, especially the difference between pretraining, R1-Zero, R1, and distillation. Paper.
  5. AlphaEvolve (2025): read the system description as a concrete meeting of generation, selection, and externally evaluated progress. Paper.
  6. Mitchell on Copycat, followed by the analogy debate: compare explicit construction of representations with learned performance and counterfactual testing. Copycat; positive LLM results; counterfactual evaluation.
  7. The 2026 harness report: look for which mechanisms still help and which became unnecessary as the model improved. Engineering report.

These are selected anchors rather than an exhaustive bibliography. A useful next experiment would hold a model fixed while changing only its evidence access, evaluation procedure, or available tools. That would let you observe directly which part of its apparent intelligence belongs to the model and which part belongs to the working arrangement.

Prepared September 6, 2026. Links identify sources supporting nearby claims. The five-development framework, comparisons, illustrative examples, and practical recommendations are analytical synthesis. The population experiment is a deliberately simplified teaching model.