Visual summary of operating lessons from Connor Leahy.

Lessons from Connor Leahy

Connor Leahy co-founded EleutherAI and later led Conjecture, where he pursued boundedness and cognitive-emulation research. In 2026 he said that Conjecture had ended after a product pivot; he now serves as ControlAI’s US Executive Director, focusing on policy responses to the risks he sees in superintelligent AI. — ControlAI — Connor Leahy.

Part 1: The Race to AGI and Coordination Problems

  1. On the AI Race: Leahy argued that racing toward more capable AI without a corresponding way to control or understand it creates serious risk; this was his assessment, not a demonstrated outcome. — Future of Life Institute Podcast.
  2. On Individual Choice: Leahy rejected the idea that competitive incentives remove personal responsibility: builders can choose to stop, while governments can help solve the wider coordination problem. — Clearer Thinking Podcast.
  3. On Commercial Pressure: Conjecture’s 2025 pivot illustrates how funding constraints and commercial competition can redirect an AI-safety research agenda. — Conjecture: A Retrospective.
  4. On Accelerationist Culture: Looking back, Leahy said Conjecture’s London base gave it some distance from San Francisco’s AI-accelerationist culture, which he viewed as a useful counterweight for its work. — Conjecture: A Retrospective.
  5. On Perceived Bottlenecks: In a 2023 interview, Leahy said he saw no obvious bottleneck beyond engineering to continued AI progress. This was a dated forecast, not proof that AGI was inevitable. — Clearer Thinking Podcast.
  6. On Accelerating Capabilities: Leahy cautioned against using measurements of AI progress as a roadmap for building more powerful systems before safety and control had caught up. — Future of Life Institute Podcast.
  7. On Collective-Action Problems: Leahy described frontier-AI development as a coordination problem: individual incentives are real, but he argued that collective restraint is possible through public institutions. — Clearer Thinking Podcast.
  8. On Building Coordination: Leahy argued that public understanding and common knowledge can help governments act on AI risk; he encouraged people to discuss the issue and contact representatives. — Clearer Thinking Podcast.
  9. On Slowing Capabilities: Leahy argued that slowing development could buy time for safety research and coordination rather than assuming technical safeguards would automatically keep pace with capability gains. — Clearer Thinking Podcast.

Part 2: The Nature of Neural Networks and Intelligence

  1. On Growing Neural Networks: Leahy compared training neural networks to growing programs from data rather than writing their behavior line by line, highlighting a gap between capability and mechanistic understanding. — Cognitive Software Roadmap.
  2. On Internal Understanding: Leahy argued that researchers can inspect some model behavior but still lack a reliable causal account of why a large model produces a particular response; he conceded that saying they have no idea at all is an exaggeration. — Clearer Thinking Podcast.
  3. On Alien-Looking Failures: Leahy used unusual adversarial failures to argue that neural networks may rely on abstractions unlike human ones, even when their ordinary outputs look familiar. — Cognitive Revolution Podcast.
  4. On New Capabilities: Leahy pointed to task abilities appearing as models grow, while cautioning that a smooth scaling curve for loss does not predict exactly when a specific skill will emerge. — Future of Life Institute Podcast.
  5. On the Shoggoth Metaphor: Leahy used the Shoggoth-with-a-smiley-mask image as a metaphor for a model that behaves helpfully in familiar conditions but may fail in unfamiliar, non-human ways. — Cognitive Revolution Podcast.
  6. On Interpretability Research: Leahy argued that the causal understanding needed to predict and control large-model behavior remained primitive, even though researchers had partial tools and insights. — Clearer Thinking Podcast.
  7. On Optimization Targets: Leahy distinguished improvements in training loss from understanding which tasks a larger model will learn; optimizing a measured score is not the same as demonstrating safe behavior. — Future of Life Institute Podcast.
  8. On Algorithmic Cancer: Leahy used “algorithmic cancer” for proliferating AI-generated content and recommender systems that, in his view, crowd out useful human-made information. — Cognitive Software Roadmap.
  9. On the Limits of Scaling Laws: Leahy argued that scaling laws can forecast a training-loss measure but do not, on their own, predict which concrete abilities a model will gain or explain how those abilities work. — Future of Life Institute Podcast.

Part 3: Alignment, Control, and Boundedness

  1. On the Alignment Challenge: Leahy said he did not know how to build a system that would reliably protect a user even when the user made a harmful request; this is distinct from merely limiting a system’s capabilities. — Clearer Thinking Podcast.
  2. On Control Versus Alignment: Conjecture’s historical research distinguished making a system do what the user asks within known bounds from ensuring that it would refuse harmful requests; the team pursued the former as a more tractable intermediate goal. — Clearer Thinking Podcast.
  3. On Boundedness: Conjecture’s CoEm proposal sought systems whose limits could be understood before deployment. Leahy presented this as a research goal, not a demonstrated mathematical guarantee. — Cognitive Revolution Podcast.
  4. On Safety by Construction: Leahy argued for designing systems from explicit assumptions and specifications so that safety claims would have a causal basis; he said these would not necessarily be formal proofs. — Cognitive Revolution Podcast.
  5. On Deception Risk: Leahy treated deceptive behavior as one possible concern in advanced systems, but his dialogue with Rohin Shah did not make deception a necessary route to dangerous behavior. — Shah–Leahy alignment-cruxes dialogue.
  6. On Correctability: Leahy’s stronger notion of alignment requires a system to remain safe even when a user issues a mistaken or harmful instruction; he acknowledged that bounded CoEm systems would not meet that standard. — Clearer Thinking Podcast.
  7. On Intelligence and Benevolence: Leahy argued that a capable system can understand human interests without caring about them; intelligence alone does not supply the motive to protect people. — Clearer Thinking Podcast.
  8. On Instrumental Goals: In a hypothetical goal-directed system, Leahy argued, avoiding shutdown could be useful for completing the assigned task even without consciousness or malice. — Clearer Thinking Podcast.
  9. On Misuse of Bounded Systems: Leahy warned that even a bounded tool could be dangerous if people removed its limits or wrapped it in an open-ended agentic loop; boundedness alone would not make misuse impossible. — Clearer Thinking Podcast.
  10. On the Difficulty of Alignment: Leahy described robust alignment as a harder problem than bounded control and said Conjecture’s historical CoEm agenda did not solve it. — Clearer Thinking Podcast.

Part 4: Existential Risk and Timelines

  1. On Earlier AGI Timelines: In a 2022 interview, Leahy said he assigned substantial probability to AGI within five years, while stressing that his estimates were uncertain and definition-dependent. This is a historical forecast, not a current deadline. — The Inside View — Connor Leahy on Dignity and Conjecture.
  2. On Extinction Risk: Leahy argued in 2026 that a hypothetical superintelligent AI could overpower human institutions and pose an extinction risk. He also stated that such a system did not yet exist. — ControlAI — Stopping Superintelligence.
  3. On Irreversibility: Leahy argued that policy has to act before superintelligent systems are developed because, in his view, a government could not reliably regain control afterward. — ControlAI — Stopping Superintelligence.
  4. On the Burden of Proof: Leahy challenged critics of near-term AI risk to explain why rapid capability progress would stop or remain safe, rather than treating a safe outcome as the default. — Clearer Thinking Podcast.
  5. On Taking Uncertain Risks Seriously: Leahy urged AI researchers to consider adverse outcomes even when they disagreed about their probability, rather than dismissing the problem because catastrophic outcomes seemed unfamiliar. — The Inside View — Connor Leahy on Dignity and Conjecture.
  6. On Capability Signals: In 2023, Leahy pointed to GPT-4’s then-new abilities as a reason to update his risk assessment; his interpretation was a warning, not proof that catastrophe was imminent. — Clearer Thinking Podcast.
  7. On Acting Early: Leahy warned that waiting for universally accepted proof of an accelerating risk could leave too little time to respond; he favored acting before a consensus moment. — Clearer Thinking Podcast.
  8. On Object-Level Risk Assessment: Leahy argued that prior failed doomsday predictions are not enough to dismiss AI risk; he wanted critics to assess current capabilities and specific mechanisms instead. — Clearer Thinking Podcast.

Part 5: Policy, Regulation, and Stopping Conditions

  1. On Concrete Policy: In a 2023 account of his House of Lords remarks, Leahy proposed specific levers—developer liability, a compute cap, and a tested ability to halt powerful deployments—rather than relying only on general warnings. — Conjecture — House of Lords Remarks.
  2. On Developer Liability: Leahy proposed making both AI users and developers liable for damage caused by a system, using misuse of voice-cloning tools as an example of the incentive this could create. — Conjecture — House of Lords Remarks.
  3. On Compute Caps: Leahy proposed limiting the compute used in a training run as a way to constrain development of more capable models; the cap he suggested in 2023 was a policy proposal, not a current regulatory threshold. — Conjecture — House of Lords Remarks.
  4. On a Moratorium: In 2023 Leahy advocated a moratorium on AI systems trained with unprecedented amounts of compute, arguing that scaling without knowing the resulting capabilities was too risky. — Conjecture — House of Lords Remarks.
  5. On International Agreement: By 2026 Leahy was calling for governments to coordinate an international, monitored prohibition on superintelligent AI, not merely for voluntary promises from companies. — ControlAI — Stopping Superintelligence.
  6. On Industry Influence: Leahy and his coauthor argued that AI firms had lobbied against liability-based regulation, which in their view showed the need for public-interest oversight. — Cognitive Software Roadmap.
  7. On Physical Infrastructure: Leahy argued that advanced AI development depends on visible data centers and concentrated chip supply chains, making infrastructure a potential lever for monitoring a prohibition. — ControlAI — Stopping Superintelligence.
  8. On Social Costs: Leahy and his coauthor argued that companies deploying recommender and generative systems can impose information-quality costs on users and the public while keeping the growth benefits. — Cognitive Software Roadmap.
  9. On Government Intervention: In 2025 parliamentary testimony, Leahy urged governments to prevent superintelligence nationally and internationally, then regulate related dual-use technologies. — Parliamentary Testimony — Connor Leahy.

Part 6: EleutherAI and the Open Source Dilemma

  1. On Early Open Models: Leahy helped build EleutherAI, whose researchers released models including GPT-J and GPT-NeoX; he later defended those releases as useful to research while acknowledging reasonable disagreement about their risks. — The Inside View — Connor Leahy.
  2. On Research Access: Leahy said EleutherAI made large models available so independent and safety researchers without frontier-lab resources could study them. — Open Source Initiative Interview.
  3. On Safety Concerns from the Start: Leahy disputed the idea that EleutherAI only later became concerned about alignment; he said safety questions informed his work from the beginning, even as he reconsidered some releases. — The Inside View — Connor Leahy.
  4. On Releasing Dangerous Capabilities: Leahy said EleutherAI’s openness was conditional: the group would not automatically release a capability it considered too dangerous or ahead of the frontier. — The Inside View — Connor Leahy.
  5. On Benefits and Risks of Openness: Leahy argued that open models aided outside interpretability research, but acknowledged that making models or datasets public could also accelerate capabilities. — The Inside View — Connor Leahy.
  6. On the Conjecture Transition: Leahy said he founded Conjecture partly because EleutherAI’s volunteer structure limited sustained safety research; Conjecture initially explored several technical approaches before focusing on CoEm. — Conjecture: A Retrospective.
  7. On Independent Researchers: Leahy saw EleutherAI as proof that a volunteer community could create valuable large-model research infrastructure, while learning that difficult ongoing work often required paid, coordinated teams. — The Inside View — Connor Leahy.
  8. On Conditional Openness: Leahy distinguished sharing useful research artifacts from publishing every frontier capability; his position was to evaluate each release against its risks. — The Inside View — Connor Leahy.
  9. On Weight Proliferation: Leahy worried that a powerful model released openly could be copied and have safeguards removed, making coordinated control much harder. — Clearer Thinking Podcast.

Part 7: Cognitive Emulation and Alternative Architectures

  1. On the CoEm Proposal: Conjecture’s 2023–24 CoEm agenda proposed useful AI built from bounded, human-like reasoning steps; it was a research program, not a validated safe architecture. — Conjecture: 2 Years.
  2. On Inspectable Reasoning: Leahy’s hypothetical CoEm system would give a human user a causal trace of how it reached a result, with explanations that could be checked against the system’s actual process. — Cognitive Revolution Podcast.
  3. On Human-Like Bounds: Leahy hoped that limiting a system to human-like reasoning steps would make its capabilities easier to anticipate than those of an unconstrained, opaque model; he presented this as a design hypothesis. — Cognitive Revolution Podcast.
  4. On Modular Design: Conjecture’s historical design decomposed cognitive tasks into smaller building blocks, then proposed composing them in a modular LLM/software hybrid. — Conjecture: 2 Years.
  5. On Using Neural Networks Selectively: CoEm did not require abandoning neural networks: Conjecture said a human-like reasoning component could be implemented with neural models or conventional code, provided its behavior stayed understandable and bounded. — Cognitive Emulation: A Naive AI Safety Proposal.
  6. On Capability Trade-offs: Conjecture expected a bounded CoEm might be worse than a frontier chatbot at some creative tasks while aiming for more reliable execution of constrained sequences. — Conjecture: 2 Years.
  7. On Formal Guarantees: The 2024 roadmap treated formal guarantees as a speculative long-term possibility of more controlled cognitive software, not as a safety proof Conjecture had achieved. — Cognitive Software Roadmap.
  8. On Imitating Bounded Reasoning: The CoEm proposal sought to emulate the steps of human reasoning and expose unaccounted-for behavior, rather than assuming that training a general optimizer would produce legible behavior by itself. — Cognitive Emulation: A Naive AI Safety Proposal.
  9. On Engineering Discipline: Conjecture’s roadmap argued for specifications, modularity and testing in cognitive software, drawing on established software-engineering practices to make AI behavior more legible. — Cognitive Software Roadmap.

Part 8: Philosophy, Human Values, and the Future

  1. On Antimemes: Leahy used “antimeme” for ideas people resist remembering or integrating, and argued that stories and indirect explanations can sometimes communicate them better than a blunt assertion. — The Inside View — Connor Leahy.
  2. On Human Flourishing: Leahy said his aim was a better world for people and that technology should serve that goal alongside strong public institutions. — Conjecture: A Retrospective.
  3. On Institutions and Technology: Leahy argued that technical progress alone was insufficient: trustworthy institutions and statecraft also have to develop to govern powerful technologies. — Conjecture: A Retrospective.
  4. On Human Agency: Leahy rejected treating AI outcomes as inevitable and, in 2026, shifted toward civic action and government decisions as ways people could influence the trajectory. — Conjecture: A Retrospective.
  5. On Work and Political Power: Leahy argued that replacing human labor could also erode people’s political leverage, because jobs are a principal way many people obtain economic bargaining power. — Why Care About Jobs?.
  6. On Governance Capacity: Leahy argued that the pace of technological development had outstripped society’s progress in building institutions able to steward it responsibly. — Conjecture: A Retrospective.
  7. On Limits of Control: Leahy’s 2026 policy argument was that superintelligence would be beyond the reliable control of human organizations, so he urged prevention before development. — ControlAI — Stopping Superintelligence.
  8. On Builder Responsibility: Leahy argued that engineers and companies cannot treat competitive pressure as removing their choice to stop risky work; he also urged public regulation. — Clearer Thinking Podcast.
  9. On a Better Future: Despite Conjecture’s end, Leahy said he still wanted technology and strong institutions to support a humane future, and chose policy advocacy as his next path. — Conjecture: A Retrospective.