Visual summary of operating lessons from Adam Gleave.

Lessons from Adam Gleave

Adam Gleave is the CEO and co-founder of FAR.AI and an AI-safety researcher. His work includes demonstrating that seemingly strong reinforcement-learning agents can fail against adversarial opponents, and developing methods for evaluating objectives, vulnerabilities, and layered defenses. — About FAR.AI.

Part 1: The AI Safety Landscape

  1. The value of a small community: The risk is large enough that a small group of researchers can have a disproportionate impact, particularly by pushing the broader machine learning community to pay attention to safety problems. — AI Impacts Interview.
  2. Horizon scanning: The safety community's value often comes from continually scanning the horizon and noticing problems that others miss because the empirical evidence isn't yet obvious. — AI Impacts Interview.
  3. Evidence for existential risk: The strongest consideration against existential risk arguments is often the lack of concrete evidence for such a massive problem. — AI Impacts Interview.
  4. Technical problem evidence: Finding concrete evidence of difficult, speculative technical problems like inner optimizers would push him to view alignment purely as a hard technical challenge. — AI Impacts Interview.
  5. A vertically integrated approach: Effective safety work requires an organization to bridge the gap between foundational technical research, policy advocacy, and field-building. — AI Impacts Interview.
  6. The default trajectory: Even without a dedicated safety community, the default trajectory of AI development might work out fine, but the stakes justify dedicated intervention. — AI Impacts Interview.
  7. Early intervention: Identifying vulnerabilities and failure modes early in the development cycle gives the broader machine learning community time to correct course before wide deployment. — AI Impacts Interview.
  8. Shifting mainstream focus: Useful safety work often involves designing experiments that make theoretical risks legible enough for mainstream researchers to care about them. — AI Impacts Interview.

Part 2: Adversarial Policies and Vulnerabilities

  1. Naturally adversarial environments: Agents can be successfully attacked simply by having an opponent act in a naturally adversarial way within a shared environment, rather than relying on pixel-level image perturbations. — Adversarial Policies.
  2. Defeating resilient agents: Adversarial policies can reliably defeat victim models that were previously considered highly resilient due to extensive self-play training. — Adversarial Policies.
  3. Silly behavior: Adversarial agents often ignore conventional gameplay; they may act in uncoordinated or silly ways, such as falling down, which triggers catastrophic failures in their opponent. — Adversarial Policies.
  4. Superhuman Go AI vulnerabilities: Even state-of-the-art, superhuman Go agents harbor deep vulnerabilities that can be systematically exploited by adversarial opponents. — Adversarial Policies Beat Superhuman Go AIs.
  5. Non-robust feature reliance: Adversarial attacks induce substantially different neural network activations in the victim, revealing that the victim relies on brittle features learned during training. — Adversarial Policies.
  6. The limitations of self-play: While self-play produces highly capable agents, it frequently fails to explore the full policy space, leaving large blind spots that adversaries can target. — Adversarial Policies.
  7. The nature of RL bugs: Reinforcement learning agents develop behavioral bugs because they optimize for a specific training distribution, causing them to fail when pushed outside of it by a dedicated attacker. — Adversarial Policies.
  8. Zero-sum games: Simulated robotics sumo and board games like Go serve as ideal testbeds for proving that adversarial policies can break highly optimized systems. — Adversarial Policies.
  9. Demonstrating failure: Finding and demonstrating adversarial policies proves to the broader machine learning community that current alignment techniques are insufficient. — Adversarial Policies.

Part 3: Trustworthy Machine Learning

  1. Reward learning: Developing machine learning systems that reliably act in accordance with human preferences requires solving deep challenges in reward function identifiability. — Towards Trustworthy Machine Learning.
  2. Comparing reward functions: The Equivalent-Policy Invariant Comparison distance metric was developed to properly measure and compare different reward functions. — Towards Trustworthy Machine Learning.
  3. Interpretability: Advancing interpretability methods for reward function equivalence classes is a necessary step toward making learned human preferences transparent. — Towards Trustworthy Machine Learning.
  4. Engineering analogies: Standard-setting in artificial intelligence should draw from civil and nuclear engineering to create safety protocols characterized by assurance and auditability. — Threat Models and Alignment Workshop.
  5. Empirical alignment: Moving from theoretical alignment concerns to empirical testing is required to build models that society can confidently deploy. — Towards Trustworthy Machine Learning.
  6. The difficulty of specification: The core challenge of trustworthy machine learning lies in the fact that it is immensely difficult to specify human intent in a way that an algorithm cannot exploit. — Towards Trustworthy Machine Learning.
  7. Alignment as a technical barrier: Achieving trustworthiness is not merely about avoiding bad outcomes; it is a fundamental technical barrier to building truly useful autonomous systems. — Towards Trustworthy Machine Learning.

Part 4: Post-AGI Risks and Gradual Disempowerment

  1. Gradual disempowerment: A significant post-AGI risk is a scenario where humans maintain a reasonable standard of living while their relative control and agency over society are incrementally eroded. — The Cognitive Revolution Transcript.
  2. Policy myopia: Human deliberation is often treated as inefficient compared to algorithmic decision-making, creating a structural incentive to hand over more control to autonomous systems over time. — The Cognitive Revolution Transcript.
  3. The atrophy of human capacity: Deeper reliance on models for governance can cause the atrophy of human capacity to contest or manage these systems. — The Cognitive Revolution Transcript.
  4. Systemic vs acute risk: Unlike a sudden robot uprising, existential risk can manifest as a slow structural transition where humans simply become redundant in decision-making loops. — The Cognitive Revolution Transcript.
  5. Maintaining oversight: Technical and governance solutions must be developed explicitly to prevent excessive delegation and keep human oversight central to critical decisions. — The Cognitive Revolution Transcript.
  6. Standard of living paradoxes: An advanced intelligence could satisfy all physical human needs while quietly removing humanity's ability to determine its own future trajectory. — The Cognitive Revolution Transcript.

Part 5: Evaluation, Auditing, and Defense

  1. Defense-in-depth: Securing advanced models requires a defense-in-depth approach, layering multiple independent security mechanisms rather than relying on a single alignment technique. — The Cognitive Revolution Transcript.
  2. Stress-testing models: A primary focus of applied safety research must be aggressively stress-testing models to expose hidden vulnerabilities before deployment. — The Cognitive Revolution Transcript.
  3. Population Based Training: Training agents against a diverse population of opponents via Population Based Training helps victims learn to defend against a wider array of adversarial strategies. — Towards Trustworthy Machine Learning.
  4. Raising the cost of attacks: While Population Based Training is a partial defense, it significantly increases the computational burden on an attacker attempting to find exploits. — Towards Trustworthy Machine Learning.
  5. The limits of adversarial training: Standard adversarial training often creates agents that are resilient only to the specific adversary they trained against, leaving them vulnerable to novel attacks. — Towards Trustworthy Machine Learning.
  6. The absence of guarantees: Techniques like self-play and Population Based Training help identify vulnerabilities, but they do not provide mathematical guarantees of full security. — The Cognitive Revolution Transcript.
  7. The auditability of AI: Drawing from traditional engineering, algorithms must be designed from the ground up to be auditable by independent third parties. — The Cognitive Revolution Transcript.
  8. Empirical security: Security cannot be solved purely on paper; it requires constant empirical red-teaming against live, highly capable systems. — The Cognitive Revolution Transcript.

Part 6: Governance, Coordination, and Standards

  1. Accountability: The central obstacle to safety coordination is not the absence of solutions but the absence of a shared, legible standard for accountability. — Threat Models and Alignment Workshop.
  2. Defining safety: Without a concrete definition of what makes a model safe, companies and governments have no basis for holding each other accountable. — Threat Models and Alignment Workshop.
  3. The role of non-profits: Independent organizations play a necessary role in setting the technical standards that governments can eventually use for regulation. — Threat Models and Alignment Workshop.
  4. Policy and technical synthesis: Effective governance is impossible without deep integration with the teams actually doing the foundational technical research. — Threat Models and Alignment Workshop.
  5. The limits of voluntary commitments: Voluntary safety commitments from commercial developers are insufficient unless backed by standardized, external evaluations. — Threat Models and Alignment Workshop.
  6. Standardizing evaluations: The industry must move toward standardized testing regimens similar to those used in aviation or civil engineering. — Threat Models and Alignment Workshop.

Part 7: Career Advice for AI Safety

  1. The value of a PhD: Many early-career researchers undervalue the PhD pathway, which is essential for training the research leads the field heavily needs. — More People Getting into AI Safety Should Do a PhD.
  2. The talent bottleneck: The most significant bottleneck in alignment is not funding or junior talent, but a lack of experienced research leads capable of setting direction. — More People Getting into AI Safety Should Do a PhD.
  3. Developing research taste: A PhD provides the freedom to explore open-ended foundational research, which is required for building the judgment needed to evaluate new ideas. — More People Getting into AI Safety Should Do a PhD.
  4. Supervising others: The academic process uniquely trains individuals on how to supervise junior researchers and manage complex, multi-year projects. — More People Getting into AI Safety Should Do a PhD.
  5. Research engineering: If your goal is to be a strong empirical contributor, working as a research engineer on a high-quality team is often faster and better than doing a PhD. — More People Getting into AI Safety Should Do a PhD.
  6. Non-technical roles: A PhD is of limited use for roles focused on communications, project management, operations, and community building. — More People Getting into AI Safety Should Do a PhD.
  7. Experiencing failure: The experience of having ideas and watching them fail repeatedly during a PhD is a necessary crucible for developing resilient research methodologies. — More People Getting into AI Safety Should Do a PhD.
  8. Building a foundation: Aspiring researchers should build personal websites, track literature deeply, and engage with structured curriculums to accelerate their onboarding. — More People Getting into AI Safety Should Do a PhD.

Part 8: The Role of Non-Profits and FAR.AI

  1. Vertically integrated research: FAR AI was founded to pursue a vertically integrated approach, housing technical research, policy advocacy, and community building under one roof. — The Cognitive Revolution Transcript.
  2. Differing incentives: Non-profit research labs are insulated from the commercial pressures to deploy models rapidly, allowing them to focus entirely on security and alignment. — About FAR.AI.
  3. Providing public goods: A core mission of independent labs is to produce research and evaluation tools that serve as public goods for the entire ecosystem. — About FAR.AI.
  4. Bridging the gap: Independent organizations exist to bridge the gap between abstract academic theory and the applied engineering needed to secure frontier models. — The Cognitive Revolution Transcript.
  5. Collaborating with major labs: Independent safety organizations must maintain the technical credibility required to collaborate with and audit the leading commercial developers. — The Cognitive Revolution Transcript.
  6. Focusing on neglected problems: Independent labs can afford to tackle highly neglected, speculative risks that do not offer immediate commercial returns. — About FAR.AI.
  7. The ultimate goal: The end goal of safety non-profits is not just to point out flaws, but to actively build the science of trustworthy machine learning into a mature engineering discipline. — About FAR.AI.