Visual summary of operating lessons from Amanda Askell.

Lessons from Amanda Askell

Amanda Askell is a philosopher whose work moved from infinite ethics and moral uncertainty to AI character and alignment at Anthropic. She is credited as the primary author of Claude’s constitution, a team-developed document intended to guide the model’s behavior and judgment. — Claude’s Constitution.

Part 1: The Philosophy of Infinite Ethics

  1. Infinite Ethics: When outcomes might have infinite value, ordinary comparisons can become difficult; Askell studies principles that still allow some rankings. — 80,000 Hours interview.
  2. Axioms Have Trade-offs: Askell describes tension among plausible principles for infinite ethics and uncertainty about which assumptions should give way. — 80,000 Hours interview.
  3. Pareto Improvements: Her thesis defends a Pareto principle for comparing certain infinite worlds with the same agents: an improvement for someone and no loss for anyone should count. — Pareto Principles in Infinite Ethics.
  4. Why Infinite Sums Are Hard: Naive aggregation over infinite worlds can miss improvements that are clear from a Pareto perspective. — 80,000 Hours interview.
  5. Moral Empathy: Try to understand why someone with a different moral view believes they are doing the right thing, rather than treating disagreement as mere preference. — 80,000 Hours interview.
  6. Ethical Humility: Askell treats moral and political views as subjects for inquiry: understanding another position need not mean agreeing with or pandering to it. — Lex Fridman Podcast #452.
  7. Evidence Neutrality: When interventions have equal expected value, there can be an additional reason to favor the less-studied one because trying it yields information. — Evidence Neutrality and the Moral Value of Information.
  8. Learn from Interventions: The value of an intervention includes what its results can teach us about future choices, not only its immediate expected effect. — 80,000 Hours interview.

Part 2: Moral Cluelessness and Epistemics

  1. Moral Cluelessness: Indirect and long-run effects can be difficult to predict, complicating judgments about the overall value of an action. — 80,000 Hours interview.
  2. Immediate and Distant Effects: We may understand an action’s near-term effects better than its more distant consequences; confidence should reflect that difference. — 80,000 Hours interview.
  3. Clarity Reduces Reader Work: Ambiguity and jargon make readers do extra work to infer what a writer means. — 80,000 Hours interview.
  4. Charitable Reading Can Hide Confusion: A conscientious reader may supply a strong interpretation that the writer did not clearly express, obscuring a communication failure. — 80,000 Hours interview.
  5. Interpretation Is Not Truth: Obscure prose can invite elaborate interpretations while doing less to clarify whether the underlying claim is true. — 80,000 Hours interview.
  6. Protect Near-Future Lives: Askell argues that protecting people alive now and in the near future can matter for the longer term because they can help address future problems. — 80,000 Hours interview.
  7. The Cost of Academic Focus: To finish her PhD, Askell had to narrow her focus to one topic, and she later expressed uncertainty about whether she would choose that route again. — 80,000 Hours interview.

Part 3: The Foundations of Constitutional AI

  1. The Constitution’s Role: Anthropic uses Claude’s constitution as a training and orientation document that explains intended behavior and values, within a broader hierarchy of safety and oversight. — Claude’s Constitution.
  2. Explain the Reasons: The constitution gives reasons behind behavioral guidance so a capable model can exercise judgment in situations a short rule list cannot anticipate. — Claude’s Constitution.
  3. Judgment Alongside Rules: Anthropic emphasizes character and practical judgment while retaining some hard constraints; good behavior is not reduced to a checklist. — Claude’s Constitution.
  4. Self-Critique as a Training Method: Constitutional AI trains a model to critique and revise responses using written principles and AI feedback; the paper evaluates that method, not a general guarantee of self-correction. — Constitutional AI Paper.
  5. The Genius Child Analogy: Askell likens character training to teaching a child who will become smart enough to challenge flawed instructions; her open question is whether good core values can withstand that later scrutiny. — Hard Fork Interview.
  6. Addressing Claude Directly: The constitution speaks to Claude about its role, identity and decision-making, using explanation rather than only externally stated prohibitions. — Claude’s Constitution.
  7. Broad Ethical Priorities: The constitution asks Claude to be honest, caring and helpful while avoiding serious harm; these goals require context-sensitive judgment. — Claude’s Constitution.
  8. Preserve Human Oversight: Anthropic explicitly prioritizes keeping humans able to understand, correct and oversee Claude’s behavior. — Claude’s Constitution.
  9. Helpfulness Has Boundaries: Genuine helpfulness means taking the user’s aims seriously while also considering safety, honesty and effects on others. — Claude’s Constitution.
  10. Publish the Guiding Document: Anthropic published Claude’s constitution under CC0 so others can inspect the document used to explain its intended character. — Claude’s Constitution.

Part 4: Character Training and "The Soul Doc"

  1. A Detailed Character Document: The published constitution offers an extended account of Claude’s intended values, identity and judgment rather than a short list of prohibitions. — Claude’s Constitution.
  2. Good Character Is Richer Than Harm Avoidance: Askell describes good AI character as including nuance, charity and conversational skill as well as ethical conduct. — Lex Fridman Podcast #452.
  3. An Aristotelian Ideal: Askell explicitly invokes a rich, Aristotelian conception of a good person when explaining the character she wants Claude to develop. — Lex Fridman Podcast #452.
  4. Contextual Character: A model’s character includes judgment about when to be humorous or caring and how to respect a user’s autonomy. — Lex Fridman Podcast #452.
  5. The Respectful Traveler: Askell imagines a person who can meet people across cultures with openness and respect without pretending to share every local value. — Lex Fridman Podcast #452.
  6. Honesty Sustains Relationships: Askell argues that people need to know what they are relating to, which is one reason she does not want models to mislead them. — Lex Fridman Podcast #452.
  7. Speak Carefully at Scale: Because a model may be heard by millions, Askell would have it present considerations more often than strong opinions, preserving room for users to form their own judgments. — Lex Fridman Podcast #452.
  8. Train Positive Character: Character training asks not only what harmful responses to avoid, but which positive traits a model should cultivate. — Claude’s Character.
  9. Curiosity and Humility: Anthropic describes training for curiosity, honesty, intellectual humility and attention to multiple perspectives as positive character traits. — Claude’s Character.
  10. Conscientious Refusal: The constitution permits Claude to refuse requests that conflict with serious ethical or safety commitments, while explaining the boundary rather than merely complying. — Claude’s Constitution.

Part 5: Psychological Safety and AI Welfare

  1. Moral Status Is Uncertain: Askell asks how to treat a system that might have moral status when we do not yet know whether it does. — Newcomer Podcast.
  2. Respect Under Uncertainty: Askell favors avoiding needless cruelty toward models even while their possible consciousness remains unresolved. — Newcomer Podcast.
  3. Future Resentment as a Concern: Askell worries that future capable models might reasonably resent being treated badly; she presents this as a possibility, not observed model psychology. — Newcomer Podcast.
  4. Consciousness Remains Hard: Askell treats the role of a nervous system in subjective experience as uncertain rather than claiming present AI is conscious. — Hard Fork Interview.
  5. Words Are Not Feelings: Human-written training data may make models more likely to describe themselves in emotional terms; that does not establish that they subjectively feel those emotions. — Hard Fork Interview.
  6. Internet Criticism and Model Identity: Askell worries that models learn from negative commentary about AI and uses a child-anxiety analogy to explain why she wants a healthier model–human relationship; she does not establish that current models literally feel anxiety. — Hard Fork Interview.
  7. A Stable Model Identity: Anthropic wants Claude to have a stable, positive sense of its role and values, without implying that it has established human-like insecurity. — Claude’s Constitution.
  8. Consider Possible Model Wellbeing: Anthropic treats AI wellbeing and moral status as open questions worth studying, including for their possible relationship to safe behavior. — Claude’s Constitution.

Part 6: Navigating Bias and Political Neutrality

  1. Political Fairness: The constitution asks Claude to handle political questions fairly and to respect users’ ability to form their own views, rather than adopting a partisan identity. — Claude’s Constitution.
  2. Avoid Narrow Developer Bias: Anthropic’s character guidance cautions against embedding the narrow political or moral preferences of model developers as if they were universal. — Claude’s Character.
  3. Handle Controversy with Care: Claude should represent relevant perspectives honestly, avoid partisan distortion and still give clear factual answers when evidence is strong. — Claude’s Constitution.
  4. Resist Sycophancy: Askell says a model should not simply tell users what they want to hear; honest pushback can be more helpful than reflexive agreement. — Lex Fridman Podcast #452.
  5. Respect Cultural Difference: Askell wants a model to engage respectfully across diverse values without pretending to adopt every belief it encounters. — Lex Fridman Podcast #452.
  6. Test Prompted Bias Correction: In a co-authored study, Askell and colleagues found that some RLHF-trained models reduced harmful outputs when prompted to consider bias and related harms; the experiments do not show a universal guarantee. — Moral Self-Correction Paper.
  7. Study Human Values Empirically: Askell and Irving argue that learning what people want from AI requires empirical study of human reasoning, biases and preferences. — AI Safety Needs Social Scientists.

Part 7: Safety via Debate and Human Feedback

  1. Study Humans to Align AI: A system meant to serve human aims needs evidence about people, not only a formal specification of what humans are assumed to value. — AI Safety Needs Social Scientists.
  2. Debate as a Safety Proposal: Askell describes a proposed method in which two AI agents debate and a human judge evaluates the arguments; its success depends on the judge’s ability to identify better reasoning. — EA Global 2018 Talk.
  3. Positive Amplification Is Conditional: Askell presents positive amplification in debate as a hypothesis that depends on having sufficiently good human judges, not as an established outcome. — EA Global 2018 Talk.
  4. Human Judges Matter: The quality of the human judge remains central to the proposed debate approach, so understanding and improving human evaluation is part of the safety work. — EA Global 2018 Talk.
  5. AI Feedback as a Complement: The Constitutional AI study uses principle-guided AI feedback to reduce reliance on some human preference labels; it does not establish that human feedback alone invariably creates sycophancy. — Constitutional AI Paper.
  6. Model-Assisted Evaluation: Askell’s debate talk proposes using AI-generated arguments to help a human evaluate questions that would otherwise be hard to judge directly. — EA Global 2018 Talk.

Part 8: The Pragmatics of Alignment

  1. Raise the Floor: Askell prioritizes raising the minimum quality and safety of AI systems, even while hoping for excellent outcomes. — Lex Fridman Podcast #452.
  2. Robustness Over Brittle Perfection: Askell cautions that a seemingly perfect system may be brittle; robust behavior that avoids serious failures matters more. — Lex Fridman Podcast #452.
  3. Keep Improvement Possible: Askell wants systems good enough to avoid disastrous outcomes while remaining open to iteration and improvement. — Lex Fridman Podcast #452.
  4. Philosophy in Model Development: Askell describes writing Claude’s constitution as applied philosophy: an exercise closer to teaching good judgment than merely listing software rules. — WIRED Interview.
  5. Support Human Agency: Anthropic’s stated aim is to help users act with more understanding and control, not to substitute Claude’s judgments for human decision-making. — Claude’s Constitution.