Aleksander Mądry is MIT’s Cadence Design Systems Professor of Computing and directs the Center for Deployable Machine Learning. His research examines reliable machine learning, adversarial robustness, and evaluation for real-world use. — MIT Faculty Profile.

Visual summary of operating lessons from Aleksander Madry.

Part 1: The Illusion of Understanding

  1. The Learning Objective: A model can score well by exploiting predictive patterns without learning the human concept that the test was meant to probe. — TED Talk.
  2. Passing the Test: Madry’s TED talk asks whether an AI system has learned a task or found a way to pass its evaluation. — TED Talk.
  3. Human and Model Features: The features useful to a classifier need not coincide with those humans perceive as meaningful. — Adversarial Examples Are Features.
  4. Benchmark Accuracy: A high benchmark score does not by itself settle whether a model will perform the broader task outside that benchmark’s construction. — ImageNet Benchmark Study.
  5. Evaluate for Deployment: The MIT Deployable Machine Learning center studies reliability under random and adversarial corruptions alongside understandability and safe deployment. — MIT Deployable ML Center.
  6. Predictive Shortcuts: Non-robust features can be genuinely predictive in a dataset, giving standard training an incentive to use them. — Adversarial Examples Are Features.
  7. Optimization and Features: Ordinary training rewards predictive accuracy; it need not favor the same features that people find robust or interpretable. — Adversarial Examples Are Features.
  8. Apparent Bugs as Features: Some adversarial examples reflect predictive non-robust features in the data rather than an implementation bug in the classifier. — Adversarial Examples Are Features.
  9. Visual Cues: An image classifier can rely on small predictive cues that are not salient to human observers; its visual evidence need not match ours. — Adversarial Examples Are Features.
  10. A Benchmark Is a Proxy: ImageNet results measure performance on a constructed benchmark, whose labels and collection process do not fully define image classification. — ImageNet Benchmark Study.

Part 2: Adversarial Vulnerability

  1. Origins of Adversarial Examples: The coauthored study finds that some adversarial vulnerability can arise from predictive but non-robust data features, not merely defective optimization. — Adversarial Examples Are Features.
  2. Non-Robust Features: A feature may help predict the right label while being fragile to small adversarial changes and difficult for humans to perceive. — Adversarial Examples Are Features.
  3. The Model’s Incentive: Standard classifiers can use non-robust predictive features because those features improve the training objective. — Adversarial Examples Are Features.
  4. Beyond Obscuring Gradients: Adaptive-attack evaluations show why hiding or complicating gradients is insufficient evidence of a robust defense. — Adaptive Attacks Study.
  5. Robust Optimization: Adversarial training can be framed as minimizing loss against a bounded adversary that chooses difficult examples within a specified threat model. — Robust Optimization Paper.
  6. Human Perception Is Not the Whole Test: Imperceptibility to humans does not mean a perturbed image lacks predictive signals for a model; the paper explains some attacks through this feature mismatch. — Adversarial Examples Are Features.
  7. Adaptive Evaluation: A defense that resists a stock attack may still fail when the attack is adapted to its mechanism. — Adaptive Attacks Study.
  8. Training on Different Features: The authors construct datasets emphasizing robust or non-robust features to test how training data changes a model’s behavior. — Adversarial Examples Are Features.
  9. Use Strong Attacks: Robustness claims need evaluations matched to a clear threat model and strong attacks, including attacks adapted to the defense. — Adaptive Attacks Study.

Part 3: The Dependableness Trade-off

  1. Accuracy Has a Cost: In the settings studied, improving adversarial robustness can reduce accuracy on unperturbed test examples. — Robustness–Accuracy Study.
  2. Why the Trade-Off Appears: The authors relate the trade-off to predictive features that aid ordinary accuracy but are vulnerable to adversarial perturbation. — Robustness–Accuracy Study.
  3. Choose a Threat Model: Deployment decisions should specify which perturbations matter; the robust-optimization objective protects only against the adversary it models. — Robust Optimization Paper.
  4. Model Capacity: The robust-optimization study reports that capacity matters for learning models resistant to its specified adversarial attacks. — Robust Optimization Paper.
  5. A Conditional Trade-Off: The coauthored work proves an accuracy–robustness tension for particular data distributions; it does not claim every task must have the same trade-off. — Robustness–Accuracy Study.
  6. Look Beyond One Metric: Benchmark accuracy is an incomplete proxy for reliability on the broader image-classification task. — ImageNet Benchmark Study.
  7. Rework the Training Data: Experiments with robust and non-robust datasets show that changing available features can change the classifier’s adversarial behavior. — Adversarial Examples Are Features.
  8. Harder Training Objective: Robust training optimizes against an adversary, adding an inner maximization problem absent from ordinary empirical-risk minimization. — Robust Optimization Paper.

Part 4: Evaluating Frontier AI Risks

  1. Test Before Risk Escalates: Madry argues for measuring frontier-model risk proactively rather than discovering dangerous capabilities only after harm occurs. — Preparedness Talk Transcript.
  2. Specify Risk Categories: His Preparedness talk names cybersecurity, chemical and biological threats, persuasion, and model autonomy as categories for evaluation. — Preparedness Talk Transcript.
  3. Track Risk Continuously: Madry says risk assessments should be repeated as capabilities change, not treated as a single checklist. — Preparedness Talk Transcript.
  4. Forecast Emerging Risks: Preparedness, as Madry describes it, aims to forecast how misuse risk may evolve, alongside measuring present risk. — Preparedness Talk Transcript.
  5. Safety Baselines: Madry describes baselines that guide when deployment or development should pause and when security should increase as risk changes. — Preparedness Talk Transcript.
  6. Accountable Governance: The Preparedness talk discusses governance for enforcing baselines and considers external input and public accountability; it does not assign the original quoted burden of proof. — Preparedness Talk Transcript.
  7. Unknown Unknowns: Madry explicitly warns that careful checks can still miss important risks and urges a continuing search for unknown unknowns. — Preparedness Talk Transcript.
  8. Deployment Conditions: The proposed framework connects measured risk to protective actions and decisions about pausing deployment or development; Madry does not claim mathematical verification of all mitigations. — Preparedness Talk Transcript.
  9. Monitor Model Autonomy: Model autonomy was one of the categories Madry said the team would track through risk evaluations. — Preparedness Talk Transcript.
  10. Make Safety Operational: Madry calls for fact-based measurement, repeatable processes, concrete protective actions, and governance rather than safety aspirations alone. — Preparedness Talk Transcript.

Part 5: The Path Before AGI

  1. Before AGI Matters: Madry’s show is organized around what society should prepare for before AGI arrives, including the near-term effects of AI development. — Before AGI Podcast.
  2. Preparation Before AGI: The show explicitly explores practical preparatory steps and societal implications before AGI rather than treating AGI as the only relevant milestone. — Before AGI Podcast.

Part 6: AI Policy and Society

  1. Policy Must Evolve: In the MIT AI Policy Forum interview, Madry argues that AI rules need to evolve with understanding of the technology. — MIT AI Policy Forum Interview.
  2. Connect Research and Policy: Madry’s MIT interview calls for research and policy responses that take account of how AI systems are actually designed. — MIT AI Policy Forum Interview.
  3. Industry Standards: Madry specifically includes industry standards alongside law and regulation among the tools for shaping AI development. — MIT AI Policy Forum Interview.
  4. Design Choices Matter: Madry emphasizes that AI is built by people making design decisions, so those choices are a meaningful focus for governance. — MIT AI Policy Forum Interview.
  5. External Scrutiny: In his Preparedness Q&A, Madry says external input and third-party auditing matter, while describing that ecosystem as nascent. — Preparedness Talk Transcript.
  6. Assess Misuse Capabilities: Madry’s Preparedness work evaluates what harmful actors might do with AI, rather than assuming benign intended use removes risk. — Preparedness Talk Transcript.

Part 7: Data, Features, and Generalization

  1. Data Shapes Behavior: Non-robust predictive features in training data can shape what a classifier learns and how it responds to adversarial inputs. — Adversarial Examples Are Features.
  2. Robustness Needs Data: The coauthored robust-generalization study finds that learning adversarially robust classifiers can require more training examples than standard classification. — Robust Generalization Study.
  3. Hidden Dataset Signals: The benchmark study finds that dataset construction can shape measured performance; this replaces the unverified camera-signature anecdote. — ImageNet Benchmark Study.
  4. Feature-Focused Datasets: The authors’ robust and non-robust dataset constructions demonstrate a way to study which features a trained model uses. — Adversarial Examples Are Features.
  5. Shifted Data, Shifted Performance: A classifier’s benchmark success need not capture performance when examples, labels, or task conditions differ from the benchmark. — ImageNet Benchmark Study.
  6. Test the Intended Task: Evaluation should consider whether the benchmark reflects the broader image-classification task, not just held-out examples from the same benchmark. — ImageNet Benchmark Study.
  7. Non-Human Visual Signals: The feature paper shows classifiers can use predictive signals that people do not readily perceive, helping explain some adversarial examples. — Adversarial Examples Are Features.

Part 8: Building Deployable Machine Learning

  1. Real-World Reliability: The MIT center’s mission includes reliability under random and adversarial corruptions as a condition for deployable machine learning. — MIT Deployable ML Center.
  2. Beyond Leaderboards: The center frames deployability around reliability, understandable behavior, and safe use, not benchmark performance alone. — MIT Deployable ML Center.
  3. Understandable Systems: The center identifies human understandability as a goal of deployable machine learning; it does not promise an exact diagnosis for every failure. — MIT Deployable ML Center.
  4. Train for the Threat Model: Robust optimization incorporates a specified adversary in the training objective rather than adding only a post-hoc defense. — Robust Optimization Paper.
  5. Consequences of Deployment: Madry’s MIT policy discussion stresses that human design choices in AI systems have societal implications that call for technical and policy attention. — MIT AI Policy Forum Interview.
  6. Limits of Robustness Claims: The robust-optimization framework gives a precise bounded threat model, while adaptive-attack research shows empirical defense claims must face tailored testing; neither establishes universal safety guarantees. — Robust Optimization Paper.
  7. Cross-Functional Decisions: Madry describes a safety advisory group bringing together research, product, and safety perspectives to govern risk decisions. — Preparedness Talk Transcript.