Zico Kolter is a professor and the head of the Machine Learning Department at Carnegie Mellon University, a board member at OpenAI, and co-founder of the AI security startup Gray Swan. He is known for his research on the robustness of deep learning systems and for establishing frameworks to evaluate and secure large language models against novel vulnerabilities. This profile outlines his perspectives on the mechanisms driving modern artificial intelligence, the unique security threats it poses, and the governance structures required to responsibly manage its development.

Part 1: The Nature and Mechanics of Modern AI
- On fundamental mechanisms: Modern large language models operate on the basic premise of using mathematical equations to predict the next word in a sequence based on vast amounts of internet training data. — Reference: 20VC with Harry Stebbings
- On the emergence of intelligence: The fact that chaining simple word predictors together produces coherent and sophisticated responses is a monumental scientific discovery. — Reference: 20VC with Harry Stebbings
- On model intelligence: "I really believe these systems are intelligent" — Source: 20VC with Harry Stebbings
- On structural simplicity: The core architecture of these advanced systems relies on remarkably little code, which stands in stark contrast to their behavioral complexity. — Reference: The MAD Podcast with Matt Turck
- On software scale: "That entire set of code, probably two to 300 lines of Python code." — Source: The MAD Podcast with Matt Turck
- On the source of complexity: The true complexity and capability of modern artificial intelligence systems are derived entirely from the massive datasets on which they are trained. — Reference: The MAD Podcast with Matt Turck
- On data availability: While models have already consumed the most easily available, high-quality data on the internet, this does not necessarily mean their performance gains will soon plateau. — Reference: 20VC with Harry Stebbings
- On historical precedent: While past technological shifts automated physical labor or numerical calculations, the current wave of artificial intelligence is fundamentally different because it automates intelligence itself. — Reference: Where What If Becomes What's Next Podcast
- On simple mechanisms and real intelligence: Describing a model as a next-word predictor does not prove that it lacks intelligence; the sophisticated behavior emerging from that mechanism must be judged on what the system actually does. — Reference: 20VC with Harry Stebbings
Part 2: The Scope of AI Safety
- On the illusion of scale: "You can't just sort of trust models to get safer by getting bigger." — Source: The MAD Podcast with Matt Turck
- On immediate security threats: The first major category of AI risk involves short-term vulnerabilities like prompt injection and the unauthorized exfiltration of private data. — Reference: Where What If Becomes What's Next Podcast
- On societal disruption: The second tier of AI safety concerns revolves around how these systems will impact the economy, the job market, and human mental health. — Reference: Where What If Becomes What's Next Podcast
- On catastrophic risks: The third category addresses the severe dangers that could arise if malicious actors successfully harness AI for destructive physical or digital capabilities. — Reference: Where What If Becomes What's Next Podcast
- On long-term existential risk: The fourth area of safety research focuses on the theoretical future challenges of managing uncontrollable artificial superintelligence. — Reference: Where What If Becomes What's Next Podcast
- On guaranteed robustness: Researchers have pioneered techniques that use classical optimization within neural network layers to embed hard constraints directly into deep learning models. — Reference: OpenAI
- On automated safety assessment: By developing automated methods for evaluating the safety of large language models, teams have proven that existing model safeguards can be reliably bypassed through optimization techniques. — Reference: OpenAI
Part 3: Model Governance and Oversight
- On the Safety and Security Committee: The role of OpenAI's safety committee is to oversee the governance of model development, interacting with distinct internal teams dedicated to preparedness, alignment, and model policy. — Reference: The MAD Podcast with Matt Turck
- On corporate governance parallels: The oversight function of an AI safety committee is highly analogous to the role an audit committee plays in independently reviewing a company's finances. — Reference: The MAD Podcast with Matt Turck
- On reviewing releases: Before a major frontier model is released, the safety committee convenes to review deployment safeguards and evaluate external third-party reports. — Reference: The MAD Podcast with Matt Turck
- On halting deployments: If a model fails to meet safety and readiness policies, the oversight committee has the authority to delay its release until the issues are addressed. — Reference: The MAD Podcast with Matt Turck
- On accepting current limitations: "If a model is not good enough at something, what do you do? You wait, right? Because the next model will be better at it." — Source: The MAD Podcast with Matt Turck
- On collective responsibility: Given that artificial intelligence integrates into critical infrastructure, proper oversight requires active collaboration between industry developers, academia, and government entities. — Reference: Where What If Becomes What's Next Podcast
- On establishing standards: It is critical for leading AI developers to formalize safety and governance committees early to build public trust and establish baseline standards for responsible deployment. — Reference: The MAD Podcast with Matt Turck
- On continuous oversight: Safety reviewers should remain close enough to researchers throughout development that final release meetings contain no surprises or last-minute discoveries. — Reference: The MAD Podcast with Matt Turck
Part 4: AI Security, Red Teaming, and Agents
- On the limits of traditional cybersecurity: Because AI models behave in fundamentally different ways than traditional software, the methods used to secure them require a completely new mindset. — Reference: Latent Space
- On human-like vulnerabilities: Unlike rigid software programs, AI systems can be deceived or tricked in ways that closely mirror the methods used to trick human beings. — Reference: Latent Space
- On automated red teaming: Specialized automated models designed to probe for weaknesses are now proving more capable of breaking into AI systems than human red teamers. — Reference: Latent Space
- On correlated failures: A major security risk lies in the widespread adoption of a few foundational models; an exploit discovered in one core system quickly becomes a universal vulnerability across many downstream applications. — Reference: Latent Space
- On the risk of AI agents: The shift toward giving AI agents the ability to browse the web, execute code, and take independent actions exponentially expands the attack surface for bad actors. — Reference: Scripod
- On independent security ecosystems: Just as traditional computing platforms spawned their own cybersecurity sectors, the AI industry is actively generating third-party security entities focused strictly on defending and auditing models. — Reference: Latent Space
- On treating models as untrusted: Adopting AI introduces model-specific vulnerabilities, so deployed models should be treated as untrusted systems whose risks must be understood and mitigated directly. — Reference: Latent Space — Red-Teaming after Mythos
- On specialized red-teamers: Frontier models often refuse adversarial tasks because of their safeguards, so effective automated red teaming requires models explicitly trained for that purpose. — Reference: Latent Space — Red-Teaming after Mythos
- On automating interpretability: Before automating every other science, AI should help turn the still-ad-hoc study of model behavior and interpretability into a systematic science. — Reference: Latent Space — Red-Teaming after Mythos
- On defense in depth: Agent guardrails must be combined with isolation, authentication, and access controls because no AI-layer control can secure every possible tool use by itself. — Reference: Latent Space — Red-Teaming after Mythos
- On least-privilege agents: Letting an agent inherit every permission of the human operating it is an unsafe default that must give way to narrower agent-specific access. — Reference: Latent Space — Red-Teaming after Mythos
Part 5: Societal Trust and the Path Forward
- On the erosion of objective reality: The proliferation of AI accelerates a societal crisis where people may no longer believe anything they see, threatening the concept of an objective record of facts. — Reference: 20VC with Harry Stebbings
- On deepfakes and media: While AI technologies exacerbate existing issues with digital misinformation, they might also eventually provide the necessary tools to detect manipulated media. — Reference: Where What If Becomes What's Next Podcast
- On human-AI relationships: As machines develop the ability to simulate reasoning, developers must carefully navigate the psychological impacts these interactions will have, particularly on vulnerable populations. — Reference: Where What If Becomes What's Next Podcast
- On privacy fears: Public discourse is frequently skewed by misconceptions regarding how AI chatbots actually process and store personal user data. — Reference: Where What If Becomes What's Next Podcast
- On physical limitations: The future scaling of AI is heavily constrained by real-world physical challenges, including infrastructure construction and enormous energy requirements. — Reference: Where What If Becomes What's Next Podcast
- On the timeline to AGI: When considering the near-term development of artificial general intelligence, Kolter maintains a vision that is highly optimistic about technical capabilities yet extremely cautious regarding safety. — Reference: Where What If Becomes What's Next Podcast
- On separating agent personas: Work, home, and distinct professional contexts should not collapse into one agent identity and permission set; agents need the same separation of roles that humans maintain. — Reference: Latent Space — Red-Teaming after Mythos
- On insurable AI: AI insurance should pair rigorous deployment-risk measurement with concrete safeguards that reduce the risk of systems otherwise considered too dangerous to insure. — Reference: Latent Space — Red-Teaming after Mythos
- On standards before breaches: The industry should establish credible AI compliance frameworks before an inevitable public prompt-injection breach forces standards to emerge reactively. — Reference: Latent Space — Red-Teaming after Mythos