Aman Khan is an AI product manager working on Google’s Agent Platform. He was previously Head of Product at Arize AI and has worked at Spotify, Cruise, Zipline, and Apple. His writing and teaching focus on evaluating AI products and learning through hands-on prototyping. — Aman Khan — Official Bio.

Part 1: Breaking into AI Product Management
- Define AI Product Quality: AI PMs need to understand model failure modes and define what good outputs look like, then turn that judgment into evaluations rather than relying only on a feature specification. — Khan — Five AI PM Skills.
- Build Technical Intuition: Khan recommends learning enough about prompts, retrieval, and fine-tuning to understand the cost and trade-offs of available solutions, while partnering closely with engineers. — Khan — Five AI PM Skills.
- Share Ownership of AI Outcomes: For AI products, Khan describes PMs and engineers jointly defining success metrics, labeling examples, debugging failures, and owning outcomes rather than passing a requirements document from one side to the other. — Khan — Five AI PM Skills.
- Prototype to Learn Model Behavior: Khan recommends building working AI prototypes to discover where an experience fails before committing to polished designs. — Khan — Five AI PM Skills.
- Demonstrate AI Product Skills: Khan's career-transition advice is to learn the tools and demonstrate AI PM knowledge through hands-on projects, rather than relying on a job-title change alone. — Khan — AI PM Roadmap.
- Learn Evaluation and Observability: Khan identifies evaluations and production observability as recurring skills for AI PMs; both help teams understand and improve unpredictable model behavior. — Khan — Five AI PM Skills.
- Design Around Observed Failures: Khan urges PMs to examine real user interactions and failure examples when defining AI quality, so product decisions reflect where the system actually disappoints users. — Khan — Beyond Vibe Checks.
- Work Closely with AI Engineers: Khan argues AI PMs should define success and label examples with engineers, rather than treating model behavior as an opaque engineering handoff. — Khan — AI PMs and Engineers on Evals.
- Measure More Than a Pass/Fail: Khan's evaluation approach combines specific output-quality checks with user feedback and business metrics instead of reducing AI quality to one binary result. — Khan — Beginner’s Guide to AI Evals.
Part 2: Moving Beyond "Vibe Checks"
- Replace Vibe Checks with Evals: Khan warns that inspecting a handful of plausible outputs is insufficient for judging AI-product quality; teams should use explicit criteria and representative examples. — Khan — Beyond Vibe Checks.
- Scale Output Reviews with Validated Judges: Human review supplies initial labels, while code checks and LLM judges can evaluate more outputs once checked against human judgments. — Khan — Beyond Vibe Checks.
- Test Changes Against a Dataset: Khan recommends rerunning evaluations over a representative dataset after prompt or model changes to detect regressions beyond the examples that inspired the change. — Khan — Beyond Vibe Checks.
- Specify What Good Means: Effective evals use explicit success criteria—such as factual correctness, tone, or adherence to instructions—rather than an undefined sense that output feels right. — Khan — Beyond Vibe Checks.
- Check Confident but Wrong Outputs: Khan highlights hallucination as an evaluation target: an answer may read plausibly while inventing facts or misusing supplied context. — Khan — Beyond Vibe Checks.
- Give Teams a Quality Baseline: Khan's eval workflow gives teams a common benchmark for comparing prompt and model changes, reducing guesswork about whether a change improved the product. — Khan — Beyond Vibe Checks.
- Connect Quality to Business Outcomes: Khan recommends relating evaluation scores to real user and business outcomes, rather than treating an isolated model score as sufficient evidence of success. — Khan — AI PMs and Engineers on Evals.
- Move from Exploration to Systematic Testing: A quick prototype can establish feasibility, but Khan distinguishes that stage from the evaluation and production monitoring needed for a reliable AI release. — Khan — AI PMs and Engineers on Evals.
- Monitor Reliability After Launch: Khan recommends continuous evaluation of live interactions and user feedback after launch because an offline test alone does not establish production quality. — Khan — Beyond Vibe Checks.
Part 3: Evaluations as the New PRD
- Use Evals as AI Product Requirements: Khan describes evaluations as a new definition of done for AI PMs: spell out desired behavior in examples and measurable criteria before declaring a feature ready. — Khan — Five AI PM Skills.
- Align PMs and Engineers on Examples: Jointly labeling outputs and refining criteria gives product and engineering a shared interpretation of what the AI should do. — Khan — AI PMs and Engineers on Evals.
- Build a Representative Labeled Dataset: Khan recommends collecting real interactions and human ground-truth labels so evaluations reflect user requests, including edge cases. — Khan — Beyond Vibe Checks.
- Validate LLM Judges Against Humans: LLM-as-judge evaluations can scale, but Khan advises comparing judge decisions with labeled human examples and refining the prompt when they disagree. — Khan — Beginner’s Guide to AI Evals.
- Expand Evals with New Failure Modes: Khan's iteration process adds newly discovered edge cases to the dataset and reruns the evaluation after product changes. — Khan — Beginner’s Guide to AI Evals.
- Evaluate Specific Behaviors: Khan names hallucination, tone, toxicity, and correctness as separate evaluation targets so a broad quality score does not hide the failure type. — Khan — Beyond Vibe Checks.
- Use Evals to Iterate: Khan frames evaluations as a feedback loop: measure an output, improve the prompt or model, and rerun the dataset to see whether the change helped. — Khan — Beyond Vibe Checks.
- Prioritize Observed Failures: Khan advises using production examples and failure patterns to decide which evaluation cases and product improvements deserve attention. — Khan — Beginner’s Guide to AI Evals.
- Translate User Needs into Criteria: The PM contributes user intent and examples while the engineer turns them into repeatable measurements; both sides refine the evaluator together. — Khan — AI PMs and Engineers on Evals.
- Use Disagreement to Refine Criteria: Joint labeling sessions surface differences between product and engineering judgments, making vague quality criteria concrete. — Khan — AI PMs and Engineers on Evals.
- Avoid AI-for-Everything Products: Khan favors narrowly scoped AI products with clear success metrics and straightforward evaluation criteria over products trying to do everything. — Khan — AI PMs and Engineers on Evals.
- Co-Write Evaluation Prompts: Khan describes evaluation-prompt writing as pair programming between PMs and engineers, not a one-way requirements handoff. — Khan — AI PMs and Engineers on Evals.
Part 4: Developing AI Product Sense
- Build AI Product Sense Through Use: Khan says AI product intuition develops through using the underlying primitives yourself, including context management, retrieval, tool calling, and reasoning. — Khan — Building AI Product Sense.
- Learn Model Capabilities by Experimenting: Khan recommends hands-on use of AI tools and prototypes to learn what they can and cannot reliably do, rather than relying on feature descriptions alone. — Raviv & Khan — AI Product Sense.
- Build a Personal AI Operating System: Khan uses a file-based personal OS as a way to experience an AI product's context, retrieval, and workflow limitations firsthand. — Khan — Building AI Product Sense.
- Match Performance to the User Experience: AI PMs should choose success criteria that reflect the user experience; Khan's evaluation examples include latency alongside response quality rather than treating either in isolation. — AI Engineer — Evaluation Framework for PMs.
- Design for Long-Tail Failures: Khan stresses collecting edge cases and real customer failures because AI systems exhibit many more failure modes than deterministic software. — Learning from Machine Learning — Aman Khan.
- Choose Problems with Testable Outcomes: Khan favors well-scoped AI products with clear success metrics and straightforward evaluation criteria. — Khan — AI PMs and Engineers on Evals.
- Manage Context, Not Just Tasks: Khan's personal OS treats the files, goals, and work context an agent reads as the material to manage; he retains the final decision. — Khan — Building AI Product Sense.
- Give AI Context Without Over-Structuring: Khan recommends giving a personal agent raw context with minimal structure, allowing it to synthesize patterns while a person makes final calls. — Khan — Building AI Product Sense.
- Watch for False Positives in Evals: The original episode flags outputs that appear correct but rest on flawed reasoning as a dangerous evaluation false positive. — Supra Insider — Predictable AI Products.
Part 5: Non-Deterministic Design Principles
- Design for Non-Determinism: Khan contrasts deterministic unit tests with AI systems' broader failure space, arguing for statistical confidence and monitoring of acceptable overall behavior. — Learning from Machine Learning — Aman Khan.
- Capture Corrections as Evaluation Data: Khan recommends gathering real user feedback and examples, then turning newly discovered errors into cases for evaluation and iteration. — Khan — Beyond Vibe Checks.
- Use Context to Improve Agent Decisions: Khan's personal OS keeps goals, meeting transcripts, and task context available to the agent so its suggestions can adapt as those inputs change. — Khan — Building AI Product Sense.
Part 6: Observability and the ML Lifecycle
- Observe AI Behavior in Production: Khan says tracing lets teams inspect an agent's inputs, outputs, and intermediate actions instead of treating a bad answer as an opaque failure. — Khan — Five AI PM Skills.
- Watch for Data Drift: In his feature-store observability deck, Khan defines drift as distribution change over time and recommends monitoring data quality, feature drift, and model performance. — Khan — Feature Store Observability.
- Trace the Failing Step: Khan gives retrieval of the wrong document as a concrete example of why tracing is more actionable than simply reporting that an AI system is broken. — Khan — Five AI PM Skills.
- Test Prompt Changes: Khan demonstrates tracing an application prompt, editing it in a playground, and comparing outputs through evaluations before adopting a change. — AI Engineer — Evaluation Framework for PMs.
- Feed Production Cases Back into Evals: Khan recommends capturing real-world interactions, labeling them, and adding new failure cases to evaluation datasets for the next iteration. — Khan — Beginner’s Guide to AI Evals.
- Use Traces to Debug a Response: Tracing records inputs, outputs, and metadata around an agent request, allowing the team to inspect the actions behind a reported bad result. — AI Engineer — Evaluation Framework for PMs.
- Monitor Product-Specific Quality: Khan's eval framework looks beyond latency or error rates to application-specific response qualities such as friendliness, factuality, and business outcome. — AI Engineer — Evaluation Framework for PMs.
- Detect Production Quality Changes: Khan recommends continuous evaluation on live requests so teams can see whether changes to the system are affecting user experience over time. — Khan — Beyond Vibe Checks.
Part 7: Prototyping with AI Agents
- Use AI Tools to Build Working PM Prototypes: Khan describes using Cursor and other AI coding tools to create functional prototypes before a full engineering build. — Khan — AI Prototyping for PMs.
- Prototype in Hours, Not Weeks: Khan reports that AI coding tools let him prototype product ideas in hours or minutes, sometimes during a meeting, before committing engineering time. — Khan — AI Prototyping for PMs.
- Learn Constraints by Building: Khan recommends building a working AI prototype to expose how the model and tools actually behave, then using that knowledge to discuss trade-offs with engineers. — Khan — Five AI PM Skills.
- Bring a Working Prototype to Engineering: Khan says a functional prototype can give his engineering team a higher-resolution starting point than a document of requirements. — AI Engineer — Evaluation Framework for PMs.
- Test Feasibility Before Polishing UI: Khan says functional prototypes can reveal flaws in an AI experience that polished mocks missed, allowing teams to revise the idea earlier. — Khan — Five AI PM Skills.
- Explore More Ideas with Lower Build Cost: Khan argues that as the cost of creating software falls, PMs can test ideas directly, putting more emphasis on choosing the right problem. — AI Engineer — Evaluation Framework for PMs.
- Learn by Using AI Tools: Khan's career advice is to get familiar with AI tools and demonstrate capability through hands-on projects. — Khan — AI PM Roadmap.
- Shift from Handoff to First Draft: Khan describes a PM bringing prototypes or mocks into a design and engineering discussion, making the first draft a concrete artifact rather than only a PRD. — Becoming an AI PM — Video Interview.
- Write Tests as Acceptance Criteria: The coauthored Cursor guide uses coding-agent tests to encode constraints, including character limits, null inputs, and typos; attribute this to Eric Xiao and Aman Khan jointly. — Xiao & Khan — Cursor for PMs.
Part 8: Lessons from Spotify, Apple, and Cruise
- Apply Evaluation Lessons from Self-Driving: Khan says his work on self-driving evaluation systems at Cruise informed his later work evaluating AI agents; both require systematic testing of non-deterministic behavior. — AI Engineer — Evaluation Framework for PMs.
- Monitor Data Quality, Drift, and Performance: Khan's feature-store presentation emphasizes data-quality checks, drift analysis, and performance measures as connected parts of ML observability. — Khan — Feature Store Observability.
- Carry Evaluation Skills Across Domains: Khan traces work on self-driving evaluations at Cruise through recommendation systems at Spotify to AI-agent evaluations at Arize, illustrating evaluation as a transferable discipline. — AI Engineer — Evaluation Framework for PMs.