Chip Huyen is an engineer, author, and educator known for mapping out the practical realities of deploying machine learning systems into production. She wrote the widely read book Designing Machine Learning Systems and has become a leading voice in treating AI not as magic, but as a rigorous software engineering discipline. This collection organizes her insights across MLOps, system complexity, large language models, and technical hiring. — Designing Machine Learning Systems.

Visual summary of operating lessons from Chip Huyen.

Part 1: The Engineering First Mindset

  1. Engineering before specialization: Huyen advises engineers interested in ML to strengthen their engineering skills, which she sees as a durable foundation for building ML products. — What I learned from looking at 200 machine learning tools.
  2. Build useful ML tools: Huyen explicitly invites engineers to build better tools for ML practitioners, identifying production tooling as an unmet need. — What I learned from looking at 200 machine learning tools.
  3. Data science and ML engineering: Huyen distinguishes data science’s traditional focus on business insight from ML engineering’s goal of turning data into products, while noting that the roles can overlap. — ML Engineer vs. Data Scientist.
  4. The combined AI stack: Huyen explains that real-world AI systems may combine conventional ML models and foundation models, so both ML engineering and AI engineering can matter. — AI Engineering Book.
  5. Own the workflow without owning every tool: Huyen argues that good infrastructure abstractions can help data scientists own projects end to end without requiring them to manage every low-level infrastructure detail. — Why Data Scientists Shouldn’t Need to Know Kubernetes.
  6. Coding is only part of ML: Huyen contrasts conventional software work with ML projects, where data, model development, deployment and tooling extend far beyond writing code. — What I learned from looking at 200 machine learning tools.
  7. Start simple to debug: Huyen recommends adding complexity gradually so teams can understand and debug a system before investing in more elaborate models. — Design a Machine Learning System.
  8. Establish a baseline: Huyen recommends comparing complex models with random, human and simple-heuristic baselines before judging whether the added complexity helps. — Design a Machine Learning System.

Part 2: The Reality of Production and Deployment

  1. Deployment begins ongoing maintenance: Huyen warns that deploying a model is not the end: performance can degrade as real-world data changes, requiring monitoring and updates. — Data Distribution Shifts and Monitoring.
  2. Bridge development and production: Huyen identifies scale and state as two differences between development and production environments that make moving an ML project into production difficult. — Why Data Scientists Shouldn’t Need to Know Kubernetes.
  3. Optimize for the actual serving requirement: Huyen distinguishes research-oriented model development from production requirements such as inference latency and operational constraints. — Research vs. Production.
  4. Operationalize beyond deployment: Huyen treats production ML as an ongoing workflow that includes serving, monitoring and maintaining models, not just training them. — Data Distribution Shifts and Monitoring.
  5. Monitor before retraining: Huyen recommends monitoring production models for performance changes and deploying updates when issues are detected; retraining is a response to evidence, not a universal automatic schedule. — Data Distribution Shifts and Monitoring.
  6. Some models need frequent deployment: Huyen notes that fast-changing data can require short development and deployment cycles, in some cases including nightly model updates. — What I learned from looking at 200 machine learning tools.
  7. Design components to work together: The CS329S course notes warn that an ML system without intentional design tying its components together can become fragile and costly to maintain. — Stanford CS329S: Machine Learning Systems Design.
  8. Plan for changing systems: Huyen notes that tutorial-built systems can age quickly as tools, business requirements and data distributions change. — Stanford CS329S: Machine Learning Systems Design.
  9. Ask whether ML is needed: Huyen cautions against defaulting to ML or generative AI where a simpler solution would meet the product goal. — Common Pitfalls in Generative AI Applications.

Part 3: Data Quality Over Algorithms

  1. Data can be the competitive advantage: In her 2020 analysis of ML tooling, Huyen argued that many applications would compete on the quality and quantity of their data rather than novel algorithms alone. — What I learned from looking at 200 machine learning tools.
  2. Reuse models where appropriate: Huyen anticipated that many companies would adopt available models and invest in the data and tooling needed to turn them into useful applications. — What I learned from looking at 200 machine learning tools.
  3. Test model assumptions: Huyen advises checking whether a proposed model’s assumptions match the available data and the problem being solved. — Design a Machine Learning System.
  4. Inspect the data early: Huyen says that even a short period spent looking directly at project data has often revealed problems that saved her hours later. — Common Pitfalls in Generative AI Applications.
  5. Escape the data catch-22: Huyen suggests launching without deep learning when a product first needs users to generate the data a deep-learning system would require. — Design a Machine Learning System.
  6. Design for data dependence: Huyen describes ML systems as unusually data-dependent: data varies from one use case to another, so a design cannot simply be copied wholesale. — Designing Machine Learning Systems.
  7. Formulate the problem with data in mind: For ML system design, Huyen asks candidates to specify what data and how much of it each proposed formulation would need for both training and evaluation. — Design a Machine Learning System.
  8. Understand data parallelism: Huyen describes data parallelism as splitting data across workers, training in parallel and aggregating gradients, with synchronization challenges to manage. — Research vs. Production.

Part 4: System Complexity and Design

  1. ML systems have many moving parts: Huyen calls ML systems complex because they involve multiple components and stakeholders, and unique because their data differs by application. — Designing Machine Learning Systems.
  2. Think beyond the model: Huyen’s ML-systems-design framework includes project setup, data pipeline, modeling and serving, rather than treating model training as the whole system. — Design a Machine Learning System.
  3. Context determines design: Huyen’s system-design guide begins with goals, users, performance constraints, available data and project constraints; design choices depend on those specifics. — Design a Machine Learning System.
  4. Define the system requirements: CS329S defines ML systems design as choosing architecture, infrastructure, algorithms and data to satisfy specified requirements. — Stanford CS329S: Machine Learning Systems Design.
  5. Beat the simple heuristic: Huyen says a more complex model should outperform a useful heuristic by enough to justify its extra complexity; her app-recommendation example uses a 70% heuristic baseline. — Design a Machine Learning System.
  6. Try rules before ML: Huyen recommends looking for effective heuristics before reaching for ML, while warning that excessively complex rules may themselves justify an ML approach. — Design a Machine Learning System.
  7. Deep learning must earn its cost: Huyen notes that deep-learning models can be expensive to train and difficult to explain, so simpler alternatives may be preferable unless performance gains justify them. — Design a Machine Learning System.
  8. Fit the solution to requirements: Huyen’s design guidance evaluates candidate approaches against goals, user experience, latency, accuracy and project constraints. — Design a Machine Learning System.

Part 5: Navigating LLMs and AI Agents

  1. The demo-to-production gap: Huyen observes that an impressive LLM demo can be easy to make while reliability, evaluation and operations make production use harder. — Building LLM Applications for Production.
  2. Treat prompts as an engineering input: Huyen recommends testing and managing prompts deliberately because small changes can alter an LLM application’s behavior. — Building LLM Applications for Production.
  3. Decide what is worth building: In a direct talk, Huyen argues that as AI makes software easier to replicate, choosing and imagining worthwhile products becomes increasingly important. — Pragmatic Engineer Talk: Building When AI Can Build.
  4. Test prompts beyond their examples: Huyen warns that an LLM can overfit few-shot examples in a prompt and recommends evaluating on separate examples. — Building LLM Applications for Production.
  5. Keep room for building for joy: Huyen says in her talk that she continues to build because she enjoys it, and hopes building for fun can remain worthwhile even as replication gets easier. — Pragmatic Engineer Talk: Building When AI Can Build.
  6. Version prompts: Huyen recommends tracking prompt changes as part of an LLM application’s development and evaluation process. — Building LLM Applications for Production.
  7. Manage the application around the model: Huyen describes LLM applications as systems that combine model behavior with programmatic components, prompts, data and evaluation. — Building LLM Applications for Production.
  8. Separate hype from production demand: In 2020 Huyen expected ML hype to cool while demand for production tooling would continue; this is her dated forecast, not a current market measurement. — What I learned from looking at 200 machine learning tools.

Part 6: Evaluation and Metrics

  1. LLM evaluation is hard: Huyen explains why open-ended language outputs and changing use cases make LLM application evaluation difficult. — Building LLM Applications for Production.
  2. Do more than a vibe check: Huyen warns against treating an impressive demo or a subjective first impression as sufficient evaluation for a production AI application. — Common Pitfalls in Generative AI Applications.
  3. Define what success means: Huyen recommends tying evaluation to the actual application goal and developing use-case-specific measurements rather than relying on a single generic benchmark. — Common Pitfalls in Generative AI Applications.
  4. Connect ML and business objectives: Huyen’s system-design guidance starts with business and product goals, then chooses ML metrics that help evaluate progress toward them. — Design a Machine Learning System.
  5. Do not rely only on AI judges: Huyen identifies relying entirely on AI judges while skipping human evaluation as a common evaluation pitfall. — Common Pitfalls in Generative AI Applications.
  6. Evaluate the evaluators: Huyen says AI judges are non-deterministic and should be tested and improved over time for the use case at hand. — Common Pitfalls in Generative AI Applications.
  7. Keep human review in the loop: Huyen reports that strong product teams she has observed supplement automated evaluation with recurring human assessment of application outputs. — Common Pitfalls in Generative AI Applications.

Part 7: Real-Time Machine Learning

  1. Distinguish online inference and continual learning: Huyen distinguishes real-time prediction from the more demanding ability to incorporate new data and update a model continually. — Machine Learning Is Going Real-Time.
  2. Streaming work does not simply finish: Huyen contrasts finite batch jobs with streams that continuously process arriving events. — Introduction to Streaming for Data Scientists.
  3. Evaluate online updates cautiously: Huyen notes that online training introduces the challenge of evaluating a newly updated model before exposing users to its predictions. — Real-Time Machine Learning: Challenges and Solutions.
  4. Use online evidence alongside offline tests: Huyen recommends using post-deployment feedback and monitoring to understand how a model actually performs for users, not relying only on offline test results. — Data Distribution Shifts and Monitoring.
  5. Design serving for latency: Huyen discusses real-time inference as an architectural choice with latency requirements distinct from training throughput. — Real-Time Machine Learning: Challenges and Solutions.
  6. Streaming features need state: Huyen explains that stream processing may need to retain state over time windows and handle events that arrive late or out of order. — Introduction to Streaming for Data Scientists.
  7. Freshness can change prediction value: Huyen explains that in use cases whose signals change quickly, fresh data can improve predictions; the required freshness depends on the application. — Real-Time Machine Learning: Challenges and Solutions.
  8. Watch for model feedback loops: Huyen describes how predictions can shape the data later used to evaluate or train models, creating self-reinforcing feedback effects. — Data Distribution Shifts and Monitoring.
  9. Real-time systems have infrastructure costs: Huyen cautions that moving from batch to real-time processing can require additional streaming infrastructure and operational complexity. — Introduction to Streaming for Data Scientists.

Part 8: Careers, Interviews, and Open Source

  1. Learn by teaching: Huyen describes teaching TensorFlow while still learning it herself, using students’ questions to deepen her understanding. — Confession of a So-Called AI Expert.
  2. Cultivate interests through practice: In career advice, Huyen describes making room for both engineering and writing rather than treating passion as a single preordained career choice. — Career Advice for CS Graduates.
  3. Protect time for your own work: Huyen argues that a career can extend beyond a paid job and advises reserving time for side projects and interests of one’s own. — Lessons from My First Full-Time Job.
  4. Make friends, not transactions: Huyen advises getting to know people for who they are rather than treating relationships solely as professional connections. — Lessons from My First Full-Time Job.
  5. Bad questions can damage hiring: Huyen argues that poorly designed interview questions reveal weak hiring processes and can lead to poor decisions. — Bad Interview Questions.
  6. Coordinate interview coverage: Huyen advises hiring managers to assign interviewers different skills to assess so the panel can build a more complete picture of a candidate. — The Interview Pipeline.
  7. Clarify ambiguous interviews: Huyen advises candidates to ask what an unclear question is intended to evaluate so they can give a more relevant answer. — Bad Interview Questions.
  8. Open source is not costless: Huyen notes that open-source tools can still be commercial and that maintaining them takes substantial time and resources. — What I learned from looking at 200 machine learning tools.
  9. Why teams open source tools: Huyen names transparency, collaboration, flexibility and customer confidence among reasons to release tools with visible source code. — What I learned from looking at 200 machine learning tools.