Visual summary of operating lessons from Alex Ratner.

Lessons from Alex Ratner

Alex Ratner is the co-founder and CEO of Snorkel AI and an affiliate assistant professor of computer science at the University of Washington. He helped develop Snorkel during his Stanford PhD to make training-data creation more programmatic. His work argues for putting data development and evaluation at the center of building reliable, specialized AI. — Chain of Thought: Every AI Agent Has an Evaluation Gap.

Part 1: The Bottleneck of Manual Labeling

  1. On making data a research problem: Ratner says training-data creation was long treated as upstream or janitorial work, even when it was the obstacle his collaborators needed help with. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  2. On the labeling bottleneck: In enterprise projects he has seen, getting a capable model can take about a day while labeling task data can take months; the old universal 90% statistic is removed. — Greylock: Jumpstarting Data-Centric AI.
  3. On changing labels: A hand-labeled dataset is costly to revise when categories or downstream requirements change, which motivates reusable programmatic supervision. — Stanford AI Lab: Weak Supervision.
  4. On scarce experts: Specialist annotations can be expensive; Ratner and coauthors use radiologists as an example of expertise that cannot be cheaply obtained at crowd-labeling scale. — Stanford AI Lab: Weak Supervision.
  5. On iteration speed: Ratner contrasts increasingly accessible models with the much longer process of preparing enterprise training data, making the latter a constraint on iteration. — Greylock: Jumpstarting Data-Centric AI.
  6. On Snorkel’s origin: Ratner says Stanford collaborators repeatedly asked for help labeling and curating data, prompting the team to turn that pain point into a research project. — Madrona: Founded & Funded with Alex Ratner.
  7. On private data: Ratner identifies private financial, insurance, medical and government data as settings where generic one-by-one labeling is especially hard to scale. — Madrona: Founded & Funded with Alex Ratner.
  8. On the model-centric era: He describes an earlier workflow in which researchers downloaded curated datasets and concentrated on model features and architecture, leaving data preparation outside the core loop. — Madrona: Founded & Funded with Alex Ratner.

Part 2: The Data-Centric AI Shift

  1. On data-centric development: For Ratner, data-centric AI makes data creation, labeling, curation and evaluation the primary levers for measuring and improving a model. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  2. On standardized models: As strong models have become easier to invoke, he argues that teams often gain more by improving task data than by inventing a new architecture; the model need not be literally fixed. — Greylock: Jumpstarting Data-Centric AI.
  3. On programming training data: Ratner and coauthors propose expressing labeling rules in code so one rule can generate supervision for many examples; the borrowed “Software 2.0” attribution is removed. — Stanford AI Lab: Weak Supervision.
  4. On data-development tools: His argument is that powerful model frameworks alone do not solve the work of building and maintaining the labeled data required for deployment. — Greylock: Jumpstarting Data-Centric AI.
  5. On failure slices: Ratner recommends identifying specific subtasks or error modes in an evaluation set so teams can see where a model needs improvement. — Building the Data Development Platform for Specialized AI.
  6. On the jagged frontier: He warns that AI can handle some difficult-looking tasks yet fail on ordinary-looking ones with messy inputs, long action sequences or complex outputs. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  7. On iterative data work: Ratner describes a loop of programming labels, training models, inspecting errors and revising the data rather than treating the dataset as a static download. — Madrona: Founded & Funded with Alex Ratner.
  8. On the foundation-model last mile: He treats general models as foundations that still need task-specific data and adaptation before they can reliably serve specialized enterprise workflows. — Greylock: Jumpstarting Data-Centric AI.
  9. On defining behavior: Ratner argues that evaluation data must define the tasks an AI system is meant to perform and the criteria by which its outputs will be judged. — Building the Data Development Platform for Specialized AI.

Part 3: Weak Supervision and Programmatic Labeling

  1. On weak supervision: Weak supervision uses higher-level but noisy inputs—such as patterns, knowledge bases or other classifiers—to generate labels programmatically. — Stanford AI Lab: Weak Supervision.
  2. On source accuracy without full labels: Snorkel’s original research models the agreements and conflicts among labeling functions to estimate source quality without requiring a large hand-labeled training set. — Snorkel: Rapid Training Data Creation with Weak Supervision.
  3. On labeling functions: A labeling function encodes a heuristic once and applies it across many examples, letting experts provide more leverage than one-label-at-a-time work. — Snorkel: Rapid Training Data Creation with Weak Supervision.
  4. On conflicting signals: The Snorkel framework combines overlapping and sometimes conflicting weak labels with a generative model rather than assuming each heuristic is correct. — Snorkel: Rapid Training Data Creation with Weak Supervision.
  5. On revisable rules: When a labeling definition changes, teams can modify the relevant rule and regenerate labels; the old million-document legal example is hypothetical, not a reported event. — Stanford AI Lab: Weak Supervision.
  6. On existing heuristics: The research treats existing rules, knowledge bases and classifiers as potential weak-supervision inputs rather than discarding them when building a trained model. — Stanford AI Lab: Weak Supervision.
  7. On experts as teachers: Ratner wants domain experts to contribute rules and feedback that teach a model across many examples, instead of spending all their time clicking individual labels. — Madrona: Founded & Funded with Alex Ratner.
  8. On many dynamic tasks: Ratner and coauthors envision systems supporting numerous related, changing tasks under weak supervision; the old amortization claim is not quantified. — Stanford AI Lab: Weak Supervision.
  9. On reusable labeling code: Representing supervision as code lets teams revise a rule and rerun it across data, making the source of labels more inspectable than an opaque hand-labeled batch. — Stanford AI Lab: Weak Supervision.

Part 4: The Role of Subject Matter Experts

  1. On domain judgment: Ratner’s enterprise examples show that the valuable bottleneck is often extracting subject-matter knowledge for the training or evaluation data, not acquiring another generic model. — Madrona: Founded & Funded with Alex Ratner.
  2. On expert heuristics: He describes gathering domain experts’ rules of thumb, then modeling their imperfections instead of assuming each rule is a definitive label. — Madrona: Founded & Funded with Alex Ratner.
  3. On experts in the loop: Ratner frames data development as a human-in-the-loop process in which experts help specify and validate supervision throughout model development. — Madrona: Founded & Funded with Alex Ratner.
  4. On translating domain knowledge: He says ML engineers and subject-matter experts need a way to turn business-specific judgment into labeled data and task definitions. — Greylock: Jumpstarting Data-Centric AI.
  5. On lower-friction expert input: A programmatic data-development interface can let experts express rules or judgments without making them design a new neural network architecture. — Madrona: Founded & Funded with Alex Ratner.
  6. On collaborative development: Ratner describes domain experts and ML practitioners working together to create and refine training data for specialized applications. — Greylock: Jumpstarting Data-Centric AI.
  7. On using expert time well: He favors capturing expert knowledge in higher-level rules and targeted reviews over relying entirely on repetitive one-example-at-a-time labeling. — Stanford AI Lab: Weak Supervision.
  8. On specifying success: In Ratner’s agent example, expert underwriters help define representative tasks and assess whether an answer meets the standards of their domain. — Data-Centric Development of an Enterprise AI Agent.

Part 5: Foundation Models in the Enterprise

  1. On foundation models as starting points: Ratner describes foundation models as strong starting points, not automatic production solutions for every specialized workflow. — Greylock: Jumpstarting Data-Centric AI.
  2. On specialized data: He argues that public-internet training alone cannot align an agent to all high-impact specialist tasks; expertise and task-specific data must be added. — Building the Data Development Platform for Specialized AI.
  3. On adaptation choices: Ratner names prompting and fine-tuning as possible ways to adapt a foundation model; the previous assertion that fine-tuning is almost always required is not supported. — Greylock: Jumpstarting Data-Centric AI.
  4. On distillation: Ratner distinguishes legitimate distillation, in which a larger model teaches a smaller one, from assuming a model can generate data that reveals all its own blind spots. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  5. On proprietary expertise: Ratner says enterprises can use their internal expertise and data to build differentiated datasets for specialized systems. — Building the Data Development Platform for Specialized AI.
  6. On avoiding architecture dependence: He argues that developing the task data gives teams a lever to improve or adapt models as model architectures change; the absolute “inoculates against lock-in” claim is removed. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  7. On evaluating generative systems: Ratner says specialist agent outputs need evaluators aligned with domain standards and expert judgment rather than an out-of-the-box vibe check. — Data-Centric Development of an Enterprise AI Agent.
  8. On stronger models and higher demands: More capable foundation models broaden what teams can try, while specialized deployment still demands careful data, evaluation and governance. — Greylock: Jumpstarting Data-Centric AI.

Part 6: The "Evaluation Gap" and Measuring AI

  1. On the evaluation gap: Ratner hypothesizes that the ability to build agents has advanced faster than the ability to define and measure their capabilities. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  2. On moving beyond demos: He warns that without precise evaluation, enterprises may be unable to deploy high-impact agents responsibly; “demo purgatory” was the host’s wording, not Ratner’s. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  3. On benchmark overfitting: Ratner acknowledges that teams can overfit to public benchmarks, but argues they remain useful alongside private, use-case-specific measurements. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  4. On long-horizon evaluation: He names autonomy horizon—the number and evolution of steps in a task—as a dimension missing from many simple agent benchmarks. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  5. On task-specific grading: For complex outputs, Ratner calls for rubrics and verifiers tailored to the task rather than relying on simple answer-key comparisons. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  6. On evaluation in the development loop: Ratner describes measurement, error inspection and data revision as continuing parts of improving an AI system, not merely an end-of-project pass/fail check. — Data-Centric Development of an Enterprise AI Agent.
  7. On high-stakes trust: He argues that inaccurate measurement is particularly consequential when enterprise agent errors carry real financial or operational costs. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  8. On useful failures: Ratner says data and environments are valuable when they expose where a model struggles, so teams can focus improvement on those gaps. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  9. On private and public benchmarks: He supports both open benchmarks that guide the field and private evaluations tailored to a deployment’s specific use case. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  10. On evaluation first: In his insurance-agent walkthrough, Ratner begins by curating representative evaluation tasks before selecting improvements to the model. — Data-Centric Development of an Enterprise AI Agent.

Part 7: From Academia to Startup

  1. On discovering messy data: Ratner says collaborators using medical, geological and enterprise data pushed the Stanford team away from model-centric research toward the harder work of labeling. — Madrona: Founded & Funded with Alex Ratner.
  2. On the Stanford-to-company path: Snorkel began as a Stanford lab project around expert labeling heuristics and later became a platform for enterprise data development. — Madrona: Founded & Funded with Alex Ratner.
  3. On research becoming a product: Ratner says years of office hours with varied users helped the team test its ideas and see what was needed beyond a research prototype. — Madrona: Founded & Funded with Alex Ratner.
  4. On customer-driven evolution: Ratner credits repeated customer needs—not merely new model architectures—with redirecting Snorkel’s work toward data-centric tools. — Greylock: Snorkel’s Obsession with Data.
  5. On open-source feedback: He says early code and Stanford office hours brought in users from consulting, biology and law, revealing how differently they tried to apply the methods. — Madrona: Founded & Funded with Alex Ratner.
  6. On customer observations: Ratner traces the original labeling project to repeated requests from lab collaborators who could not use another algorithm until they had suitable data. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  7. On remaining adaptable: As models and use cases change, he favors developing the data and evaluation layer so teams can measure new failure modes and adapt. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  8. On the Stanford research setting: Ratner says the lab’s weekly office hours exposed the team to varied real users while they developed weak-supervision theory and code. — Madrona: Founded & Funded with Alex Ratner.

Part 8: The Future of AI Agents

  1. On agentic tasks: Ratner characterizes newer agents by multi-step action in complex environments, with outputs that cannot always be checked by a simple answer key. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  2. On domain-specific agents: He argues that agents intended for specialized work need expert-derived tasks and standards reflecting the actual operating environment. — Building the Data Development Platform for Specialized AI.
  3. On supervising agent behavior: In Ratner’s evaluation framework, a task payload can include an environment, tools, reference answer, rubric and verifier—not just a final response. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  4. On the expert’s continuing role: Ratner argues that expert judgment remains necessary to generate, curate and review agent-training and evaluation data, even when automation accelerates the process. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  5. On long chains of action: Ratner says benchmarks should test extended tool use and evolving objectives because simple one-turn questions miss where agent behavior may fail. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  6. On expert-accelerated work: His preferred data-development model puts humans and AI together: automation accelerates generation and review, while experts supply knowledge machines cannot invent. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  7. On realistic context: Ratner says meaningful agent evaluation may require the actual mix of codebases, internal tools, tickets and instructions that forms a task’s environment. — Chain of Thought: Every AI Agent Has an Evaluation Gap.
  8. On the next frontier: Ratner expects agent progress to depend on expert-created data and environments that reveal blind spots and support more reliable evaluation and tuning. — Chain of Thought: Every AI Agent Has an Evaluation Gap.