Ankur Goyal is an entrepreneur and engineering leader who has navigated the transition from high-performance infrastructure to the vanguard of AI development. After serving as VP of Engineering at MemSQL, now SingleStore, he founded Impira, a data and AI platform acquired by Figma. He subsequently led AI work at Figma before founding Braintrust, an AI observability and evaluation platform where he is currently CEO. His career reflects a deep commitment to software craftsmanship, rigorous evaluation, and solving the complex bottlenecks that prevent technical products from reaching production-grade quality.

Visual summary of operating lessons from Ankur Goyal.

Part 1: The Engineering Mindset & Software Quality

  1. On Extreme Paranoia: Quality is the hard part: "taking something from 95% to 99% or a hundred percent is very challenging and not something they teach you in school." — Reference: First Round Review, "What Braintrust got right about product-market fit" (Ankur Goyal podcast).
  2. On High-Bar Users: His years at MemSQL (now SingleStore) shaped his view of building for high-bar users, whose standards make quality non-negotiable. — Reference: First Round Review, "What Braintrust got right about product-market fit" (Ankur Goyal podcast).
  3. On Defining Specs vs. Prompts: When a formalized behavior specification and the runtime context disagree, the specification must take precedence to ensure the agent aligns with intended outcomes. — Braintrust Blog
  4. On Managing Trace Volume: "In an industry flooded with futuristic workflow promises, we've been fascinated by a more fundamental problem: building automation that can manage the volume of traces our customers generate, in a way that doesn't blow out token costs or break down at serious scale." — Braintrust Blog: AI observability is active observability
  5. On AI Observability Limitations: Traditional tools designed for microservices cannot handle the scale of AI engineering, requiring a fundamental shift in how databases are built to manage megabyte-sized traces. — Braintrust Blog: Brainstore

Part 2: The AI Era: Evals & Non-Determinism

  1. On The "Holy Shit" Moment: After systematic evals, a new model "wasn't perfect… but in aggregate, it was just way better. And that was a holy shit moment for me." — Reference: Latent Space, "Production AI Engineering starts with Evals — with Ankur Goyal of Braintrust" (2024).
  2. On UI and Design Convergence: "There's this very interesting convergence happening between UI engineering and design." — Reference: Latent Space, "Production AI Engineering starts with Evals — with Ankur Goyal of Braintrust" (2024).
  3. On Behavior Specifications: A behavior deserves to be formalized into a spec only when it governs a recurring choice, can be evaluated across an agent's trajectory, and is critical enough to test continuously. — Braintrust Blog
  4. On Retiring Specs: "Teams should retire behavior specs once the model reliably exhibits the behavior without extra guidance." — Braintrust Blog
  5. On The Purpose of Behavior Specs: "A behavior spec is an opinion, formed from experience, about how the agent should work." — Braintrust Blog
  6. On Distilling Production Data: It is cost-prohibitive to give coding agents access to all production data; instead, teams must build tools that automatically distill traffic down to the most critical insights. — Braintrust Blog: AI observability is active observability
  7. On AI in Production: "Since founding Braintrust, AI has rapidly moved from experimentation to production. Agents are no longer demos or side projects, but are embedded in nearly every engineering workflow." — Braintrust Blog: Announcing Series B

Part 3: Startup Strategy & Product-Market Fit

  1. On Testing the Thesis: Before building Braintrust, the team was "very skeptical" and wrote down "the conditions… that would have to be true" for it to be a good idea, then checked them with about 50 companies. — Reference: First Round Review, "What Braintrust got right about product-market fit" (Ankur Goyal podcast).
  2. On Dying in Obscurity: Of two failure modes, dying in obscurity or being overhyped and failing, "it's much better to be overhyped and then fail," within the bounds of integrity. — Reference: First Round Review, "What Braintrust got right about product-market fit" (Ankur Goyal podcast).
  3. On Funding vs. Fit: Early on he learned that "raising money was no indication of the quality of your idea or product market fit." — Reference: First Round Review, "What Braintrust got right about product-market fit" (Ankur Goyal podcast).
  4. On The Terrible Product Test: Zapier engineers used Braintrust's prototype five days after it was built, even though it was "terrible… so ugly, barely ran." That usage was an early sign. — Reference: First Round Review, "What Braintrust got right about product-market fit" (Ankur Goyal podcast).
  5. On Building Custom Infrastructure: Because traditional analytic databases are fundamentally unsuited for AI workloads, it is necessary to build custom database infrastructure rather than attempting to work around existing limitations. — Braintrust Blog: Brainstore

Part 4: Leadership, Culture & Recruiting

  1. On Internal Meetings: Braintrust holds one company-wide meeting per week. People who need frequent meetings to do their work "are not going to enjoy" it there. — Reference: First Round Review, "What Braintrust got right about product-market fit" (Ankur Goyal podcast).
  2. On Rigorous Interviewing: It's "important to have a rigorous and challenging interview, even for people that you know are good." His brother Manu, Braintrust's first hire, was interviewed heavily too. — Reference: First Round Review, "What Braintrust got right about product-market fit" (Ankur Goyal podcast).
  3. On Long-Term Recruiting: Recruiting great engineers "is not a transactional process." They would sometimes spend years recruiting someone, keeping a list and meeting quarterly. — Reference: First Round Review, "What Braintrust got right about product-market fit" (Ankur Goyal podcast).
  4. On Being Specific About Culture: At Impira the culture "regressed to the mean," so at Braintrust they have been "much more specific about what our culture is and how we operate." — Reference: First Round Review, "What Braintrust got right about product-market fit" (Ankur Goyal podcast).

Part 5: Customer Obsession & The Future of AI

  1. On Solving Your Own Frustration: Braintrust grew out of his frustration building evals while leading AI at Figma after it acquired Impira. — Reference: First Round Review, "What Braintrust got right about product-market fit" (Ankur Goyal podcast).
  2. On Process Supervision: "To build long-horizon agents you can trust, you have to supervise the process, not only the result." — Braintrust Blog
  3. On Active Observability: "We think of this next chapter as moving from 'AI observability' to 'active observability'. We work behind the scenes, constantly, to find answers to questions without you having to ask them ad-hoc." — Braintrust Blog: AI observability is active observability
  4. On AI as an Operating System: Because AI behaves like a constantly changing operating system that humans cannot fully inspect, teams must rely on scalable infrastructure to maintain clarity and accountability with every update. — Braintrust Blog: Announcing Series B
  5. On Customer Obsession: "Customers guide what we do at Braintrust. I make sure to talk to them every day, listening to their difficulties and celebrating their wins. If Braintrust makes their lives easier and their products better, I know we are doing our job." — Braintrust Blog: Announcing Series B