Anton Troynikov is CEO of American Terawatt and a co-founder of Chroma, an open-source database for AI applications. His work includes industrial power infrastructure and software tools for AI memory and retrieval. These lessons cover embeddings, long-context reliability, machine-learning engineering, and his views on the social effects of AI. — Anton Troynikov — Career and Current Work. — The Cognitive Revolution — Embeddings and AI. — Chroma Research — Context Rot.

Visual summary of operating lessons from Anton Troynikov.

Embeddings, Retrieval, and AI Memory

  1. On Use Geometry to Operate on Meaning: Representing data as vectors makes geometric operations such as distance, density, clustering, and fitted surfaces available to applications. — The Cognitive Revolution — Embeddings and AI.
  2. On Let Embeddings Carry Semantic Structure: Modern embedding models place semantically related sentences near one another, creating an AI-native representation of data. — The Cognitive Revolution — Embeddings and AI.
  3. On Retrieve Knowledge the Model Never Saw: Private or current information can be embedded and retrieved for a general model even when it was absent from the model's training data. — The Cognitive Revolution — Embeddings and AI.
  4. On Treat Memory as a Platform Layer: A vector database is most useful when treated as a knowledge and storage layer for applications that keep language models in the loop. — The Cognitive Revolution — Embeddings and AI.
  5. On Study Behavior Through Representation: A model's behavior becomes easier to investigate when developers examine how it represents the data it receives. — The Cognitive Revolution — Embeddings and AI.
  6. On Combine Semantic and Exact Filters: Useful retrieval systems should support semantic similarity while also restricting results through exact metadata or phrase filters. — The Cognitive Revolution — Embeddings and AI.
  7. On Match Search Method to Dataset Size: Use exact nearest-neighbor search when a dataset is small enough for it to be practical; consider approximate methods as the collection grows. — The Cognitive Revolution — Embeddings and AI.
  8. On Tune the Precision-Recall Tradeoff: Tune approximate search around the workload: faster queries or inserts may come at the cost of missing some relevant neighbors. — The Cognitive Revolution — Embeddings and AI.
  9. On Cluster for Semantic Discovery: Geometric clusters in embedding space can expose groups of material that share human-readable themes. — The Cognitive Revolution — Embeddings and AI.
  10. On Learn Relevance Without Retraining: Human feedback can improve retrieval by learning a transformation over embedding space without retraining the underlying model. — The Cognitive Revolution — Embeddings and AI.

Building AI Infrastructure for Developers

  1. On Package Capability with the Right Affordances: A powerful model matters less if people cannot quickly understand what it enables or begin experimenting with it. — The Cognitive Revolution — Embeddings and AI.
  2. On Optimize for Rapid Experimentation: Early AI infrastructure should preserve development velocity instead of requiring a full engineer merely to operate one storage component. — The Cognitive Revolution — Embeddings and AI.
  3. On Build Around the Actual Workflow: Developer tooling should fit how applications are built rather than conforming to an abstract category definition. — The Cognitive Revolution — Embeddings and AI.
  4. On Make the Storage Layer AI-Native: An AI application store needs good inserts, updates, deletes, embedding integration, and output shaped for the model's context window. — The Cognitive Revolution — Embeddings and AI.
  5. On Turn Internal Pain into a Product: A strong infrastructure product can begin as an internal tool, then become broadly useful when many developers encounter the same obstacle. — The Cognitive Revolution — Embeddings and AI.
  6. On Prefer Empirical Machine-Learning Research: Machine-learning research should spend more time observing system behavior before forcing it into a theory. — The Cognitive Revolution — Embeddings and AI.
  7. On Separate Intuition from Engineering Truth: Useful intuition can produce a working system before the underlying engineering principles are fully understood. — The Cognitive Revolution — Embeddings and AI.
  8. On Use AI to Remove Peripheral Friction: Troynikov says coding assistants help with unfamiliar libraries and debugging, freeing him to focus on work he knows well; their suggestions can still be wrong. — The Cognitive Revolution — Embeddings and AI.

Long Context and Context Engineering

  1. On Reject Uniform Context Assumptions: Chroma’s experiments show that the tested models do not use long context uniformly, even on deliberately simple tasks. — Chroma Research — Context Rot.
  2. On Expect Reliability to Decline with Length: In Chroma’s tested tasks and model families, performance becomes less reliable as input length grows; the report does not cover every real-world use case. — Chroma Research — Context Rot.
  3. On Do Not Generalize from Lexical Retrieval: Needle-in-a-haystack benchmarks test narrow lexical matching and do not represent the ambiguity and reasoning of real applications. — Chroma Research — Context Rot.
  4. On Hold Difficulty Constant When Testing Length: A good long-context experiment changes input size without also making the underlying task harder. — Chroma Research — Context Rot.
  5. On Test Semantic Ambiguity: In Chroma’s retrieval experiments, lower similarity between the question and the relevant passage increases the rate of performance decline as context grows. — Chroma Research — Context Rot.
  6. On Treat Distractors as Unequal: In Chroma’s experiments, related distractors differ in how much they hurt retrieval, and their effects become more pronounced with longer inputs. — Chroma Research — Context Rot.
  7. On Expect Model-Family Differences: In the tested distractor conditions, some model families often abstain under uncertainty while others more often give confident but incorrect answers. — Chroma Research — Context Rot.
  8. On Recognize That One Distractor Can Matter: In Chroma’s experiments, one related distractor lowers performance relative to the baseline, and adding four compounds the decline. — Chroma Research — Context Rot.
  9. On Account for Semantic Blending: Chroma found that semantic overlap between the relevant passage and its surrounding content can affect retrieval, but its two tested topics do not establish a universal rule. — Chroma Research — Context Rot.
  10. On Account for Document Structure: In Chroma’s controlled retrieval experiments, models perform better on shuffled sentences than on logically coherent surrounding text; this is not a recommendation to shuffle every prompt. — Chroma Research — Context Rot.
  11. On Prefer Focused Context: In Chroma’s conversational-memory test, models answer more accurately when given only the relevant excerpts than when asked to find them in the full history. — Chroma Research — Context Rot.
  12. On Engineer Presentation, Not Only Presence: Chroma’s findings support careful context construction: where and how information appears can affect performance, though the report does not explain the underlying mechanisms. — Chroma Research — Context Rot.

Understanding Machine-Learning Systems

  1. On Do Not Project Human Meaning onto Models: Troynikov cautions that a model’s learned representation may differ from human perception because it is optimized for its training objective. — The Cognitive Revolution — Embeddings and AI.
  2. On Dissect the Machine: Anthropomorphic explanations conceal mechanisms; investigate where the model moves in latent space and why behavior changes. — The Cognitive Revolution — Embeddings and AI.
  3. On Train Models to Find and Compose Knowledge: Troynikov anticipates smaller models trained to retrieve and compose external knowledge rather than store all of it in their weights. — The Cognitive Revolution — Embeddings and AI.
  4. On Use Noise to Support Generalization: Troynikov points to dropout and training-data augmentation as ways to limit overfitting; his further explanation of noisy image-caption pairs was an intuition, not a demonstrated result. — The Cognitive Revolution — Embeddings and AI.
  5. On Extend Embeddings Across Modalities: Troynikov describes embeddings for text, images and audio, and sees potential for video, robotic actions, genomes and protein structures. — The Cognitive Revolution — Embeddings and AI.
  6. On Demand Higher Precision in Biology: Troynikov contrasts approximate media search with biological retrieval, where finding the specific protein can be essential; this is a precision requirement, not a claim of clinical efficacy. — The Cognitive Revolution — Embeddings and AI.
  7. On Frame Alignment as Control: Troynikov frames alignment as an engineering-control problem: test whether the machine actually performs the intended task rather than relying on anthropomorphic explanations. — The Cognitive Revolution — Embeddings and AI.

AI, Media, and Human Adaptability

  1. On Expect New Media to Destabilize Society: Troynikov worries that AI, like earlier communication technologies, could destabilize society before people develop ways to recognize manipulation. — The Cognitive Revolution — Embeddings and AI.
  2. On Build Mimetic Antibodies: Troynikov argues that societies need time to build shared literacy about manipulation in new media, which he calls mimetic antibodies. — The Cognitive Revolution — Embeddings and AI.
  3. On Remember the Capabilities Overhang: In the 2023 interview, Troynikov argues that people can discover useful capabilities in existing models without training new ones. — The Cognitive Revolution — Embeddings and AI.
  4. On Separate Information Risk from Physical Capability: Troynikov distinguishes generating dangerous information from carrying out a harmful plan in the physical world, where deployment remains difficult; he does not dismiss misuse risk. — The Cognitive Revolution — Embeddings and AI.
  5. On Use AI to Accelerate Adaptation: Troynikov hopes individualized reasoning tools could help more people approach the research frontier and respond to unfamiliar crises; this is an aspiration, not a demonstrated outcome. — The Cognitive Revolution — Embeddings and AI.