Ron Alfa, a physician-scientist and the co-founder of Noetik, wants to fix the high failure rate of cancer clinical trials. He builds foundation models of human biology from massive, purpose-built datasets to predict how individual patients will respond to treatment. The lessons below detail his strategy for moving new therapies out of the lab and into successful clinical use.

Part 1: The Translation Bottleneck and Clinical Failure
- On the true cause of clinical trial failure: Most drug failures in clinical trials happen not because the molecules are poorly designed, but because the trials enroll broad populations without identifying the specific patients who would actually benefit. — Reference: Biotech companies are being reshaped by AI - SiliconANGLE
- On predicting clinical success: The most impactful problem to solve in drug discovery today is translation—predicting whether a therapeutic will succeed when moving from early discovery and preclinical models into actual human patients. — Reference: Foundation Model Series: Harnessing Multimodal Data to Advance Immunotherapies with Ron Alfa from Noetik
- On redefining patient populations: Instead of relying on traditional trial-and-error, large foundation models can mathematically define the exact patient population where a drug will be effective before the trial even begins. — Reference: Interview Questions for Ron Alfa, NOETIK - PMWC Precision Medicine World Conference
- On unlocking the value of existing drugs: By improving our understanding of unique cancer biologies and accurately predicting response, we can significantly increase the success rates of treatments that have already been developed but failed in broader trials. — Reference: 🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
- On evaluating historical trial cohorts: Because inference can be performed on standard H&E pathology slides, AI models can retrospectively analyze past clinical trials to identify which specific biological profiles distinguished responders from non-responders. — Reference: Ron Alfa — Latent Space Summary
- On actionable output for pharma: Instead of just recommending a drug, AI can identify novel patient subtypes based on molecular patterns and provide pharmaceutical partners with a specific biomarker signature and a predicted probability of response. — Reference: 🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik - Deep Analysis & Key Insights | Metapodcast
Part 2: Moving Beyond Reductionist Biology
- On the limits of traditional models: The biotech industry has spent decades relying on reductionist cell lines and proxy animal models that lack context and do not accurately reflect the complexity of human disease. — Reference: Interview Questions for Ron Alfa, NOETIK - PMWC Precision Medicine World Conference
- On prioritizing human biology: Rather than trying to refine animal models or organoids to improve translation, the most direct and effective approach is to build models based directly on real human tissue data. — Reference: Lessons from Ron Alfa, Co-Founder and CEO of Noetik, on Bringing AI-Native Precision Oncology to Every Cancer Patient
- On simplistic biomarkers: Traditional biomarkers are often too simplistic, relying on a single mutation or protein stain, which fails to capture the intricate spatial and immune patterns driving cancer biology. — Reference: 🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik - Deep Analysis & Key Insights | Metapodcast
- On avoiding human bias in biomarkers: Using self-supervised learning on multimodal data allows models to autonomously discover hidden cancer subtypes and genetic patterns that human-defined approaches have previously missed. — Reference: 🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik - Deep Analysis & Key Insights | Metapodcast
- On validating models with known handholds: When building entirely new systems, it is strategic to start with diseases like lung cancer that have established biological "handholds," allowing you to verify that self-supervised models are actually learning meaningful biology. — Reference: Lessons from Ron Alfa, Co-Founder and CEO of Noetik, on Bringing AI-Native Precision Oncology to Every Cancer Patient
- On scaling in vivo validation: To validate human-trained models without relying on simple cell lines, researchers can use multiplexed CRISPR platforms to inject numerous barcoded cancer variants into a single mouse, generating hundreds of distinct tumors to test predictions against known biology. — Reference: Ron Alfa — Latent Space Summary
Part 3: Generating Fit-for-Purpose Data
- On the inadequacy of public datasets: Publicly available datasets are insufficient for training frontier biological models because they lack the multimodal, high-fidelity human tumor data required to understand real disease biology. — Reference: Interview Questions for Ron Alfa, NOETIK - PMWC Precision Medicine World Conference
- On integrating biological scales: To train models that truly understand human biology, datasets must span multiple scales, incorporating information from the tissue level down to protein architecture, cells, and the genome. — Reference: Foundation Model Series: Harnessing Multimodal Data to Advance Immunotherapies with Ron Alfa from Noetik
- On controlling for batch effects: To prevent models from learning technical noise, patient samples should be distributed across dozens of slides and processed in different batches, allowing biological signals to be statistically disentangled from instrument drift or staining differences. — Reference: 🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik - Deep Analysis & Key Insights | Metapodcast
- On the timeline of data generation: Collecting the necessary data is a massive undertaking; an AI biotech startup must often spend up to two years establishing pipelines and generating fit-for-purpose data before accumulating enough volume to train a functional model. — Reference: Lessons from Ron Alfa, Co-Founder and CEO of Noetik, on Bringing AI-Native Precision Oncology to Every Cancer Patient
- On the consequences of data scarcity: Biological foundation models require a critical threshold of data to succeed; dropping the training data volume to even 10-40% can cause substantial failures in the model's ability to generalize to new cancer types. — Reference: Ron Alfa — Latent Space Summary
- On adjusting data modality assumptions: Real-world data collection requires iteration; early assumptions that protein-level data would be most predictive may give way to transcriptomic or cellular data if those prove easier to interpret and validate. — Reference: Lessons from Ron Alfa, Co-Founder and CEO of Noetik, on Bringing AI-Native Precision Oncology to Every Cancer Patient
Part 4: Spatial Context and Model Architecture
- On the necessity of spatial context: Biology cannot be fully understood by examining cells in isolation, because a cell's function and behavior change significantly depending on its physical location and its interactions with neighboring cells. — Reference: Interview Questions for Ron Alfa, NOETIK - PMWC Precision Medicine World Conference
- On the complexity of spatial transcriptomics: Spatial transcriptomic mapping allows models to observe tens of thousands of genes in their native architecture, treating each data point more like a 20,000-channel image than a standard visual input. — Reference: Ron Alfa — Latent Space Summary
- On the accessibility of H&E inference: While models require expensive, high-density multimodal data for training, they can be designed to run inference solely on standard H&E stains, democratizing access to high-resolution biological insights for virtually every patient. — Reference: 🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik - Deep Analysis & Key Insights | Metapodcast
- On autoregressive scaling in biology: Applying next-token autoregressive training to spatial transcriptomics mimics the scaling behavior of large language models, indicating that biological models improve dynamically as they are fed more data. — Reference: Ron Alfa — Latent Space Summary
- On tissue context length: Just as language models need longer context windows to understand complex documents, biological foundation models require longer tissue context lengths—observing larger tissue regions simultaneously—to capture nonlinear spatial relationships. — Reference: Ron Alfa — Latent Space Summary
Part 5: Startup Building and the Era of Simulation
- On shifting experimentation to the GPU: By building virtual cells and patients, researchers can simulate how a tissue will respond to a perturbation in silico, moving the core of experimental testing from the physical wet lab to GPU clusters. — Reference: Interview Questions for Ron Alfa, NOETIK - PMWC Precision Medicine World Conference
- On recognizing early signals: When building unproven technologies, founders must learn to differentiate between early, imperfect results that are "30% working" and pointing in the right direction, versus approaches that are completely failing. — Reference: Lessons from Ron Alfa, Co-Founder and CEO of Noetik, on Bringing AI-Native Precision Oncology to Every Cancer Patient
- On recognizing data arbitrage: Identifying a compound through independent computational methods that previously showed abandoned activity in humans serves as a powerful proof point and data arbitrage for an AI platform's effectiveness. — Reference: Artificial Intelligence Lights a Beacon to New Medicine for Neurofibromatosis type 2 | by Ron Alfa | Recursion