
Lessons from Aakanksha Chowdhery
Aakanksha Chowdhery is an AI researcher whose work spans wireless systems, large-scale language models, multimodal and embodied AI, and reinforcement learning for autonomous coding. She was the first author of PaLM, co-authored PaLM-E, worked as a lead researcher on Gemini, and now develops autonomous coding systems at Reflection AI. — AI Engineer — RL for Autonomous Coding.
Part 1: Scaling Large Language Models
- Scale Requires Systems Co-Design: PaLM showed that reaching 540 billion parameters was not simply a modeling decision: the model, data pipeline, and distributed training system had to be designed together. — PaLM: Scaling Language Modeling with Pathways.
- Distribute Work Without Losing Utilization: Training across 6,144 TPU v4 chips made communication and placement first-class design constraints; scale is useful only when the accelerators remain productively occupied. — Stanford MLSys Seminar #69 — PaLM.
- Scale Can Improve Few-Shot Learning: PaLM’s larger variants improved across many few-shot tasks, reducing the amount of task-specific tuning needed for new problems. — PaLM: Scaling Language Modeling with Pathways.
- Watch for Discontinuous Gains: Some BIG-bench tasks improved sharply only at the largest scale, a reminder that capability curves can contain thresholds rather than smooth progress. — PaLM: Scaling Language Modeling with Pathways.
- Treat Data Mixture as a Design Choice: PaLM combined filtered web pages, books, Wikipedia, conversations, and code; the training mixture and vocabulary were deliberate parts of the system, not incidental inputs. — Google Research — PaLM.
- Evaluation Must Expand with Capability: As models gain reasoning, multilingual, and code abilities, evaluation has to cover a broader set of tasks and include responsible-AI analysis rather than rely on one aggregate benchmark. — PaLM: Scaling Language Modeling with Pathways.
- Expect Instability to Appear at the Largest Scale: PaLM’s 540-billion-parameter run encountered roughly twenty irregular loss spikes that did not appear in the smaller models, and the team had no principled general fix during that training run. — PaLM: Scaling Language Modeling with Pathways.
- Measure Useful Hardware Work: PaLM reported 57.8% hardware FLOPs utilization at unprecedented scale, showing why utilization—not chip count alone—is a meaningful systems metric. — Google Research — PaLM.
- Audit Memorization Alongside Generalization: The PaLM work measured training-data memorization as scale increased instead of treating benchmark gains as sufficient evidence of generalization. — PaLM: Scaling Language Modeling with Pathways.
Part 2: The PaLM and Pathways Architecture
- Infrastructure Can Remove Cluster Boundaries: Pathways enabled one training job to span two TPU v4 Pods, letting PaLM use 6,144 chips as a coordinated system. — Google Research — PaLM.
- Separate Orchestration from Accelerator Execution: Pathways coordinates distributed computation while accelerator programs perform the dense numerical work, allowing the training system to manage a much larger topology. — Stanford MLSys Seminar #69 — PaLM.
- Hide Input and Communication Latency: At large scale, data delivery and collective communication must overlap useful computation so expensive accelerators are not left waiting. — Stanford MLSys Seminar #69 — PaLM.
- Make Checkpoint Restarts Deterministic: PaLM’s training framework was designed so that restarting from an earlier checkpoint reproduced the subsequent training trajectory bit for bit. — PaLM: Scaling Language Modeling with Pathways.
- Combine Parallelism Strategies: PaLM used model and data parallelism together because neither alone efficiently fits and trains a model of this size. — Google Research — PaLM.
- Design Around the Physical Topology: The placement of model partitions and communication across chips and Pods materially affects achievable training efficiency. — Stanford MLSys Seminar #69 — PaLM.
- Trade Stored Activations for Recomputation: PaLM saved one fully sharded activation tensor per layer and rematerialized the remaining activations during backpropagation because that delivered better throughput at larger batch sizes. — PaLM: Scaling Language Modeling with Pathways.
- Engineer Long Runs for Failure: Large distributed training runs must tolerate accelerator and infrastructure failures without losing the entire run. — Stanford MLSys Seminar #69 — PaLM.
Part 3: Multimodal AI and Gemini
- Build Multimodality into the Model: Gemini was trained jointly across text, image, audio, and video rather than treating non-text modalities only as add-ons. — Gemini: A Family of Highly Capable Multimodal Models.
- Ground Reasoning in Multiple Modalities: A multimodal model should connect language with visual and other signals so it can reason about information that is not present in text alone. — Stanford MLSys Seminar #90 — PaLM-E and Gemini.
- Preserve Information Lost by Transcription: Gemini can ingest audio features directly, allowing it to retain nuances that may disappear when speech is first converted into text. — Gemini: A Family of Highly Capable Multimodal Models.
- Use Modality-Specific and Combined Evaluations: Gemini was evaluated across text, code, image, audio, and video benchmarks because multimodal capability cannot be summarized by a text score. — Gemini: A Family of Highly Capable Multimodal Models.
- Match Model Size to Deployment Context: Gemini’s Ultra, Pro, and Nano variants were designed for different capability and efficiency envelopes, from data centers to devices. — Gemini: A Family of Highly Capable Multimodal Models.
- Train on Interleaved Context: Interleaved multimodal inputs let a model learn relationships among text, images, audio, and video within one sequence. — Gemini: A Family of Highly Capable Multimodal Models.
- Video Adds Temporal Reasoning: Understanding video requires relating information across frames rather than treating each frame as an isolated image. — Gemini: A Family of Highly Capable Multimodal Models.
- Allocate Visual Compute Where Detail Demands It: Gemini supports variable input resolution so finer-grained visual tasks can receive more computation. — Gemini: A Family of Highly Capable Multimodal Models.
Part 4: Embodied AI and Robotics
- Ground Language in Perception: PaLM-E connects a language model to continuous sensor inputs so words and plans can be grounded in what a robot perceives. — PaLM-E: An Embodied Multimodal Language Model.
- Encode Sensors as Multimodal Sentences: PaLM-E interleaves visual, state-estimation, and text encodings in a shared sequence that the model can process end to end. — PaLM-E: An Embodied Multimodal Language Model.
- Update Plans from Observations: Embodied reasoning depends on repeated observations of the world, not a single static prompt, so plans can reflect the current state. — Stanford MLSys Seminar #90 — PaLM-E and Gemini.
- Joint Training Can Transfer Across Domains: PaLM-E benefited from joint training across robotics, language, vision, and vision-language data, showing positive transfer rather than simple interference. — PaLM-E: An Embodied Multimodal Language Model.
- Separate High-Level Decisions from Motor Control: PaLM-E generates textual decisions or subgoals that condition existing low-level policies rather than directly replacing the robot’s motor controller. — PaLM-E: An Embodied Multimodal Language Model.
- Close the Loop with Failure Detection: PaLM-E uses current observations to detect whether a skill succeeded and can replan after disturbances or low-level-policy failures. — PaLM-E: An Embodied Multimodal Language Model.
- One Model Can Serve Multiple Embodiments: The same PaLM-E model handled several embodied tasks, observation modalities, and robot embodiments instead of requiring one model per platform. — PaLM-E: An Embodied Multimodal Language Model.
- Preserve General Abilities While Adding Grounding: At larger scale, PaLM-E added embodied and visual capabilities while retaining general language performance. — PaLM-E: An Embodied Multimodal Language Model.
Part 6: Agentic AI and Autonomous Systems
- Agents Need Different Training Signals: Systems that act over multiple steps need training data and objectives that reward useful trajectories, not only plausible next-token predictions. — TWIML — Rethinking Pre-Training for Agentic AI.
- Train for Recovery, Not Only First-Try Success: Agentic competence includes noticing an error, using feedback, and trying a better path rather than merely producing one polished response. — TWIML — Rethinking Pre-Training for Agentic AI.
- Tool Use Must Be Learned in Context: An agent has to decide when and how to invoke tools, interpret their outputs, and incorporate the result into the next action. — TWIML — Rethinking Pre-Training for Agentic AI.
- Trajectories Are Training Data: For agents, the sequence of observations, actions, feedback, and revisions can be more informative than the final answer alone. — TWIML — Rethinking Pre-Training for Agentic AI.
- Pre-Training Should Prepare Models to Act: If agency is postponed entirely to post-training, the base model may lack representations needed for interaction, planning, and recovery. — TWIML — Rethinking Pre-Training for Agentic AI.
- Long Horizons Compound Small Errors: As workflows lengthen, small mistakes accumulate; evaluation and training must therefore cover extended sequences rather than isolated steps. — TWIML — Rethinking Pre-Training for Agentic AI.
- Evaluate in Interactive Environments: Static question sets miss how an agent changes its environment, reacts to feedback, and recovers during a task. — TWIML — Rethinking Pre-Training for Agentic AI.
Part 7: Reinforcement Learning for Coding
- Exploit Execution Feedback: Code is attractive for reinforcement learning because tests and execution can provide concrete feedback on whether a candidate solution works. — AI Engineer — RL for Autonomous Coding.
- Verifiability Makes Rewards Scalable: Domains such as code and mathematics support automated checks, making it possible to reward outcomes without hand-labeling every reasoning step. — AI Engineer — RL for Autonomous Coding.
- Spend Inference Compute on Alternatives: Sampling multiple candidate solutions can uncover a correct path that one-shot generation misses, provided the system can select or verify it. — AI Engineer — RL for Autonomous Coding.
- Search Is Useful Only with Selection: Generating more attempts raises coverage, but the system still needs a reliable signal to identify the useful solution among them. — AI Engineer — RL for Autonomous Coding.
- Autonomous Coding Is Bigger Than Code Generation: Real software engineering requires a system to combine reasoning, tools, context, verification, and multiple model calls across a workflow. — AI Engineer — RL for Autonomous Coding.
- Reinforcement Learning Is Also a Systems Problem: RL at scale coordinates generation and optimization loops, multiple model copies, and accelerator placement; algorithmic gains depend on infrastructure efficiency. — AI Engineer — RL for Autonomous Coding.
- Train in Imperfect Environments: Real engineering rollouts contain incomplete signals and messy outcomes, so autonomous systems must learn beyond perfectly specified benchmark tasks. — AI Engineer — RL for Autonomous Coding.
Part 8: Early Research: Networks and Wireless Systems
- Design the Whole Wireless Video System: The surveillance work treated cameras, wireless links, scheduling, and video delivery as one deployed system rather than optimizing a single isolated component. — The Design and Implementation of a Wireless Video Surveillance System.
- Mobility Changes the Channel: Measurements across more than twenty flights showed that drone movement creates time- and frequency-selective wireless conditions for ground clients. — Aerial Channel Prediction and User Scheduling in Mobile Drone Hotspots.
- Predict Before Scheduling: The drone-hotspot system predicts per-client subcarrier signal quality before choosing which users to serve. — Aerial Channel Prediction and User Scheduling in Mobile Drone Hotspots.
- Optimize Scheduling for Network Utility: The scheduling method selects clients using predicted signal quality to maximize overall network utility rather than serving users blindly. — Aerial Channel Prediction and User Scheduling in Mobile Drone Hotspots.
- Model the Physical Cause of Variation: The aerial-channel model uses constructive and destructive interference among line-of-sight and reflected paths, tying the algorithm to measured propagation behavior. — Aerial Channel Prediction and User Scheduling in Mobile Drone Hotspots.