
Lessons from Ashish Vaswani
Ashish Vaswani is a co-author of Attention Is All You Need, the 2017 paper that introduced the Transformer architecture. The authors contributed equally and their listing order was random. After his research at Google Brain, he co-founded Adept and Essential AI. These lessons cover his research on attention, efficient computation and human–computer collaboration. — WIRED — The Transformer’s Eight Authors.
Part 1: The Origins of the Transformer
- On Machine Context: Word representations need to change with context: the same word can carry different meanings in different sentences. Vaswani describes this shift as a reason to learn contextual rather than fixed representations. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On the Motivation: The original ambition was to read and generate translations in parallel through iterative refinement. That decoding approach proved difficult, so the team kept parallel input processing while returning to autoregressive output generation. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Early Hypotheses: The Transformer made attention the central mechanism for sequence transduction, dispensing with recurrence and convolution in the proposed architecture. — Attention Is All You Need.
- On Collaboration: The Transformer emerged from an eight-author collaboration that combined self-attention ideas, implementation work and rapid experiments. The oral history describes shared offices, hallway conversations and an intensive final sprint—not a lone-inventor breakthrough. — WIRED — The Transformer’s Eight Authors.
- On the Name: Llion Jones proposed the title as a play on the Beatles’ All You Need Is Love. It reflected the team’s decision to rely on attention rather than recurrence or convolution. — WIRED — The Transformer’s Eight Authors.
- On Initial Training Runs: Removing recurrence allowed the model to parallelize computation within training examples. This does not mean that autoregressive translation output is generated all at once. — Attention Is All You Need.
- On Algorithmic Simplicity: The paper sought a simpler sequence-transduction architecture built from attention and position-wise feed-forward layers, rather than recurrent or convolutional components. — Attention Is All You Need.
- On Breaking Paradigms: Without recurrence or convolution, the model needed explicit positional information. The original Transformer added positional encodings to its input embeddings so it could use token order. — Attention Is All You Need.
Part 2: Breaking the Recurrence Bottleneck
- On Sequential Processing: Recurrent models compute each hidden state from earlier states, limiting parallelization within a training example. The paper identifies this as a particular constraint for longer sequences and memory-limited batching. — Attention Is All You Need.
- On Computational Efficiency: A self-attention layer connects positions with a constant number of sequential operations, whereas a recurrent layer requires operations proportional to sequence length. This is a parallelism advantage, not a claim that full attention has linear computational cost. — Attention Is All You Need.
- On Path Lengths: Self-attention shortens the computational paths between distant positions. The authors argue that shorter paths make long-range dependencies easier to learn. — Attention Is All You Need.
- On Matrix Multiplications: Dot-product attention can be expressed as matrix multiplication, letting the implementation use optimized accelerator kernels. Vaswani describes compatibility with existing hardware as an important practical reason for that formulation. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Hardware Utilization: An architecture’s useful operations must also run efficiently on real accelerators. Vaswani emphasizes that hardware constraints influence which attention formulations are practical; he does not present the Transformer as a design that maximizes every device’s throughput. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Training Speed: In the paper’s 2017 translation experiments, the Transformer achieved strong results with substantially less training cost than the compared systems. Those task-specific results should not be read as a universal speed guarantee. — Attention Is All You Need.
- On Self-Attention Utility: Encoder self-attention lets each token form a representation by drawing on other positions in the input through content-based addressing. Decoder attention is masked, so it cannot use future output tokens. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Future Hardware: Hardware improvements can change the practicality of an algorithm. Vaswani notes that greater memory bandwidth might make previously difficult sparse-attention approaches more feasible. — Stanford CS25 — Vaswani’s Transformer Lecture.
Part 3: The Philosophy of Attention
- On Information Retrieval: Attention maps a query and a collection of key-value pairs to a weighted combination of values. Compatibility between each query and key determines the weights. — Attention Is All You Need.
- On Multi-Head Attention: Multi-head attention applies separate learned projections to queries, keys and values, allowing different heads to attend to different positions and representation subspaces. — Attention Is All You Need.
- On Subspace Representations: A single attention head can blend information that needs to remain distinct. Vaswani explains multi-head attention as a way to select different parts of the input independently instead of averaging them together. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Scaled Dot-Product Attention: Scaled dot-product attention divides scores by the square root of the key dimension. The paper motivates this by the risk that large dot products push softmax into regions with very small gradients. — Attention Is All You Need.
- On Interpretability: Attention visualizations can suggest what individual heads are doing, but they are not a complete explanation of a model’s decisions. Vaswani cautions against reading too much into these patterns. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Syntactic Structure: The paper’s attention visualizations show heads apparently tracking long-distance dependencies, anaphora and sentence structure. These examples suggest differentiated behavior, rather than proving that every head has a clear linguistic role. — Attention Is All You Need.
- On Masking: Decoder self-attention masks subsequent positions. Combined with shifted output embeddings, this ensures that a prediction depends only on already-known output positions. — Attention Is All You Need.
- On Feed-Forward Networks: Each attention layer is paired with a position-wise feed-forward network comprising two linear transformations and a ReLU between them. The same transformation is applied separately at each position within a layer. — Attention Is All You Need.
- On Generalization: Vaswani describes attention as a general content-based mechanism rather than a text-only operation. His lecture discusses image attention and music generation as examples, while retaining the need for suitable positional structure. — Stanford CS25 — Vaswani’s Transformer Lecture.
Part 4: The Path to Adept AI and Human-Machine Collaboration
- On Building Tools: Vaswani frames tool-using models as a way to strengthen human–machine collaboration. Useful software interfaces can connect models to capabilities people have already built. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Action Models: At Adept, the proposed next step beyond language generation was a Transformer trained to take actions in software. ACT-1 demonstrated browser actions such as clicking, typing and scrolling; these were capability previews, not proof of universal automation. — Adept — ACT-1: Transformer for Actions.
- On Software Interfaces: Natural-language interfaces can let a user state a goal instead of navigating every software control manually. Adept presented this as an ambition for action models, illustrated with existing tools. — Adept — ACT-1: Transformer for Actions.
- On Enterprise Use Cases: Vaswani described enterprise workflows as an initial application area, starting with company data and helping more people ask useful analytical questions. That is a product direction, not a claim that enterprise automation always has the highest return. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On User Intent: An action model must connect a high-level request with operations in the relevant tool. Adept’s preview shows contextual intent inference and identifies asking clarifying questions as a future improvement. — Adept — ACT-1: Transformer for Actions.
- On Knowledge Work: Adept’s vision was for action models to collaborate with people in their existing software. The release presented an AI teammate as an ambition and acknowledged that ACT-1 did not yet know how to do everything. — Adept — ACT-1: Transformer for Actions.
- On Multimodality: ACT-1 combined a user’s instruction with observations of the browser viewport and actions on available UI elements. This gave the model access to the state of the interface, rather than relying on text instructions alone. — Adept — ACT-1: Transformer for Actions.
- On Practical Applications: Product use can expose capability gaps and provide feedback for improving a model. Vaswani argues that interaction with real workflows can inform training in ways that isolated development may miss. — Stanford CS25 — Vaswani’s Transformer Lecture.
Part 5: Rethinking Scaling and the Compute Paradigm
- On Premature Scaling: Scaling decisions should be tested rather than assumed to transfer. In work co-authored by Vaswani, strategies effective in one compute region did not necessarily generalize to larger models, and pretraining quality did not reliably predict downstream performance. — Scale Efficiently.
- On Compute Dependence: An AI race can push companies to scale established ideas while neglecting riskier alternatives. Vaswani argues for keeping those alternative research paths alive rather than treating additional compute as the only route forward. — Economic Times — Essential AI Interview.
- On Efficiency: Model shape can improve efficiency without simply increasing size. The co-authored study found redesigned configurations that matched downstream quality with fewer parameters and faster training in its tested T5 setting. — Scale Efficiently.
- On Data Quality: Vaswani identifies data as another source of improvement alongside architecture. His lecture describes large models trained on carefully curated text and leaves room for further gains from better data; it does not establish a universal law tying data quality to model size. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Architectural Complacency: The popularity of Transformers is not a reason to restrict research to them. Vaswani encourages exploring unsolved problems and leaves open the possibility of future architectures with better scaling properties. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Post-Training: Vaswani’s stated working hypothesis is that much of a model’s ability is acquired during pretraining, with reinforcement learning amplifying abilities already present. He therefore emphasizes improving the base model, rather than claiming equal importance for all training stages. — AMD — Vaswani on Open Models.
- On Hardware Constraints: Reducing mathematical operations is not enough if memory movement dominates runtime. Vaswani discusses attention’s memory traffic and the communication requirements of large distributed models as practical design constraints. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Diminishing Returns: Evaluate compute by useful performance relative to total cost of ownership. Vaswani’s considerations include interconnects, software flexibility, expected experiments and reliability—not merely the number of accelerators purchased. — Accel — Vaswani on Compute and Research.
- On Evaluation: Evaluation must measure the behavior that matters downstream. The co-authored scaling study found that strong pretraining perplexity could be a misleading guide to fine-tuned task quality; this supports checking actual task performance, not the original claim about benchmark saturation. — Scale Efficiently.
Part 6: Building Essential AI and Open Models
- On Open Collaboration: Vaswani argues that limited collaboration can concentrate AI capabilities. He favors open frontier-model work that others can inspect and build on, while presenting openness as a strategy rather than a guarantee against inequality. — Economic Times — Essential AI Interview.
- On Defining Reasoning: Vaswani describes reasoning research in terms of a model’s ability to iterate on its thinking. He says the team studied how this behavior develops during pretraining, treating that account as its research hypothesis rather than a settled definition of reasoning. — AMD — Vaswani on Open Models.
- On Enterprise Integration: Action models can work through existing software interfaces. Adept’s preview used a browser extension and demonstrated workflows spanning multiple programs; it did not promise that enterprise deployment requires no infrastructure changes. — Adept — ACT-1: Transformer for Actions.
- On Purpose-Built Models: Essential AI’s stated focus was frontier capability in software engineering and scientific discovery. This is a case for choosing target domains, not evidence that every enterprise task should use a smaller model. — AMD — Vaswani on Open Models.
- On Collaborative Ecosystems: Publishing research lets others extend an idea beyond what its original team could achieve. Vaswani credits open exchange with the continuing development of Transformer-based methods and worries that closed advances would slow collective progress. — AMD — Vaswani on Open Models.
- On System Architecture: Large models require attention to the whole computing system, not just neural-network architecture. Vaswani describes distributed Transformers in terms of interconnects, congestion and infrastructure as well as model design. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Custom Silicon: Vaswani describes close collaboration with AMD on making large-scale training work on MI300X clusters. The example concerns accelerator integration and software feedback; it does not show that Essential AI designed custom silicon. — AMD — Vaswani on Open Models.
- On Accessibility: Vaswani argues that people should be able to use powerful AI tools to address their own problems. He favors open models for that purpose, without claiming that openness alone eliminates compute costs. — Essential AI’s Open-Science Vision — Interview.
Part 7: Advice for Future Researchers
- On Long-term Research: Short-term commercial pressure can displace longer-term research. Vaswani warns about companies redirecting resources from R&D toward immediate revenue and argues for maintaining long-term bets alongside a sustainable business. — Economic Times — Essential AI Interview.
- On First Principles: Attention asks which information a position should draw from based on content. Vaswani explains the mechanism through content-based addressing, contrasting it with fixed local interactions. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Simplicity: Vaswani finds autoregressive models’ simplicity compelling even while considering iterative-refinement and diffusion approaches. He treats simplicity as a reason to keep exploring those models, not as proof that complex architectures cannot be adopted. — AMD — Vaswani on Open Models.
- On Choosing Problems: Research plans should inform compute planning: anticipated model sizes and the experiments needed to answer a question help determine resource needs and how those needs will grow. Vaswani does not prescribe choosing problems just beyond current hardware. — Accel — Vaswani on Compute and Research.
- On Empirical Validation: The Transformer paper tested architectural choices through ablations, varying attention heads, dimensions and other components. Those experiments provide evidence for specific design decisions rather than relying only on architectural intuition. — Attention Is All You Need.
Part 8: The Broader Trajectory of AI Development
- On AI Apprentices: Vaswani hopes scientists will work with AI models as computational collaborators or apprentices. He also wants knowledge from those interactions to feed back into open models so later users can tackle harder problems. — AMD — Vaswani on Open Models.
- On Scientific Discovery: Scientific discovery is one of Vaswani’s stated priorities for AI development. He discusses possible contributions to mathematics and drug discovery as future aims, not demonstrated discoveries or a dismissal of language generation. — AMD — Vaswani on Open Models.
- On General Intelligence: Vaswani identifies reasoning and planning as areas where adaptive inference and external planners may help. He asks whether smaller models could match larger ones by doing more thinking, leaving the answer open rather than asserting a proven recipe for general intelligence. — Stanford CS25 — Vaswani’s Transformer Lecture.
- On Agentic Behavior: Adept’s ACT-1 preview showed repeated observations and actions used to pursue one goal, including tasks spanning several tools. It illustrated multi-step agent behavior while acknowledging limitations and the need for human feedback. — Adept — ACT-1: Transformer for Actions.
- On Modality Convergence: Vaswani recalled an early ambition to bring speech, audio and vision under a single architecture. The paper’s future-work section likewise proposed exploring non-text modalities; this was a research direction, not a guarantee that all modality distinctions will disappear. — WIRED — The Transformer’s Eight Authors.
- On the Future: Vaswani describes AI as still early despite years of capability gains. His outlook leaves substantial research to do, including understanding what models learn and developing new methods for learning. — AMD — Vaswani on Open Models.