Akshat Bubna is the co-founder and CTO of Modal, a code-first, usage-based compute platform for AI and data workloads. Modal developed its own filesystem, container runtime, and scheduler to support rapid deployment and elastic scaling. This profile covers Bubna’s approach to systems engineering, agent sandboxes, and technical talent. — Modal Series B Announcement.

Visual summary of operating lessons from Akshat Bubna.

Part 1: Infrastructure and the Agent Experience

  1. The agent cloud era: Cloud infrastructure designed for human operators is being adapted for agents that write code, run it, inspect results, and revise environments; programmatic interfaces and clear feedback matter more as a result. — Latent Space Interview.
  2. Agent iteration speed: Agents need fast execution and feedback loops, with enough context from outputs and errors to make their next change intelligently. — Latent Space Interview.
  3. Software-defined environments: Modal’s code-first interface lets developers and agents define runtime requirements in code rather than hand-managing large Kubernetes configurations. — Latent Space Interview.
  4. Scaling agent sandboxes: Reinforcement-learning and coding-agent workloads can demand very large numbers of short-lived isolated sandboxes; Bubna describes deployments at the scale of 100,000 sandboxes. — Latent Space Interview.

Part 2: Rethinking Serverless and GPUs

  1. Spiky inference workloads: Inference demand is variable, making elastic capacity more valuable than keeping fixed GPU allocations idle between peaks. — How We Achieved Truly Serverless GPUs.
  2. Static provisioning costs: Provisioning GPUs for peak demand can leave expensive capacity underused during quieter periods. — How We Achieved Truly Serverless GPUs.
  3. The serverless cold-start problem: Elastic GPU inference depends on reducing the time it takes to start a replica and load its model; the coauthored Modal article describes why ordinary model initialization is too slow. — How We Achieved Truly Serverless GPUs.
  4. Pooling compute: Modal’s cofounders describe a global, usage-based compute pool intended to give developers capacity without managing individual machines. — Modal Series B Announcement.
  5. Optimizing GPU startup: Modal’s coauthored technical account describes cloud buffers, a custom filesystem, and CPU/GPU checkpoint restore; one documented workload fell from roughly 2,000 seconds to about 50 seconds of startup time. — How We Achieved Truly Serverless GPUs.
  6. Sharing GPU infrastructure knowledge: In a coauthored technical post, Bubna and colleagues argue that sharing techniques for efficient GPU use can help expand effective capacity rather than weakening Modal’s position. — How We Achieved Truly Serverless GPUs.
  7. Anticipating GPU demand: Bubna says Modal added GPU support before ChatGPT’s launch, initially expecting uses such as computer vision and XGBoost before generative AI demand surged. — Latent Space Interview.
  8. Usage-based infrastructure: Modal’s cofounders positioned the product around charging for compute actually used, rather than asking developers to reserve machines for bursty jobs. — Modal Series B Announcement.

Part 3: Systems Engineering and Building from Scratch

  1. Avoiding legacy abstractions: Bubna argues that Kubernetes was not designed for the scale and placement needs of modern ML and GPU workloads, motivating Modal to build purpose-built infrastructure. — Times of India Interview.
  2. Docker’s startup limitation: Bubna recounts that unpacking large Docker image tarballs before starting workloads was too slow for Modal’s desired container startup path. — Amplify Partners Interview.
  3. A lazy-loading filesystem: Modal used a FUSE-based filesystem to present container files on demand, avoiding a full image transfer before the process can begin. — Amplify Partners Interview.
  4. Iterating infrastructure: Bubna describes developing the filesystem in stages: a simple network mount, a Python prototype, then a Rust implementation once the performance and safety requirements were clear. — Amplify Partners Interview.
  5. Acquiring adjacent systems expertise: Modal’s Jamsocket acquisition brought in a team experienced with stateful runtimes, real-time backends, and secure AI code execution—capabilities adjacent to Modal’s platform. — Jamsocket Is Joining Modal.
  6. Changing compute profiles: Bubna discusses workloads that move between CPU-heavy and GPU-heavy phases, making resource orchestration and efficient handoffs important. — Latent Space Interview.

Part 4: Operating the Agent Cloud at Scale

  1. Observability for agent-written code: As agents generate more code, Bubna expects operators to rely more on observability to understand execution and diagnose failures. — Latent Space Interview.
  2. Aggregating fragmented capacity: Bubna describes Modal operating across 17 cloud and bare-metal providers, using software to present scattered compute capacity as a usable pool. — Latent Space Interview.
  3. Separating queue guarantees: Bubna explains why Modal split its input plane into a low-latency, at-most-once path and a durable path suited to backlog and retry. — Amplify Partners Interview.
  4. Multi-tenant isolation: For untrusted workloads, Bubna describes using gVisor to narrow the system-call surface rather than relying on ordinary containers alone. — Amplify Partners Interview.
  5. Runtime abuse controls: Bubna describes detecting crypto-mining abuse through runtime signals such as system calls, network behavior, and executable fingerprints. — Amplify Partners Interview.
  6. Removing hidden round trips: Bubna’s performance examples emphasize avoiding kernel and disk round trips, keeping useful state in memory, and tuning network behavior. — Amplify Partners Interview.
  7. Portable distributed networking: Bubna explains that multi-node GPU work has to bridge cloud-specific RDMA implementations and their associated networking and startup constraints. — Amplify Partners Interview.
  8. Inference as a systems problem: Bubna frames KV-cache movement and transferring model weights between training and inference as memory-movement and scheduling problems. — Latent Space Interview.
  9. Automating inference optimization: Bubna describes an internal auto-inference harness that sweeps configurations, profiles with NVIDIA tools, and tests GPU choices to automate repeated tuning work. — Latent Space Interview.
  10. Scheduling and price tiers: Bubna connects control of capacity and scheduling to lower-cost batch options for workloads that can tolerate more latency. — Latent Space Interview.
  11. Bounded agent capability: Bubna argues that agents need powerful execution tools with hard permission boundaries rather than relying on a model to enforce access rules. — Latent Space Interview.
  12. Shared responsibility for sandbox security: Modal’s incident note distinguishes an exposed, unauthenticated customer endpoint from a breach of Modal’s isolation boundary; application access controls remain essential. — Modal Note on the Hugging Face Agent Incident.

Part 5: Teams, Talent, and Company Building

  1. Human capital remains a bottleneck: Bubna says building advanced AI infrastructure still depends on highly capable engineers, despite predictions of automation replacing them. — Times of India Interview.
  2. Competitive-programming talent: Bubna says Modal recruits from Olympiad and competitive-programming circles because those backgrounds can develop strong problem-solving skills. — Times of India Interview.
  3. A solvable-problems mindset: Bubna credits programming competitions with teaching a habit of treating difficult problems as tractable. — Times of India Interview.
  4. Incremental development: Bubna links his competitive-programming experience to building a simple working version first and improving it iteratively at Modal. — Times of India Interview.
  5. A culture of hard problems: Bubna’s public description of Modal’s culture emphasizes solving hard problems and building work the team can be proud of. — Latent Space Interview.
  6. Software as an infrastructure moat: Bubna describes Modal as capital-light relative to owners of data centers, with differentiation in software spanning multiple providers. — Latent Space Interview.

Part 6: Scale AI, Career, and Ecosystem Perspective

  1. Learning from Scale AI: Bubna says repeated infrastructure problems at Scale AI helped convince him that existing cloud tools were not built for modern ML work. — Times of India Interview.
  2. An integrated infrastructure stack: Bubna describes Modal building its filesystem, networking, and queue architecture as parts of one compute platform rather than treating each as a separate product. — Amplify Partners Interview.
  3. Self-belief for technical ambition: When advising young engineers, Bubna emphasizes confidence to attempt technically difficult problems. — Times of India Interview.
  4. Strengthening research infrastructure: Bubna calls for stronger deep-tech research infrastructure in India so talented people can pursue advanced work without having to leave for resources. — Times of India Interview.