ML Engineer, Training Infrastructure for Foundational Models at Sarvam AI (Bengaluru)

Application ends: November 7, 2026
Apply Now

Job Description

Sarvam AI is hiring an ML engineer for LLM training infrastructure in Bengaluru. You will own the systems that Sarvam’s next family of foundational models is trained on. The work covers distributed training on large GPU clusters, parallelism, GPU kernels and the reliability of long training runs.

At a glance

  • Position: ML Engineer (Training Infra), Foundational Models
  • Where: Bengaluru, Karnataka, India
  • Type: Full time
  • Experience: 3 or more years; strong early-career candidates are also considered
  • Apply: open until filled (re-checked 7 November 2026)

What the training infrastructure engineer does

  • Build, run and keep improving the distributed training stack across large GPU clusters
  • Design data, tensor, pipeline, sequence and expert parallelism, and decide which mix suits which model and scale
  • Profile and speed up training end to end: kernels, communication overlap, memory layout, checkpointing and data loading
  • Write and tune custom GPU kernels in CUDA or Triton where standard ones are too slow
  • Keep long training jobs reliable: fault tolerance, checkpoint integrity, clean restarts and spotting slow or faulty nodes
  • Work with researchers so new model designs can be trained efficiently

Why this role matters

Few teams in India train their own large foundation models, and Sarvam is one of them. Training at this scale is a hard systems problem: a slow data loader or a bad node can waste weeks of GPU time. This role gives you hands-on work that is usually done only at a handful of labs abroad.

What Sarvam is looking for

  • A BS or MS in computer science or a close field, or equal proven experience
  • 3 or more years building ML training infrastructure or large distributed systems
  • Hands-on experience training large models with Megatron-LM, DeepSpeed, FSDP, NeMo or similar, including being on call for a real pretraining run
  • Deep knowledge of GPU architecture, the CUDA programming model and GPU profiling tools such as Nsight and the PyTorch profiler
  • Strong knowledge of PyTorch internals
  • Meaningful open-source contributions to training infrastructure projects such as Megatron, DeepSpeed, PyTorch, vLLM, Triton or NCCL

Nice to have

  • Custom CUDA or Triton kernels with measured gains on real workloads
  • Training models of 10B+ parameters or on clusters of 1,000+ GPUs
  • Slurm or Kubernetes at scale
  • Mixed precision (BF16, FP8) or quantisation-aware training in production
  • First-author papers or technical reports on training systems or model efficiency

Pay and location

The posting does not state pay. The role is based in Bengaluru.

How to apply

Apply through the official Sarvam job posting. Sarvam’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 7 November 2026. Check the posting for location and work-authorisation details before you apply.

See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.

Hiring institution: Sarvam AI

Official advertisement: jobs.ashbyhq.com

How to prepare for this application

  • List your open-source pull requests to training frameworks with links, near the top of your CV
  • Prepare numbers from a real run you worked on: model size, GPU count, MFU and what you changed to raise it
  • Revise how tensor, pipeline and expert parallelism split memory and communication, and when each one wins
  • Be ready to debug a slow training step on a whiteboard: data loading, NCCL traffic, kernel time or a straggler node
  • Learn about Sarvam's published Indian-language models before the interview

About Sarvam AI

Sarvam AI is a Bengaluru company that builds AI models and products for India, with a focus on Indian languages. Its posting says it is backed by Lightspeed, Peak XV and Khosla Ventures.

We send one confirmation email first. Every alert has an unsubscribe link.