ML Engineer (Data), Foundational Models at Sarvam AI (Bengaluru)

Application ends: October 12, 2026
Apply Now

Job Description

Sarvam is hiring an ML Engineer to own the data infrastructure for its next foundation models in Bengaluru: petabyte-scale curation and filtering pipelines, and the systems that decide what goes into a training run, in what proportion and in what order.

At a glance

  • Role: ML Engineer (Data), Foundational Models
  • Where: Bengaluru, India
  • Experience: 3+ years building large-scale data systems (exceptional early-career candidates considered)
  • Focus: pre-training and post-training data pipelines, quality filtering and mixture design
  • Apply: open until filled (re-checked 12 October 2026)

About the training data role

  • Build pre-training and post-training pipelines at petabyte scale: ingestion, parsing, normalisation, filtering, deduplication, tokenisation and packing
  • Develop quality filtering, including model-based quality classifiers and contamination detection
  • Own data mixture design, curriculum and annealing strategies with the research team
  • Build tools that let researchers analyse, slice, attribute and debug the data
  • Scale the pipeline to multilingual corpora, code, maths and multi-source web data

Indian-language AI is a distinct research problem: many languages have little training data, scripts vary widely, and users mix languages in a single sentence. Building models, evaluations and serving systems that work well for these users is what makes roles at an Indian foundation-model company different from similar roles elsewhere.

What Sarvam is looking for

  • A BS or MS in computer science or a related field, or equivalent experience
  • 3+ years building large-scale data systems such as petabyte-scale or distributed pipelines
  • Hands-on data curation and filtering for LLM training, and the ability to defend the choices behind a corpus you built
  • Deep familiarity with Spark, Ray, Beam, Dask or similar, and the storage underneath them
  • Strong Python, comfort with tokenisation, sharding, packing and IO performance, and open-source contributions in data tooling

Nice to have

  • Building or working with large open pre-training corpora
  • Multilingual data collection, normalisation, quality scoring and mixing
  • Model-based quality classifiers, contamination detection or data attribution
  • Tokenisation research, and first-author papers or technical reports

Pay and location

Sarvam’s posting does not state a salary. The role is based in Bengaluru.

How to apply

Apply through the official Sarvam job posting. Sarvam’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.

See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.

Hiring institution: Sarvam AI

Official advertisement: jobs.ashbyhq.com

How to prepare for this application

  • Corpus walk-through: prepare to explain a dataset you built end to end and why each filter was there.
  • Deduplication: revise MinHash and near-duplicate detection at scale.
  • Quality classifiers: know how model-based filters are trained and how they can bias a corpus.
  • Indic data: think about script normalisation and quality scoring for many Indian languages.
  • Mixtures: have a view on curriculum and annealing, and how you would measure their effect.

About Sarvam

Sarvam is building what it calls the bedrock of sovereign AI for India: a full-stack platform spanning research, foundation models, infrastructure and applications, with a focus on making AI work for India's languages and institutions. Headquartered in Bengaluru, it works with leading enterprises and public institutions, is backed by Lightspeed, Peak XV and Khosla Ventures, and partners with brands such as Tata Capital, SBI Life, CRED, IDFC and LIC.

We send one confirmation email first. Every alert has an unsubscribe link.