Job Description
Sarvam is hiring an ML Researcher for its foundational models in Bengaluru, to answer the open questions that decide what its next models look like: architecture, optimisation, data composition and training dynamics, tested through ablations at meaningful scale and turned into decisions for production training runs.
At a glance
- Role: ML Researcher, Foundational Models
- Where: Bengaluru, India
- Level: PhD in ML, CS or a related field, with 3+ years of post-PhD research (exceptional early-career candidates considered)
- Focus: pre-training research: architecture, optimisation, scaling and post-training recipes
- Apply: open until filled (re-checked 12 October 2026)
About the foundation model research role
- Drive open-ended research on architecture, optimisation, scaling behaviour, training stability and post-training recipes
- Design and run ablations at scales that inform large-run decisions, including end-to-end pre-training experiments
- Turn findings into concrete proposals for the next training run and own them through to production
- Work closely with the infrastructure and data teams on questions at that boundary
- Read broadly, write internally and publish externally when the work merits it
Indian-language AI is a distinct research problem: many languages have little training data, scripts vary widely, and users mix languages in a single sentence. Building models, evaluations and serving systems that work well for these users is what makes roles at an Indian foundation-model company different from similar roles elsewhere.
What Sarvam is looking for
- A PhD in machine learning, computer science or a closely related field (or nearly complete)
- 3+ years of post-PhD research experience or equivalent depth; exceptional early-career candidates with a strong record will be considered
- First-author papers at top ML venues such as NeurIPS, ICML, ICLR, ACL, EMNLP or COLM
- Hands-on experience pre-training transformer language models from scratch, ideally at 7B+ parameters, including debugging a run you owned
- Meaningful open-source LLM contributions, fluent PyTorch and distributed training, and strong experimental judgement
Nice to have
- Novel architectures such as mixture-of-experts, state-space or hybrid models
- Multilingual or multimodal pre-training
- Post-training research (RLHF, RLVR, distillation, reasoning models)
- Taking a research idea from prototype to a shipped capability
Pay and location
Sarvam’s posting does not state a salary. The role is based in Bengaluru.
How to apply
Apply through the official Sarvam job posting. Sarvam’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.
See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.
Hiring institution: Sarvam AI
Official advertisement: jobs.ashbyhq.com
How to prepare for this application
- Own a training run: prepare a detailed story of a pre-training run you ran, what broke and how you fixed it.
- Scaling: know scaling-law basics and how to decide whether a small-scale result will hold.
- Architecture choices: have views on MoE, state-space and hybrid models for multilingual data.
- Disagree with evidence: the posting expects you to defend your reasoning; practise doing that with data.
- Open source: link your LLM code, models or datasets.
About Sarvam
Sarvam is building what it calls the bedrock of sovereign AI for India: a full-stack platform spanning research, foundation models, infrastructure and applications, with a focus on making AI work for India's languages and institutions. Headquartered in Bengaluru, it works with leading enterprises and public institutions, is backed by Lightspeed, Peak XV and Khosla Ventures, and partners with brands such as Tata Capital, SBI Life, CRED, IDFC and LIC.