Job Description
Sarvam is hiring a Researcher in vision-language models (VLMs) in Bengaluru, to work across the whole model lifecycle, from data and training to evaluation and production, with a particular focus on multilingual models and Indic multimodal tasks.
At a glance
- Role: Researcher, Vision
- Where: Bengaluru, India
- Focus: multilingual vision-language models and Indic multimodal benchmarks
- Background: track record of good research through publications, technical reports or shipped work
- Apply: open until filled (re-checked 12 October 2026)
About the vision research role
- Research vision-language architectures: encoders, fusion mechanisms, pretraining objectives and scaling
- Design training methods (pretraining, SFT, RLHF, DPO) adapted for multilingual VLMs
- Investigate data mixtures, quality signals and synthetic data approaches that improve models
- Build evaluation frameworks and benchmarks, especially for Indic multimodal tasks, and study failure modes, robustness and interpretability
- Prototype with engineers so ideas can be tested at scale, and contribute to open source and research collaborations
Indian-language AI is a distinct research problem: many languages have little training data, scripts vary widely, and users mix languages in a single sentence. Building models, evaluations and serving systems that work well for these users is what makes roles at an Indian foundation-model company different from similar roles elsewhere.
What Sarvam is looking for
- Deep understanding of vision-language models: training dynamics, architecture trade-offs and failure modes
- A track record of good research through publications, technical reports or impactful shipped work
- Rigorous experimental design and strong PyTorch skills, running experiments end to end
- Willingness to work across data, training and evaluation problems
Nice to have
- A PhD or Master’s with relevant research in ML, computer vision, NLP or a related field
- Papers at A/A* venues
- Multilingual or low-resource language modelling
- Document understanding, OCR or structured visual prediction, and large-scale data curation
Pay and location
Sarvam’s posting does not state a salary. The role is based in Bengaluru.
How to apply
Apply through the official Sarvam job posting. Sarvam’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.
See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.
Hiring institution: Sarvam AI
Official advertisement: jobs.ashbyhq.com
How to prepare for this application
- VLM architectures: compare how image encoders are connected to language models and the trade-offs of each approach.
- Post-training: be ready to discuss SFT, RLHF and DPO for multimodal models.
- Indic documents: think about OCR and document understanding for Indian scripts; it is listed as a bonus.
- Benchmarks: propose an Indic multimodal benchmark and how you would avoid contamination.
- Experiments: walk through an ablation you designed and what it changed.
About Sarvam
Sarvam is building what it calls the bedrock of sovereign AI for India: a full-stack platform spanning research, foundation models, infrastructure and applications, with a focus on making AI work for India's languages and institutions. Headquartered in Bengaluru, it works with leading enterprises and public institutions, is backed by Lightspeed, Peak XV and Khosla Ventures, and partners with brands such as Tata Capital, SBI Life, CRED, IDFC and LIC.