Job Description
Sarvam AI is hiring a Performance Engineer for on-device inference in Bengaluru to take its models from research hand-off to production-ready artifacts on at least two of its target chipsets: Intel, ARM and Apple processors, and Nvidia or AMD GPUs.
At a glance
- Role: Performance Engineer, On-Device Inference
- Location: Bengaluru
- Experience: 3+ years on ML systems
- Ownership: 1–2 (model, chipset) pairs, end to end
- Apply: open until filled (re-checked 12 October 2026)
What you will do
- Own one or two (model, chipset) pairs from research hand-off to production
- Quantise models, validate accuracy, benchmark and document the results
- Write the deployment workbook for each pair you own
- Embed part-time with the teams that consume the models, debugging performance and accuracy issues alongside them
- Maintain and extend the team’s benchmark harness
On-device inference means running a model on the user’s own hardware, such as a laptop, phone or edge box, rather than in a data centre. Memory, power and speed limits are much tighter there, so the choice of export path, quantisation scheme and runtime decides whether a model is usable, and accuracy has to be re-checked after every compression step. In this role you partner with the app team that consumes each model while it is integrated.
What Sarvam AI is looking for
- 3+ years on ML systems
- Solid PyTorch and ONNX export experience, including the tricky parts: dynamic shapes, control flow and custom ops
- Quantisation in production on at least one real model
- Comfort with at least two of ONNX Runtime, TensorRT, CoreML, OpenVINO, QNN and LiteRT
- Profiling fluency on at least one platform
Nice to have
- Custom op authoring in any runtime
Pay and location
Sarvam’s posting does not state a salary. The role is based in Bengaluru.
How to apply
Apply through the official Sarvam AI job posting. Sarvam AI’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.
See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.
Hiring institution: Sarvam AI
Official advertisement: jobs.ashbyhq.com
How to prepare for this application
- Export pitfalls: prepare a story about an ONNX export that broke and how you fixed it.
- Quantisation: compare post-training quantisation with quantisation-aware training, and how you check accuracy.
- Runtimes: know the trade-offs between two of the listed runtimes you have used.
- Profiling: be ready to read a profile and find the bottleneck operator.
- Documentation: the role writes deployment workbooks; bring a sample of clear technical writing.
About Sarvam
Sarvam is building what it calls the bedrock of sovereign AI for India: a full-stack platform spanning research, foundation models, infrastructure and applications, with a focus on making AI work for India's languages and institutions. Headquartered in Bengaluru, it works with leading enterprises and public institutions, is backed by Lightspeed, Peak XV and Khosla Ventures, and partners with brands such as Tata Capital, SBI Life, CRED, IDFC and LIC.