Performance Engineer, LLM Inference at Sarvam AI (Bengaluru)

Application ends: October 12, 2026
Apply Now

Job Description

Sarvam is hiring a senior Performance Engineer to own its production serving path for large distributed models: disaggregated prefill-decode, distributed KV-cache transfer, routing and scheduling, and speculative decoding with draft models the engineer trains. The posting lists Bengaluru or Chennai, hybrid or on-site.

At a glance

  • Role: Performance Engineer, Inference (senior)
  • Where: Bengaluru or Chennai, India (hybrid or on-site, per the posting)
  • Experience: 5+ years in ML systems, with 2+ years on production inference serving
  • Stack: SGLang, vLLM, NVIDIA Dynamo or TensorRT-LLM
  • Apply: open until filled (re-checked 12 October 2026)

About the inference role

  • Own the end-to-end serving path for large distributed models across a multi-node, multi-tenant fleet
  • Modify serving frameworks at source level where stock behaviour does not fit Sarvam’s workloads
  • Run and extend disaggregated prefill-decode, distributed KV-cache transfer and cross-node routing and scheduling
  • Build and train speculators (draft models distilled from target models) and tune acceptance rates against live traffic
  • Integrate artefacts from the model and kernel teams into the running stack

The two performance postings form a vertical stack: the kernels team writes microsecond-level GPU code, and the inference team integrates it into a running serving stack and owns system-level numbers. Sarvam serves several model families, including LLMs, mixture-of-experts models, Indic speech and multimodal models, on a multi-node fleet of NVIDIA Hopper and Blackwell GPUs.

What Sarvam is looking for

  • 5+ years in ML systems with 2+ years on production inference, with concrete outcomes such as throughput or p99 gains
  • Serving 100B+ parameter models across multi-node tensor, pipeline or expert parallelism
  • Source-level fluency in one of SGLang, vLLM, Dynamo or TensorRT-LLM (having modified its scheduler, KV allocator or disaggregation path)
  • Trained your own speculative decoding draft models and tuned them against a real serving distribution
  • Deep KV-cache knowledge, TP/PP/EP with NCCL, multi-tenant serving (MIG/MPS) and profiling with Nsight Systems and py-spy

Pay and location

Sarvam’s posting does not state a salary. It lists the location as Bengaluru or Chennai, hybrid or on-site.

How to apply

Apply through the official Sarvam job posting. Sarvam’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.

See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.

Hiring institution: Sarvam AI

Official advertisement: jobs.ashbyhq.com

How to prepare for this application

  • Serving numbers: prepare outcomes in tokens per second, cost per token or p99 latency from systems you ran.
  • Disaggregation: explain when splitting prefill and decode pays off and what KV transfer costs.
  • Speculative decoding: know EAGLE-style drafts and how acceptance rate interacts with batch size.
  • Framework internals: be ready to discuss the scheduler and KV allocator of the framework you know best.
  • Multi-tenancy: think through co-locating models with MIG or MPS without hurting latency.

About Sarvam

Sarvam is building what it calls the bedrock of sovereign AI for India: a full-stack platform spanning research, foundation models, infrastructure and applications, with a focus on making AI work for India's languages and institutions. Headquartered in Bengaluru, it works with leading enterprises and public institutions, is backed by Lightspeed, Peak XV and Khosla Ventures, and partners with brands such as Tata Capital, SBI Life, CRED, IDFC and LIC.

We send one confirmation email first. Every alert has an unsubscribe link.