Performance Engineer, GPU Kernels at Sarvam AI (Bengaluru)

Application ends: October 12, 2026
Apply Now

Job Description

Sarvam is hiring a senior Performance Engineer to own its GPU kernel layer, writing custom CUDA, DSL-based and PTX kernels where stock libraries such as cuBLAS, FlashAttention or out-of-the-box Triton leave performance on the table. The posting lists Bengaluru or Chennai, hybrid or on-site.

At a glance

  • Role: Performance Engineer, Kernels (senior)
  • Where: Bengaluru or Chennai, India (hybrid or on-site, per the posting)
  • Experience: 5+ years in ML systems, with 2+ years writing production CUDA kernels
  • Hardware: H100, H200 and B200 GPUs
  • Apply: open until filled (re-checked 12 October 2026)

About the kernels role

  • Author custom CUDA, CUTLASS/CuTe DSL and PTX kernels that beat stock libraries on real production workloads
  • Own the explanation of every performance change your kernels make to production latency
  • Work on attention kernels and other hot paths across the serving stack
  • Collaborate with the inference team, which integrates kernels into the serving system

The two performance postings form a vertical stack: the kernels team writes microsecond-level GPU code, and the inference team integrates it into a running serving stack and owns system-level numbers. Sarvam serves several model families, including LLMs, mixture-of-experts models, Indic speech and multimodal models, on a multi-node fleet of NVIDIA Hopper and Blackwell GPUs.

What Sarvam is looking for

  • 5+ years in ML systems, including 2+ years authoring production CUDA kernels, with a kernel in production that beat its baseline
  • CUDA at kernel-authoring level: block sizing, shared-memory layout, warp primitives, async copies (cp.async, TMA) and MMA selection
  • CUTLASS/CuTe DSL at modify-and-extend level, and PTX at debug-and-modify level
  • Nsight Compute and Systems fluency, including reading roofline plots
  • Hands-on attention kernel work (FlashAttention-family, paged, MLA, sliding-window or sparse) and awareness of Hopper-to-Blackwell changes

Nice to have

  • Communication kernels: NCCL/NVSHMEM, custom collectives, expert-parallel dispatch or KV transfer
  • Open-source kernel contributions, which the posting calls the highest-yield signal
  • tcgen05, TMA, CTA clusters, distributed shared memory, async pipelining and FP4/microscaling paths
  • Host-path optimisation on GH200/GB200

Pay and location

Sarvam’s posting does not state a salary. It lists the location as Bengaluru or Chennai, hybrid or on-site.

How to apply

Apply through the official Sarvam job posting. Sarvam’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.

See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.

Hiring institution: Sarvam AI

Official advertisement: jobs.ashbyhq.com

How to prepare for this application

  • Show a kernel: bring a kernel you wrote, the baseline it beat, and the profiler evidence.
  • Attention internals: be ready to whiteboard tiling and softmax rescaling in FlashAttention.
  • Hopper vs Blackwell: know what TMA, WGMMA and tcgen05 change.
  • Roofline reasoning: practise deciding whether a kernel is compute- or memory-bound and what to do next.
  • GitHub: tidy your public kernel work; the posting says it is the strongest signal.

About Sarvam

Sarvam is building what it calls the bedrock of sovereign AI for India: a full-stack platform spanning research, foundation models, infrastructure and applications, with a focus on making AI work for India's languages and institutions. Headquartered in Bengaluru, it works with leading enterprises and public institutions, is backed by Lightspeed, Peak XV and Khosla Ventures, and partners with brands such as Tata Capital, SBI Life, CRED, IDFC and LIC.

We send one confirmation email first. Every alert has an unsubscribe link.