Embedded Data Scientist, Chanakya at Sarvam AI (Delhi)

Application ends: October 12, 2026
Apply Now

Job Description

Sarvam AI is hiring an Embedded Data Scientist for its Chanakya team in Delhi. Working at client sites alongside Strategic Deployment Engineers, you will turn large, messy, multimodal client data into structures that AI systems can reason over reliably.

At a glance

  • Role: Embedded Data Scientist, Chanakya
  • Location: Delhi
  • Experience: 2–5 years
  • Data: documents, imagery, audio, geospatial data and structured records
  • Apply: open until filled (re-checked 12 October 2026)

What you will do

  • Understand the client’s data sources, formats, workflows and terminology across documents, imagery, audio, geospatial data and records
  • Design domain ontologies for entities, relationships, hierarchies and operational concepts
  • Define document segmentation and chunking strategies that keep meaning intact and support retrieval
  • Decide how each data type should be indexed, embedded and linked, and help deployment engineers turn this into ingestion pipelines
  • Evaluate how well the AI system retrieves and reasons over client data, and refine the structures
  • Help define benchmarks that reflect real deployments, and feed insights back to product and engineering teams

The posting explains that this role decides how data is represented inside the AI system: how documents are split, what metadata exists, how entities and relationships are modelled, and how different data types connect. The work often involves classified or operationally sensitive datasets in places where standard tooling may not exist, and you own the quality of the data layer in your accounts.

What Sarvam AI is looking for

  • 2–5 years in data science, applied machine learning or large-scale data analysis
  • Strong Python, including pandas, NumPy and modern NLP or LLM tooling
  • Solid machine learning fundamentals, enough to shape evaluation and work with a models team on training and benchmarks
  • Experience with large unstructured datasets such as documents, transcripts, reports or operational records
  • Familiarity with LLM systems, retrieval pipelines or vector search
  • Experience designing data schemas, metadata frameworks, entity models or semantic structures

Nice to have

  • Knowledge graphs, ontologies or semantic data modelling
  • Multimodal datasets (text, imagery, audio, geospatial or structured)
  • Work in constrained or air-gapped environments

Pay and location

Sarvam’s posting does not state a salary. The role is based in Delhi, with work at client sites.

How to apply

Apply through the official Sarvam AI job posting. Sarvam AI’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.

See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.

Hiring institution: Sarvam AI

Official advertisement: jobs.ashbyhq.com

How to prepare for this application

  • Chunking: be ready to argue for a segmentation strategy for long, structured reports.
  • Ontology design: practise modelling entities and relationships for an unfamiliar domain.
  • Retrieval evaluation: know recall@k and how to build a labelled test set.
  • Messy data: bring an example of structure you built from chaotic data.
  • Communication: explain a data insight to a non-technical stakeholder.

About Sarvam

Sarvam is building what it calls the bedrock of sovereign AI for India: a full-stack platform spanning research, foundation models, infrastructure and applications, with a focus on making AI work for India's languages and institutions. Headquartered in Bengaluru, it works with leading enterprises and public institutions, is backed by Lightspeed, Peak XV and Khosla Ventures, and partners with brands such as Tata Capital, SBI Life, CRED, IDFC and LIC.

We send one confirmation email first. Every alert has an unsubscribe link.