Senior Software Engineer, Infrastructure Agent Systems at Together AI (Remote, India)

September 11, 2026
Application ends: October 12, 2026
Apply Now

Job Description

Together AI is hiring a Senior Software Engineer for Infrastructure Agent Systems, remote in India, to build production AI agents that diagnose hardware failures and investigate incidents across its GPU fleet, along with the knowledge graphs, retrieval and orchestration platform those agents run on.

At a glance

  • Role: Senior Software Engineer, Infra Agent Systems
  • Where: Remote, based in India
  • Experience: 5+ years building production backend, distributed or infrastructure systems
  • Focus: AI agents, knowledge graphs and retrieval for autonomous infrastructure
  • Apply: open until filled (re-checked 12 October 2026)

About the infrastructure agents role

  • Design and build production AI agents that diagnose, investigate and remediate infrastructure issues across a very large GPU fleet
  • Build the distributed services, orchestration framework, knowledge graph and retrieval systems that power the agents
  • Develop fleet intelligence that combines telemetry, infrastructure state, operational knowledge and past incidents
  • Integrate with observability, incident management, ticketing, inventory, source control and chat systems
  • Deliver, operate and support the software in production

Running tens of thousands of GPUs reliably is a software problem as much as a hardware one: failures are constant at that scale, and every idle GPU costs money. Together AI’s India roles build the automation and AI agents that keep that fleet healthy and busy.

What Together AI is looking for

  • 5+ years building production backend, distributed or infrastructure systems, owning major systems from design to production
  • Depth in at least one of: AI agent systems (orchestration, tool use, evaluation, grounding), knowledge graphs, or search and retrieval (ranking, RAG, semantic search)
  • Strong backend engineering: API design, service boundaries, data modelling and integrations
  • Kubernetes, GitOps such as ArgoCD, infrastructure-as-code and cloud platforms
  • Comfort in Go, TypeScript, Python or Rust

Nice to have

  • GPU infrastructure, datacentres, bare-metal systems, hardware failure modes, BMC/IPMI or cluster schedulers
  • Graph databases and event-driven systems such as NATS or Kafka
  • Observability with Prometheus and Grafana
  • Evaluation frameworks for LLM-powered systems

Pay and location

Together AI’s posting does not state a salary. The role is remote, based in India.

How to apply

Apply through the official Together AI job posting. Together AI’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.

See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.

Hiring institution: Together AI

Official advertisement: job-boards.greenhouse.io

How to prepare for this application

  • Trustworthy agents: be ready to explain how you would stop an agent taking a harmful action on production systems.
  • Knowledge graphs: design a graph linking hosts, components, incidents and runbooks.
  • Retrieval: discuss grounding an agent in logs and past incidents with RAG.
  • Evaluation: how would you measure whether an agent's diagnosis was right?
  • Ownership: the team runs what it builds, so bring an on-call or production-support story.

About Together AI

Together AI describes itself as the AI Native Cloud, purpose-built for AI engineers. It offers high-performance inference, fine-tuning and reinforcement learning, and large-scale pre-training around a marketplace of open models. Customers named in its posting include Cursor, Decagon, ElevenLabs, Salesforce and Zoom, and the company says it serves more than 400 trillion tokens a month.

We send one confirmation email first. Every alert has an unsubscribe link.