Job Description
Together AI is hiring an AI Infrastructure Systems Engineer in Bangalore to build the automation that runs one of the world’s largest GPU fleets for frontier model training and inference, including AI agents that deploy, diagnose and repair infrastructure with minimal human intervention.
At a glance
- Role: AI Infrastructure System Engineer
- Where: Bangalore, India
- Experience: 3+ years building distributed systems, infrastructure platforms or large-scale backend software
- Languages: Python, Go or Rust
- Apply: open until filled (re-checked 12 October 2026)
About the AI infrastructure role
- Build fleet automation that provisions, validates, deploys, upgrades, repairs and retires GPU clusters
- Build AI infrastructure agents that automate deployment, root-cause failures, triage incidents and remediate problems
- Develop fleet intelligence platforms that monitor hardware health, firmware, networking, storage, thermals and workload performance to predict failures
- Create automated validation for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage and distributed AI workloads
- Build internal platforms and developer tools, working with hardware, networking, platform and AI teams
Running tens of thousands of GPUs reliably is a software problem as much as a hardware one: failures are constant at that scale, and every idle GPU costs money. Together AI’s India roles build the automation and AI agents that keep that fleet healthy and busy.
What Together AI is looking for
- 3+ years building distributed systems, infrastructure platforms or large-scale backend software
- Strong software engineering in Python, Go or Rust
- Experience building platforms, automation systems or developer infrastructure
- Linux, Kubernetes, Terraform, Ansible or similar
- Systems thinking across hardware and software, and an automation-first mindset
Nice to have
- GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch
- InfiniBand or RoCE networking
- Bare-metal provisioning and lifecycle management
- Large-scale AI training or inference clusters, hardware health monitoring, distributed storage, and AI agents for operations
Pay and location
Together AI’s posting does not state a salary. The role is based in Bangalore.
How to apply
Apply through the official Together AI job posting. Together AI’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.
See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.
Hiring institution: Together AI
Official advertisement: job-boards.greenhouse.io
How to prepare for this application
- Automation stories: bring an example of replacing a manual operations task with software.
- GPU failure modes: read about common GPU, NVLink and InfiniBand faults and how they are detected.
- Kubernetes: be ready to discuss operators and controllers for bare-metal or GPU workloads.
- Agents for ops: think about how an AI agent could triage an incident safely, with humans in the loop.
- Coding: expect systems-design and coding interviews in Python, Go or Rust.
About Together AI
Together AI describes itself as the AI Native Cloud, purpose-built for AI engineers. It offers high-performance inference, fine-tuning and reinforcement learning, and large-scale pre-training around a marketplace of open models. Customers named in its posting include Cursor, Decagon, ElevenLabs, Salesforce and Zoom, and the company says it serves more than 400 trillion tokens a month.