Job Description
Together AI is hiring software engineers at junior, senior or staff level, remote in India, to build the Kubernetes-native control plane that provisions and runs its GPU inference fleet, so the inference team can request clusters, model deployments or capacity changes through an API and let controllers handle the rest.
At a glance
- Role: Software Engineer (junior, senior or staff), Inference / Compute Infrastructure
- Location: Remote in India
- Languages: Go, Python, Rust or similar
- Focus: provisioning state machine, self-service API and self-healing
- Apply: open until filled (re-checked 12 October 2026)
What you will do
- Build the provisioning state machine that tracks each physical host from discovery and GPU driver/CUDA bring-up to health checks and decommissioning
- Design declarative APIs so the inference team can create, scale and tear down clusters with one call and no tickets
- Automate self-healing: detect failing nodes, drain them safely, trigger repair and return healthy capacity to the pool
- Make the pipeline dependable with idempotency, retries, rollback and drift detection
- Encode the cluster shapes the inference team needs (topology, interconnect, scheduling constraints) as platform abstractions
- Engineer infrastructure code like a product, with strong typing, tests, review, versioning and CI/CD
The goal the posting describes is decoupling: people building on the platform should not need to know which serving stack, scheduler or hardware pool is doing the work. The same systems should keep the fleet efficient, with defragmentation, rebalancing and bin-packing that raise GPU utilisation without hurting latency. Engineers own what they build, including running it in production.
What Together AI is looking for
- Strong software engineering in Go, Python, Rust or similar
- Durable workflow orchestration such as Temporal or Cadence for long-running workflows that survive failures
- Building control planes or orchestration systems that reconcile state over time, such as Kubernetes controllers or operators
- Event-driven design with queues or streams such as Kafka, NATS or SQS
- A product mindset: internal platforms or APIs used by other engineering teams
Nice to have
- Bare-metal provisioning (PXE/iPXE, Redfish/IPMI, BMC) or networking basics (VLANs, BGP, fabric design)
- GPU cluster software (NCCL, CUDA, InfiniBand/RoCE)
- Work at a hyperscaler, GPU cloud or data-centre-scale organisation
- Systems programming in Rust or Go
Pay and location
Together AI’s posting does not state a salary. The role is remote within India.
How to apply
Apply through the official Together AI job posting. Together AI’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.
See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.
Hiring institution: Together AI
Official advertisement: job-boards.greenhouse.io
How to prepare for this application
- Reconciliation loops: explain how a controller converges actual state to desired state.
- Durable workflows: know how Temporal replays history to resume after a crash.
- Bin-packing: think about GPU fragmentation and how to consolidate workloads.
- Level: the posting hires at several levels; say clearly which you are applying for.
- Ownership: prepare a story about operating software you wrote in production.
About Together AI
Together AI describes itself as the AI Native Cloud, purpose-built for AI engineers. It offers high-performance inference, fine-tuning and reinforcement learning, and large-scale pre-training around a marketplace of open models. Customers named in its posting include Cursor, Decagon, ElevenLabs, Salesforce and Zoom, and the company says it serves more than 400 trillion tokens a month.