Staff Engineer, Distributed Storage and HPC & AI Infrastructure at Together AI (Bangalore)

Application ends: October 12, 2026
Apply Now

Job Description

Together AI is hiring a Staff Engineer for distributed storage and HPC/AI infrastructure in Bangalore to run, scale and tune the multi-petabyte storage systems behind its GPU training and inference workloads, and to set the company’s storage roadmap as the fleet grows.

At a glance

  • Role: Staff Engineer, Distributed Storage and HPC & AI Infrastructure
  • Location: Bangalore
  • Experience: 8+ years in storage engineering
  • Technologies: Vast, Weka, Ceph, Lustre; Kubernetes; Go and Python
  • Apply: open until filled (re-checked 12 October 2026)

What you will do

  • Set the technical strategy and storage roadmap as Together’s GPU fleet scales
  • Engineer multi-petabyte storage with Vast, Weka and Ceph, cutting cost through automated tiering and lifecycle policies
  • Design caching and tiered storage for very high IOPS and cluster-wide throughput
  • Tune storage isolation at network layers 2 and 3 for secure multi-tenancy
  • Write Kubernetes storage operators and controllers for self-service provisioning and quota enforcement
  • Profile and benchmark data paths, contributing to open-source storage projects and internal tools

AI training and inference need data delivered fast: model weights, checkpoints and datasets must reach GPUs quickly enough to keep them busy. The posting talks about data paths of 10 GB/s or more per GPU node, multi-tier caching, intelligent prefetching and model-weight distribution across thousands of nodes.

What Together AI is looking for

  • 8+ years in storage engineering at multi-petabyte scale, including high-performance storage for GPU or HPC clusters
  • Deep expertise in a parallel filesystem such as Ceph, WekaFS, Lustre, Vast or GPFS, plus object storage such as S3, MinIO, Ceph or R2
  • Kubernetes storage: CSI drivers, StatefulSets, PersistentVolumes and custom controllers
  • Strong Go and Python for production systems and tooling
  • RDMA/InfiniBand and parallel-filesystem tuning for GPU workloads, and the Linux storage stack (filesystems, LVM, NVMe, RAID)
  • Infrastructure as code (Terraform, Ansible, Helm, ArgoCD) and observability (Prometheus, Grafana, Thanos)
  • A BS or MS in computer science or engineering, or equivalent experience, and a history of technical leadership

Nice to have

  • GPU Direct Storage, NVMe-oF and RDMA implementations
  • ML storage patterns such as model weights, checkpointing and dataset caching
  • Benchmarking tools such as fio, iperf3, iostat and blktrace

Pay and location

Together AI’s posting does not state a salary. The role is based in Bangalore.

How to apply

Apply through the official Together AI job posting. Together AI’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.

See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.

Hiring institution: Together AI

Official advertisement: job-boards.greenhouse.io

How to prepare for this application

  • Parallel filesystems: be ready to compare Lustre, Weka and Ceph for checkpoint-heavy training.
  • Throughput maths: estimate the bandwidth needed to load a large model onto many nodes at once.
  • Kubernetes storage: explain how a CSI driver and a custom operator fit together.
  • Leadership: prepare an example where your design changed reliability or cost significantly.
  • Benchmarks: know how you would use fio to find a storage bottleneck.

About Together AI

Together AI describes itself as the AI Native Cloud, purpose-built for AI engineers. It offers high-performance inference, fine-tuning and reinforcement learning, and large-scale pre-training around a marketplace of open models. Customers named in its posting include Cursor, Decagon, ElevenLabs, Salesforce and Zoom, and the company says it serves more than 400 trillion tokens a month.

We send one confirmation email first. Every alert has an unsubscribe link.