Job Description
Together AI is hiring a Staff Engineer for distributed storage and HPC/AI infrastructure in Bangalore to run, scale and tune the multi-petabyte storage systems behind its GPU training and inference workloads, and to set the company’s storage roadmap as the fleet grows.
At a glance
- Role: Staff Engineer, Distributed Storage and HPC & AI Infrastructure
- Location: Bangalore
- Experience: 8+ years in storage engineering
- Technologies: Vast, Weka, Ceph, Lustre; Kubernetes; Go and Python
- Apply: open until filled (re-checked 12 October 2026)
What you will do
- Set the technical strategy and storage roadmap as Together’s GPU fleet scales
- Engineer multi-petabyte storage with Vast, Weka and Ceph, cutting cost through automated tiering and lifecycle policies
- Design caching and tiered storage for very high IOPS and cluster-wide throughput
- Tune storage isolation at network layers 2 and 3 for secure multi-tenancy
- Write Kubernetes storage operators and controllers for self-service provisioning and quota enforcement
- Profile and benchmark data paths, contributing to open-source storage projects and internal tools
AI training and inference need data delivered fast: model weights, checkpoints and datasets must reach GPUs quickly enough to keep them busy. The posting talks about data paths of 10 GB/s or more per GPU node, multi-tier caching, intelligent prefetching and model-weight distribution across thousands of nodes.
What Together AI is looking for
- 8+ years in storage engineering at multi-petabyte scale, including high-performance storage for GPU or HPC clusters
- Deep expertise in a parallel filesystem such as Ceph, WekaFS, Lustre, Vast or GPFS, plus object storage such as S3, MinIO, Ceph or R2
- Kubernetes storage: CSI drivers, StatefulSets, PersistentVolumes and custom controllers
- Strong Go and Python for production systems and tooling
- RDMA/InfiniBand and parallel-filesystem tuning for GPU workloads, and the Linux storage stack (filesystems, LVM, NVMe, RAID)
- Infrastructure as code (Terraform, Ansible, Helm, ArgoCD) and observability (Prometheus, Grafana, Thanos)
- A BS or MS in computer science or engineering, or equivalent experience, and a history of technical leadership
Nice to have
- GPU Direct Storage, NVMe-oF and RDMA implementations
- ML storage patterns such as model weights, checkpointing and dataset caching
- Benchmarking tools such as fio, iperf3, iostat and blktrace
Pay and location
Together AI’s posting does not state a salary. The role is based in Bangalore.
How to apply
Apply through the official Together AI job posting. Together AI’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.
See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.
Hiring institution: Together AI
Official advertisement: job-boards.greenhouse.io
How to prepare for this application
- Parallel filesystems: be ready to compare Lustre, Weka and Ceph for checkpoint-heavy training.
- Throughput maths: estimate the bandwidth needed to load a large model onto many nodes at once.
- Kubernetes storage: explain how a CSI driver and a custom operator fit together.
- Leadership: prepare an example where your design changed reliability or cost significantly.
- Benchmarks: know how you would use fio to find a storage bottleneck.
About Together AI
Together AI describes itself as the AI Native Cloud, purpose-built for AI engineers. It offers high-performance inference, fine-tuning and reinforcement learning, and large-scale pre-training around a marketplace of open models. Customers named in its posting include Cursor, Decagon, ElevenLabs, Salesforce and Zoom, and the company says it serves more than 400 trillion tokens a month.