Job Description
Perplexity is hiring an AI Inference Engineer in London to build and run the inference engine behind every Perplexity query, serving dozens of model architectures with tight latency and cost budgets. The stack is Rust, Python, CUDA and NVIDIA’s CuTe DSL, and the work runs from new model support and GPU kernels to a Rust serving runtime and production reliability.
At a glance
- Team: Inference
- Location: London, United Kingdom
- Stack: Rust, Python, CUDA, CuTe DSL
- Apply: open until filled (re-checked 14 October 2026)
What you would do
- Add support for new transformer-based retrieval, text-generation and multimodal models, from weight loading to scheduling and KV-cache management
- Port in-house CUDA kernels to CuTe DSL for current and next-generation NVIDIA racks
- Develop the Rust-based inference server to handle fast-growing traffic
- Profile and remove bottlenecks from network ingress to continuous batching and kernels
- Build dashboards, alerts and automated remediation, and learn from incidents
Why this role matters
Every answer Perplexity gives passes through its inference engine, so its speed and reliability shape the product. The team is moving kernels to NVIDIA’s CuTe DSL and writing its serving runtime in Rust, which makes the role a chance to work on modern GPU and systems programming in a fast-growing London team.
What Perplexity is looking for
- Deep experience with GPU programming and performance (CUDA, Triton, CUTLASS or similar)
- Understanding of modern LLM architectures and bringing them up in production
- Experience building and running production distributed systems under real load
- Comfort across Rust, Python and CUDA
Nice to have
- Other deep systems programming experience
Pay and location
Perplexity does not show a London pay range in this listing; check the posting for details.
How to apply
Apply through the official Perplexity job posting. Perplexity’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 14 October 2026. Check the posting for location and work-authorisation details before you apply.
See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.
Hiring institution: Perplexity
Official advertisement: jobs.ashbyhq.com
How to prepare for this application
- Kernels first: show CUDA, Triton or CUTLASS work with measured speed-ups.
- Serving systems: batching, KV-cache or scheduling work is directly relevant.
- Rust: any production Rust experience is a plus.
- Incidents: examples of running systems in production matter.
About Perplexity
Perplexity builds AI-powered search and agent products, including its Sonar models, Deep Research and the Comet browser agent, serving hundreds of millions of queries. Its research and inference teams train and serve many model architectures at scale.