Member of Technical Staff, AI Inference Engineer at Perplexity (London)

Application ends: October 14, 2026
Apply Now

Job Description

Perplexity is hiring an AI Inference Engineer in London to build and run the inference engine behind every Perplexity query, serving dozens of model architectures with tight latency and cost budgets. The stack is Rust, Python, CUDA and NVIDIA’s CuTe DSL, and the work runs from new model support and GPU kernels to a Rust serving runtime and production reliability.

At a glance

  • Team: Inference
  • Location: London, United Kingdom
  • Stack: Rust, Python, CUDA, CuTe DSL
  • Apply: open until filled (re-checked 14 October 2026)

What you would do

  • Add support for new transformer-based retrieval, text-generation and multimodal models, from weight loading to scheduling and KV-cache management
  • Port in-house CUDA kernels to CuTe DSL for current and next-generation NVIDIA racks
  • Develop the Rust-based inference server to handle fast-growing traffic
  • Profile and remove bottlenecks from network ingress to continuous batching and kernels
  • Build dashboards, alerts and automated remediation, and learn from incidents

Why this role matters

Every answer Perplexity gives passes through its inference engine, so its speed and reliability shape the product. The team is moving kernels to NVIDIA’s CuTe DSL and writing its serving runtime in Rust, which makes the role a chance to work on modern GPU and systems programming in a fast-growing London team.

What Perplexity is looking for

  • Deep experience with GPU programming and performance (CUDA, Triton, CUTLASS or similar)
  • Understanding of modern LLM architectures and bringing them up in production
  • Experience building and running production distributed systems under real load
  • Comfort across Rust, Python and CUDA

Nice to have

  • Other deep systems programming experience

Pay and location

Perplexity does not show a London pay range in this listing; check the posting for details.

How to apply

Apply through the official Perplexity job posting. Perplexity’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 14 October 2026. Check the posting for location and work-authorisation details before you apply.

See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.

Hiring institution: Perplexity

Official advertisement: jobs.ashbyhq.com

How to prepare for this application

  • Kernels first: show CUDA, Triton or CUTLASS work with measured speed-ups.
  • Serving systems: batching, KV-cache or scheduling work is directly relevant.
  • Rust: any production Rust experience is a plus.
  • Incidents: examples of running systems in production matter.

About Perplexity

Perplexity builds AI-powered search and agent products, including its Sonar models, Deep Research and the Comet browser agent, serving hundreds of millions of queries. Its research and inference teams train and serve many model architectures at scale.

We send one confirmation email first. Every alert has an unsubscribe link.