Job Description
Cohere’s Model Efficiency team is hiring a Staff Research Engineer to push the limits of LLM inference efficiency, which the ad calls the main bottleneck. The team works across the model execution stack: model architecture and mixture-of-experts routing, decoding algorithms, software–hardware co-design for GPUs and performance tuning without hurting quality. The team is concentrated in the Eastern and Pacific time zones.
At a glance
- Team: Model Efficiency
- Location: New York preferred (EST/PST), other Cohere offices possible
- Focus: LLM inference efficiency
- Apply: open until filled (re-checked 14 October 2026)
What you would do
- Develop and prototype efficiency improvements across model architecture and MoE routing
- Improve decoding and inference-time algorithms
- Co-design software and hardware for GPU acceleration
- Ship performance gains without reducing model quality
Why this role matters
Serving large language models is expensive, and every gain in inference efficiency lowers cost and latency for customers. Cohere’s efficiency team works on the whole stack, from model design to GPU kernels, so a staff engineer here sees the full picture of how modern LLMs run in production. The team sits mainly in the Eastern and Pacific time zones, which the ad names as preferred locations.
What Cohere is looking for
- Deep experience in ML systems or model efficiency research
- Strong GPU and performance engineering skills
- A record of shipping improvements in production models
Nice to have
- Experience with speculative decoding, quantisation or MoE systems
Pay and location
Cohere lists pay ranges on the posting that vary by location; check the posting for the New York range.
How to apply
Apply through the official Cohere job posting. Cohere’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 14 October 2026. Check the posting for location and work-authorisation details before you apply.
See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.
Hiring institution: Cohere
Official advertisement: jobs.ashbyhq.com
How to prepare for this application
- Show speed-ups: quote latency or throughput gains you delivered.
- Name the techniques: MoE routing, quantisation, speculative decoding or kernels.
- Quality: explain how you checked that efficiency changes kept quality.
- Staff level: show technical leadership across teams.
About Cohere
Cohere builds foundation language models and AI products for businesses, with a focus on security and private deployment. It is headquartered in Toronto, with offices in London, New York, San Francisco, Montreal, Paris, Berlin and Seoul, and several teams work remotely within set time zones.