Job Description
Cohere’s integration team is hiring a research engineer to develop and scale machine-learning algorithms and infrastructure for LLM post-training, with a focus on large, distributed reinforcement learning. The role improves the post-training codebase, optimises algorithms and scales distributed RL. It is listed for Paris, and applicants can work anywhere between UTC−6 and UTC+1.
At a glance
- Team: Integration / RL (post-training)
- Location: Paris, London, Toronto, San Francisco or New York offices, or remote within UTC−6 to UTC+1
- Focus: Distributed RL for LLM post-training
- Apply: open until filled (re-checked 14 October 2026)
What you would do
- Design and write high-performance, scalable software for training models
- Build tools that make post-training research easier
- Optimise post-training algorithms and scale distributed RL
- Write design documents and careful experiments with the wider team
Why this role matters
Reinforcement learning after pre-training is now where much of a language model’s usefulness is shaped, and running it across thousands of GPUs is a hard systems problem. This role improves the shared code that Cohere’s researchers use every day, which multiplies the whole team’s output. The flexible time-zone rule makes it open to strong candidates across Europe and the Americas.
What Cohere is looking for
- Strong software engineering in Python and ML frameworks
- Experience with distributed training or large-scale ML systems
- Interest in reinforcement learning for language models
Nice to have
- Experience with RL post-training (RLHF, RL from verifiable rewards) at scale
Pay and location
Cohere lists pay ranges on the posting that vary by location; check the posting for the range that applies to you.
How to apply
Apply through the official Cohere job posting. Cohere’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 14 October 2026. Check the posting for location and work-authorisation details before you apply.
See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.
Hiring institution: Cohere
Official advertisement: jobs.ashbyhq.com
How to prepare for this application
- Distributed systems: show training jobs you scaled across many GPUs.
- RL for LLMs: post-training experience is the heart of this team.
- Code quality: the role is about improving a shared codebase.
- Time zones: you must be able to work within UTC−6 to UTC+1.
About Cohere
Cohere builds foundation language models and AI products for businesses, with a focus on security and private deployment. It is headquartered in Toronto, with offices in London, New York, San Francisco, Montreal, Paris, Berlin and Seoul, and several teams work remotely within set time zones.