Senior Research Scientist, Model Evaluation at Cohere (Toronto or Remote)

September 11, 2026
250000 - 535000 / year
Application ends: October 12, 2026
Apply Now

Job Description

Cohere is hiring a Senior Research Scientist in model evaluation to create the next generation of methods and infrastructure for measuring large language model progress, including new benchmarks and trained LLM judges. The role can be based in Toronto or remotely, with no minimum in-office requirement.

At a glance

  • Role: Senior Research Scientist, Model Evaluation
  • Where: Toronto, Canada, or remote (no minimum in-office requirement)
  • Pay: CA$250,000–535,000, with equity (multiple ranges listed)
  • Focus: LLM evaluation benchmarks, LLM judges and evaluation infrastructure
  • Apply: open until filled (re-checked 12 October 2026)

About the evaluation research role

  • Create ambitious new evaluation benchmarks that push the limits of what Cohere’s models can do
  • Work with cross-functional teams to turn model feedback into trustworthy, repeatable evaluations
  • Research better LLM evaluation methods, including training LLM judges, refining LLM-based data synthesis pipelines and improving evaluation efficiency
  • Build scalable, reusable tools for digging into model performance

As models become better than humans at many tasks, measuring them properly gets harder. Evaluation research designs benchmarks and judges that track what models can really do, and sets targets for what the next models should be able to do.

What Cohere is looking for

  • Enjoyment of rapidly building prototypes that show the boundaries of LLM capabilities, and experience building resources to measure them
  • Dozens of hours spent reviewing complex data and LLM outputs to ensure quality
  • Rigour about measuring AI capabilities and making sure measurements align with the capabilities that matter
  • Strong software engineering skills

Pay and location

Cohere lists pay of CA$250,000–535,000 with equity, across multiple ranges. Its listed perks include a weekly lunch stipend, health and dental benefits, six weeks of paid vacation, parental leave top-up and learning stipends.

How to apply

Apply through the official Cohere job posting. Cohere’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 12 October 2026. Check the posting for location and work-authorisation details before you apply.

See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.

Hiring institution: Cohere

Official advertisement: jobs.ashbyhq.com

How to prepare for this application

  • Benchmark design: propose a benchmark for a capability current evals miss, and explain how you would stop it saturating.
  • LLM judges: know how judges are trained and validated, and their failure modes.
  • Data quality: bring an example where manual review of outputs changed your conclusions.
  • Validity: be ready to discuss contamination, variance and whether a metric tracks the intended skill.
  • Tooling: show evaluation code or dashboards you built that others used.

About Cohere

Cohere describes itself as a security-first enterprise AI company that trains and deploys frontier foundation models and end-to-end products for businesses. It is headquartered in Toronto, with key offices in London, New York, San Francisco, Montreal, Paris, Berlin and Seoul.

We send one confirmation email first. Every alert has an unsubscribe link.