Research Engineer, Frontier Evals and Environments at OpenAI (San Francisco)

September 14, 2026
US$295,000 - US$380,000 / year
Application ends: October 14, 2026
Apply Now

Job Description

OpenAI’s Agent Post-Training team is hiring a research engineer for frontier evaluations and environments: building the ‘north star’ tasks and environments that steer its biggest training runs. Earlier open evaluations from this line of work include GDPval, SWE-bench Verified, MLE-bench, PaperBench and SWE-Lancer. It suits someone who wants to see model progress first-hand and shape where it goes.

At a glance

  • Team: Agent Post-Training, Frontier Evals and Environments
  • Location: San Francisco
  • Pay: $295,000 – $380,000 a year plus equity, as listed
  • Apply: open until filled (re-checked 14 October 2026)

What you would do

  • Build environments and evaluations that guide OpenAI’s most ambitious training runs
  • Decide with research, product, infrastructure and safety teams what goes into major model runs
  • Create the data, graders and feedback loops that teach agents new abilities
  • Track and interpret how quickly model capabilities improve

Why this role matters

Evaluations decide what a lab optimises for: a good benchmark shows where models fall short and steers the next training run. OpenAI has released several widely used evaluations from this line of work, so the role combines research judgement about what matters with engineering of environments, graders and data at very large scale.

What OpenAI is looking for

  • Strong research engineering skills in machine learning
  • Experience building benchmarks, environments or evaluation pipelines
  • Good judgement about what measures real progress

Nice to have

  • Experience with agents, coding models or tool use

Pay and location

OpenAI lists $295,000 to $380,000 a year plus equity for this San Francisco role.

How to apply

Apply through the official OpenAI job posting. OpenAI’s posting does not give a closing date, so the role is open until filled; ResearchJobs.in will re-check this listing on 14 October 2026. Check the posting for location and work-authorisation details before you apply.

See all our artificial intelligence jobs, or browse more research jobs on ResearchJobs.in.

Hiring institution: OpenAI

Official advertisement: jobs.ashbyhq.com

How to prepare for this application

  • Show evals you built: benchmarks or environments, and what they revealed.
  • Know the named benchmarks: be ready to discuss SWE-bench, MLE-bench or PaperBench.
  • Graders: experience with model-based or programmatic grading is relevant.
  • Agents: tool-use or computer-use work fits the team's focus.

About OpenAI

OpenAI is an AI research and deployment company behind the GPT and o-series models, ChatGPT, Codex and the OpenAI API. Its research teams work on training, reasoning, alignment, safety and evaluation of frontier models, mostly from San Francisco.

We send one confirmation email first. Every alert has an unsubscribe link.