Mercor is seeking PhD/Master's scientists to author AI evaluation tasks and build benchmarks for scientific computing. You will craft original, executable research problems that current frontier models struggle with, sourcing material from papers, datasets, or open-source repositories. Collaboration with AI labs is expected. Responsibilities include writing prompts, designing grading criteria, and calibrating models to ensure robust evaluation.

Also on the board Same function, level within a rung

Level

Senior

Location

San Diego, CA

Occupation

Chemists

Industry

Research and Development in the Physical, Engineering, and Life Sciences (except Nanotechnology and Biotechnology)

Posted

today

Apply for this role →