Research Engineer, Benchmarks

Clera Singapore, Singapore Publicerat 4 september 2026
full_timeonsitemid
ABOUT THE ROLE You will own the design and implementation of rigorous, domain-specific benchmarks used to evaluate frontier AI agents on realistic workflows. Sitting within a small, highly technical team of researchers and engineers, this role is central to delivering evaluations that AI labs and enterprise customers genuinely trust and rely on. WHAT YOU'LL DO - Design, implement, and maintain high-quality internal benchmarks for evaluating frontier agents on domain-specific tasks. - Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks. - Build and operate reliable infrastructure to run models and agents against benchmark tasks at scale. - Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes. - Validate that benchmark performance correlates with real-world evaluations and customer needs. - Write clear technical documentation and benchmark reports for research and engineering audiences. WHAT WE'RE LOOKING FOR - 2 to 4 years of experience in research engineering or ML engineering, with a focus on AI benchmarks, evaluation infrastructure, or agent environments. - Strong proficiency in Python, Docker, and Linux for building research or production infrastructure. - Demonstrated experience designing and running benchmarks or evaluation environments for AI agents or large language models. - Experience building infrastructure to reliably run AI models or agents against evaluation tasks at scale. - Experience developing metrics or validation studies to assess benchmark difficulty, reliability, and real-world correlation. - Ability to collaborate with domain experts and translate complex workflows into evaluation criteria. - Strong attention to detail, with a habit of spotting subtle inconsistencies and edge cases. - Comfort working independently in fast-paced, early-stage startup environments with unstructured problem spaces. - Excellent written communication skills for technical documentation and cross-timezone collaboration. - Published papers or technical writing on AI benchmarking, model evaluation, or failure modes is a strong plus. - Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation is a plus. COMPENSATION & BENEFITS Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available. LOCATION On-site in Singapore.

Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.

Visa jobbet och ansök

Ursprunglig annons: jobs.ashbyhq.com