Research Engineer, Synthetic Data
ABOUT THE ROLE
This is a hands-on research engineering role focused on building synthetic data pipelines that turn domain-specific workflows into scalable training tasks for AI agents. You will join a roughly 15-person engineering team of Olympiad medalists and published researchers, working at the frontier of reinforcement learning and AI alignment. The work you do here directly expands what AI models are capable of.
WHAT YOU'LL DO
- Design and build end-to-end synthetic data pipelines that convert domain-specific workflows into structured, realistic, and challenging training tasks.
- Collaborate with subject-matter experts to create synthetic tasks for AI agents across professional and technical domains.
- Develop synthetic task generation methods that produce diverse, realistic, and learnable outputs.
- Build tooling to mutate, validate, and iteratively improve synthetic tasks at scale.
- Analyze model and agent performance on synthetic tasks to understand what the tasks teach and where they break down.
- Define and implement metrics to quantify synthetic task diversity, realism, learnability, and overall quality.
WHAT WE'RE LOOKING FOR
- 2 to 4 years of experience in software engineering, machine learning engineering, or AI research, with a focus on data pipelines, ML infrastructure, or synthetic data systems.
- Proficiency in Python and hands-on experience in Linux environments using containerization tools such as Docker.
- Demonstrated experience applying synthetic data research methods to build end-to-end data generation pipelines for AI or ML applications.
- Strong understanding of synthetic data quality criteria and evaluation metrics, including diversity, realism, and learnability, along with awareness of their inherent limitations.
- Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
- Proven ability to independently own and deliver technical projects end-to-end with minimal predefined requirements or roadmap.
- Sharp eye for edge cases, subtle inconsistencies, and quality issues in synthetic or algorithmically generated datasets.
- Ability to reason from first principles about task design, scoring, and failure modes.
- Familiarity with reinforcement learning training paradigms, agentic AI workflows, or LLM post-training pipelines is a plus.
- Strong communication skills for effective remote collaboration across time zones.
COMPENSATION & BENEFITS
Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.
LOCATION
On-site in Singapore.
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: jobs.ashbyhq.com