Principal / Senior Research Scientist, Foundation Models for Credit
ABOUT THE ROLE
This is a founding research role at the center of a new effort to reimagine consumer credit scoring from the model architecture up. You will own the foundation models that underpin the entire research program, designing and pre-training purpose-built encoder architectures for tabular and event-sequence credit data. The work is high-stakes, technically deep, and shapes the direction of the team from day one.
WHAT YOU'LL DO
- Design, implement, and pre-train encoder-first transformer architectures for heterogeneous tabular and sequential credit-file data, spanning tokenization schemes through training objectives.
- Drive the research agenda on representation learning for credit, including self-supervised objectives, handling of missingness and censoring, temporal drift, and transfer to downstream tasks.
- Own the training infrastructure end-to-end: distributed training, experiment tracking, evaluation harnesses, and rigorous ablation discipline.
- Collaborate with causal and explainability researchers to build interpretability into architectures by construction, through attention structures, bottlenecks, and monotonicity constraints.
- Benchmark rigorously against gradient-boosting incumbents and published tabular foundation models, and communicate results to internal, regulatory, and external audiences.
WHAT WE'RE LOOKING FOR
- 8 to 10 or more years of experience, with at least 5 years specifically building and training transformer encoder architectures from scratch, including attention mechanisms, positional and temporal encodings, and tokenization of non-text data at scale.
- Demonstrated experience pre-training foundation models on large datasets (millions or more of examples) using distributed training infrastructure such as PyTorch Distributed, DeepSpeed, or equivalent.
- Deep familiarity with tabular and sequence foundation model literature and a well-formed, defensible view of its limitations.
- Strong hands-on proficiency in Python and PyTorch, including CUDA-level distributed training and experiment tracking systems.
- Experience designing evaluation harnesses, ablation studies, and reproducible benchmarking pipelines that move models from research notebooks to multi-GPU training.
- Background with event-sequence or time-series transformers on transaction, clinical, or clickstream data is a strong plus.
- Prior work in credit, lending, financial services, or other regulated domains where model governance and explainability constrained architecture choices is highly valued.
- Publications at top-tier ML venues such as NeurIPS, ICML, ICLR, or KDD, or equivalent significant open-source contributions to foundation model research.
LOCATION
This is an on-site role. Primary location is San Francisco, CA, with additional office options in New York, NY and Washington, DC.
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: jobs.ashbyhq.com