Research Scientist (Philosophy)
RESEARCH SCIENTIST, PHILOSOPHY PROGRAM
ABOUT RESOLUTION
Resolution does research on how to align artificial superintelligence (ASI). ASI may be developed in the next few years, but it is unclear whether alignment is on track to be ready in the same timeframe. We aim at higher a priori confidence in aligned outcomes by pursuing a portfolio of theory and empirics bets, any one of which — if it succeeds — would meaningfully advance the field. We invest heavily in research automation to accelerate progress, and we believe that stronger alignment theory unlocks higher automation: more principled approaches give us better filters for which directions of automated research are promising.
Resolution was founded in 2026 by researchers from UK AISI's Alignment Team, who ran the £30m Alignment Project https://alignmentproject.aisi.gov.uk/, and Timaeus https://timaeus.co/, who pioneered applying singular learning theory to alignment.
For more information, see our announcement https://resolution.org/launch/.
ABOUT PHILOSOPHY AT RESOLUTION
Work on AI alignment raises philosophical questions, including what concepts like honesty mean, what character is, what moral competence looks like, which moral theories should guide model behavior, and what research ought (or ought not) be delegated to models. Resolution’s philosophy team seeks to advance alignment by examining these kinds of questions and testing empirical claims where experimentation can help. Some projects will be primarily conceptual; others will involve collaboration with empirical researchers.
Our initial research agenda is under development. The following gives a flavor of the areas of research we are considering:
- Conceptual engineering. Alignment research depends on concepts such as honesty, harm, autonomy, and deference, each of which carries multiple senses in ordinary language. We want to understand what it means for a model to represent or operationalize one sense rather than another; how that representation can be changed through training data, context, or other interventions; and which version of a concept(s) a model should employ in a given setting. The goal is to characterize alignment-relevant concepts precisely enough to inform training targets, experimental hypotheses, and evaluation criteria.
- Character, traits, and constitutions. AI developers increasingly attempt to align models by shaping their character rather than specifying comprehensive sets of rules. But the underlying notions, including trait, persona, disposition, and character, remain poorly understood. We want to clarify the relationships among these concepts; determine which, if any, provides the right unit of analysis for alignment targets; investigate which traits advanced systems should possess; and study how those traits persist, generalize, and interact across contexts.
- Normative evaluation of models. What would moral competence in a model consist in? Does it require distinctive capacities, or does it emerge from more general forms of conceptual and practical competence? We want to design experiments that clarify how models make normative judgments and develop evidential standards for determining when, and in which contexts, reliance on those judgments may be warranted.
- Moral theories and alignment. Alignment methods presuppose some normative commitments and assumptions. These assumptions can enter system design through objectives, constitutions, training examples, or evaluation criteria. Sometimes, an explicit commitment to some moral theory is made in model development. Systems shaped by different normative frameworks may behave similarly at current capability levels while diverging substantially as their capabilities increase. We want to understand how different normative structures and moral theories are instantiated in AI systems, and what evidence could help us compare their likely alignment properties as capabilities improve.
- Epistemic and safety risks of automating conceptual reasoning. Conceptual reasoning work is increasingly being delegated to models, including within alignment research itself. It's an open question whether this should happen at all, and, if so, where the limits should be. We want to map the risks and help determine which forms of conceptual work can be delegated, under what conditions, and with what forms of oversight. We are also interested in methods for validating model-generated conceptual work, and in the scientific norms and evidential standards that can help govern AI-assisted alignment research.
We are also interested in proposals for new philosophical research that could improve our understanding of alignment or alter how alignment research is conducted.
ABOUT THE ROLE
The team is new, and the people we hire now will play a significant role in shaping what it becomes. We are hiring research scientists across multiple levels. The scope of the role will depend on the candidate’s experience, research interests, and level of seniority.
Research scientists may:
- Lead or contribute to original research projects in philosophy and AI alignment.
- Identify important but underspecified problems and turn them into projects.
- Connect philosophical analysis to empirical hypotheses, experimental designs, training approaches, or evaluation methods.
- Collaborate closely with researchers from machine learning and other relevant disciplines.
- Communicate research clearly through papers, internal reports, presentations, and external engagement.
- Contribute to the intellectual culture and development of the philosophy team and Resolution.
Depending on seniority, research scientists may also:
- Define and lead broad or ambiguous areas of research.
- Help set the team’s research agenda and strategic priorities.
- Initiate and lead collaborations across Resolution and with external researchers.
- Supervise, mentor, or manage more junior researchers.
- Represent the philosophy team in organization-wide res
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: jobs.ashbyhq.com