Posted last month
Builds evaluation infrastructure for Anthropic's safeguards agent, designing experiments, datasets, and metrics to measure misuse detection, and productionizes research into pipelines. Works across ML research and engineering to ensure trustworthy AI safety systems.