The opportunity
WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence.
What you'll do
Lead an independent, high-impact research agenda on large language models and: agentic systems, owning projects from early hypothesis through model training, evaluation, and production deployment
Design and execute large-scale post-training experiments using supervised: fine-tuning, reinforcement learning from human feedback (RLHF), RLAIF, DPO, and emerging alignment techniques β with a focus on improving multi-step reasoning, planning, and tool use in enterprise agentic workflows
Build novel evaluation benchmarks and methodologies that push beyond existing: limitations, establishing rigorous measures for how well models perform on complex, real-world enterprise tasks
Develop scalable data synthesis and curation pipelines that generate the: high-quality training signal driving model improvement β including LLM-as-judge frameworks, synthetic data generation, and adversarial dataset construction
Shape WRITER's model architecture and training roadmap by translating your: research insights into concrete improvements to our enterprise-grade LLMs, working hand-in-hand with research engineering and product teams
Publish and present original research at top-tier venues: NeurIPS, ICLR, ICML, ACL, and others β representing WRITER at the frontier of the field and contributing to the broader scientific community
What they're looking for
- Mentor and uplevel fellow researchers and engineers on the team, helping set: a high bar for scientific rigor, experimental design, and research culture
- + years of hands-on ML research experience, with deep expertise in large: language model pre-training and post-training β you've trained models at scale, debugged distributed jobs, and shipped improvements that made a measurable difference
- A Ph.D. in Computer Science, Machine Learning, NLP, or a related field: or equivalent demonstrated research experience with a strong portfolio of independent, published work
- Expert-level knowledge of post-training methods including SFT, RLHF, RLAIF,: DPO, GRPO, and related alignment and reasoning techniques, with a track record of applying them to real, production-grade systems