- Job type
- Full-time
- Work mode
- On-site
- Level
- Not listed
- Department
- Engineering
- Experience
- Not listed
- Posted
- Sep 24, 2026
About the role
About the role
Make our models performant, resource efficient and reliable at scale by developing the core frameworks for model training and evaluation, in close partnerships with fellow researchers and engineers.
- Build our training stack across model, layer, and kernel levels; optimize workloads through parallelism, quantization, and custom kernels.
- Profile end-to-end training runs on large GPU clusters; eliminate bottlenecks and failures; monitor throughput, utilization, and uptime.
- Ensure new model architectures and training recipes scale efficiently, from early experiments to frontier-scale runs.
- Make our ML training stack maximally reliable: fault tolerance, checkpointing, and deterministic orchestration for long-running, large-scale jobs.
Chai's models are moving beyond protein structure prediction into real-world therapeutic engineering. This is a chance to push the frontier of AI drug design, working alongside a rigorous and craft-obsessed team.
About you
Ideal backgrounds include deep industry experience working with top AI/ML teams on the kinds of problems and systems we describe above—with strong software system design skills, proficiency in Python, and Pytorch or JAX fluency. We look for technical spikes where you have gone deep and demonstrated exceptional impact on real-world problems and systems.
We offer
The opportunity to work at the vanguard of AI research and frontier biology, with world-class people, on a mission that matters. We protect & promote a culture of high velocity and ownership. We compensate our team accordingly.