About the role
About the Role
Work alongside an AI research team to turn experimental code into production-ready systems. You will build and improve shared post-training infrastructure that supports research data generation, processing, and evaluation.
What You'll Do
* Turn research prototypes into tested, reusable data generation, training, and evaluation pipelines.
* Build distributed experiment support, including data loading, checkpointing, and collection of agent interactions.
* Profile workloads and improve GPU utilization, memory efficiency, and data throughput.
* Develop tests, experiment tracking, and debugging tools while preserving the integrity of research results.
* Collaborate with researchers to translate ideas into maintainable software.
What We're Looking For
* Experience building ML training, inference, or data pipelines in Python.
* Strong Python skills and hands-on experience with PyTorch, JAX, or comparable machine learning frameworks.
* Understanding of machine learning experiments, including how data, numerical precision, and implementation choices affect results.
* Sound software engineering practices, including testing, profiling, version control, and documentation.
* Distributed training or data processing experience with tools such as PyTorch Distributed, DeepSpeed, or Ray is valuable.
* Familiarity with Hugging Face Transformers, vLLM, or experiment tracking platforms such as Weights & Biases or MLflow is useful.
* Experience with reinforcement learning systems or multimodal datasets is a plus.
Compensation & Benefits
Competitive compensation and equity are offered. Visa sponsorship is available.
Location
On-site in Zürich, Switzerland.