About the role
ABOUT THE ROLE
This role owns how non-text modalities enter the pretraining run: the data we train on, the encoders and fusion architecture that carry it into the language model, and the capabilities we get out.
You will build and scale multimodal data pipelines (images, video, 3D, physical / scientific data), run the architecture research that decides how modalities are tokenized, encoded, and interleaved with text, and define the evaluations and ablations for validating your data and architecture. The work spans the full stack, from a data or architecture hypothesis to a controlled training experiment to a verdict that lands in the flagship recipe.
WHAT YOU'LL DO
- Build and scale multimodal data pipelines across images, video, 3D, and physical / scientific data
- Research and design how modalities are tokenized, encoded, and fused or interleaved with text in the pretraining architecture
- Define and run the evaluations and ablations that validate whether a data or architecture change actually improves multimodal capability
- Run controlled training experiments that test specific data or architecture hypotheses at scale
- Own the full loop: hypothesis, pipeline or architecture change, controlled experiment, verdict, flagship recipe change
- Partner closely with pretraining, data, and evals teams to land multimodal changes in the model that ships
WHAT WE'RE LOOKING FOR
- Strong track record in multimodal machine learning: vision-language models, encoders, tokenization, or modality fusion at pretraining scale
- Experience building and scaling data pipelines for non-text modalities (image, video, 3D, or scientific / physical data)
- Depth in architecture research, with the judgment to design experiments that isolate whether a change in encoding or fusion actually helps
- Comfort owning a problem end to end, from experimental design through to a recommendation that changes the flagship recipe
- Strong software engineering fundamentals for building and operating data and training pipelines at scale
- Prior experience at a frontier lab or similar large-scale multimodal training environment preferred
WHY THIS ROLE MATTERS
Multimodal capability is one of the clearest frontiers left in model quality, and this role controls the full path from raw data to shipped capability. The decisions made here determine what the model can see, watch, and understand beyond text.