Back
Member of Technical Staff, Research
Full-Time · San Francisco
You will train the Wake series, Conway's foundation models over event streams. The work spans pretraining over billions of events, architecture and objective design, and the distributed training systems required to run decisive experiments at scale. You own the backbone that every customer model inherits from.
Sample projects include:
- • Pretrain a transformer over pooled event streams and beat the supervised-from-scratch baseline
- • Take a training run from replicated data parallelism to sharded training across a cluster, and debug it when throughput collapses
- • Design ablations that attribute a downstream gain to a specific data, objective, or architecture decision
Requirements:
- • Hands-on experience pretraining or continued-pretraining models at meaningful scale. You have owned a multi-GPU run end to end and eaten its failure modes: loss spikes, dataloader stalls, silent data bugs
- • Strong grounding in the modern training stack — PyTorch internals, distributed training, mixed precision
- • You reason from downstream performance, not training loss. You kill your own ideas quickly when the evidence turns
Nice to haves:
- • Experience with sequence models over non-text data: events, time series, tabular, audio
- • Optimizer or training-efficiency research
- • First-author publications at NeurIPS, ICML, ICLR, or comparable venues
- • You've read our technical report and have opinions about it