Back

Member of Technical Staff, ML Systems

Full-Time · San Francisco

You will build Conway's model factory: the system that turns one foundation model into a multitude of deployed ones. The work spans training infra, adapter modularity and serving, and the routing algorithms that decides which model versions reach production. Frontier systems are powerful but brittle, and this role is where correctness and speed stop being tradeoffs. Every model Conway ships passes through systems you own.

Sample projects include:

  • Build a serving layer where one resident backbone composes with per-customer and per-task adapters at inference time
  • Design the versioning and lineage system that lets any production output be reconstructed exactly
  • Identifying bottlenecks throughout the model training and inference pipelines to optimize throughput, latency, and costs

Requirements:

  • Experience building infrastructure for training or serving large models — you know where the GPUs, the network, or the dataloader becomes the bottleneck, and how to find out which
  • Strong systems fundamentals: distributed computing, storage, inference engineers, profiling systems (Nsight, XProf, etc.)
  • Familiarity with descending the stack to trace bugs — from a training loss anomaly down to a kernel, a network error, or a quantization bug
  • Proficiency in Python

Nice to haves:

  • Experience with multi-tenant model serving, adapter systems, or inference optimization
  • Experience with at least one systems language
  • You've built ML platforms used by teams other than your own
  • Contributions to open-source ML infrastructure
  • Experience in at least one fast moving startup

Apply for this position

Choose File (PDF, DOC, DOCX)
Conway