ML Infrastructure Engineer
- Location
- Redwood City, China
- Workplace
- On-site
- Compensation
- $220k – $350k + equity
About this role
We are looking for an ML Infrastructure Engineer with 7+ years of experience to own training infrastructure end-to-end and turn a multi-cloud GPU fleet into a world-class training engine for massive multimodal models. You'll be the connective tissue between researchers and compute – designing distributed training systems, optimizing GPU utilization, and building the pipelines that ingest terabytes of multimodal robot data. This is a high-ownership role where your work directly accelerates the path from model to deployed robot. We need someone who has led technical projects in HPC or ML infrastructure and is genuinely passionate about the robotics space.
What you will be doing
Architecting and scaling distributed training infrastructure across large GPU clusters – implementing sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) for multimodal models
Building researcher-friendly tooling and job scheduling systems (Kubernetes/SLURM) that prioritize fast iteration, automated retries, and seamless failure recovery
Designing high-throughput data pipelines to ingest and transform terabytes of multimodal robot data (video, proprioception, 3D signals) so dataloaders never starve the GPUs
Building low-latency inference pipelines for real-time robot control – applying quantization, distillation, and model compilation (TensorRT, Triton) to move models from lab to physical world
Deep systems profiling – diving into GPU utilization, I/O bottlenecks, and memory fragmentation to squeeze maximum performance out of an expanding compute fleet
What happens next
Skip the application pile. I get you in front of the people who decide.
Confirm the fit
A few questions to make sure this role is the right shape for you. Two minutes.
I pitch you to the company
I write the intro, send it to the founder, and handle the back-and-forth.
A meeting lands on your calendar
When the company wants to meet, I get the call on your calendar. You just show up.
Culture & values
Values low ego, knowledge sharing, and genuine passion for the physical AI space
Engineers have significant say in technical roadmap and strategy
Many roles offer 0-to-1 building opportunities with high ownership and autonomy
Cross-functional collaboration with researchers, hardware engineers, operations teams, and leadership
Casual approach to scheduling with work-from-home flexibility when needed
No clock-watching culture
Team filled with people genuinely excited about physical AI and robotics
Many team members follow the space closely and build projects outside of work
Open-mindedness and collaborative problem-solving prioritized over hierarchy or rigid opinions
Positions itself as a product-research company where all research is done with deployment in mind
Company moves extremely fast with speed in interview and decision-making processes
Operates with startup intensity and casual, flexible approach to work
Know someone who'd be great for this?
