ML Infrastructure Engineer

Location
Redwood City, China
Workplace
On-site
Compensation
$220k – $350k + equity

About this role

We are looking for an ML Infrastructure Engineer with 7+ years of experience to own training infrastructure end-to-end and turn a multi-cloud GPU fleet into a world-class training engine for massive multimodal models. You'll be the connective tissue between researchers and compute – designing distributed training systems, optimizing GPU utilization, and building the pipelines that ingest terabytes of multimodal robot data. This is a high-ownership role where your work directly accelerates the path from model to deployed robot. We need someone who has led technical projects in HPC or ML infrastructure and is genuinely passionate about the robotics space.

What you will be doing

Architecting and scaling distributed training infrastructure across large GPU clusters – implementing sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) for multimodal models

Building researcher-friendly tooling and job scheduling systems (Kubernetes/SLURM) that prioritize fast iteration, automated retries, and seamless failure recovery

Designing high-throughput data pipelines to ingest and transform terabytes of multimodal robot data (video, proprioception, 3D signals) so dataloaders never starve the GPUs

Building low-latency inference pipelines for real-time robot control – applying quantization, distillation, and model compilation (TensorRT, Triton) to move models from lab to physical world

Deep systems profiling – diving into GPU utilization, I/O bottlenecks, and memory fragmentation to squeeze maximum performance out of an expanding compute fleet

What happens next

Skip the application pile. I get you in front of the people who decide.

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

A meeting lands on your calendar

When the company wants to meet, I get the call on your calendar. You just show up.

Culture & values

Values low ego, knowledge sharing, and genuine passion for the physical AI space

Engineers have significant say in technical roadmap and strategy

Many roles offer 0-to-1 building opportunities with high ownership and autonomy

Cross-functional collaboration with researchers, hardware engineers, operations teams, and leadership

Casual approach to scheduling with work-from-home flexibility when needed

No clock-watching culture

Team filled with people genuinely excited about physical AI and robotics

Many team members follow the space closely and build projects outside of work

Open-mindedness and collaborative problem-solving prioritized over hierarchy or rigid opinions

Positions itself as a product-research company where all research is done with deployment in mind

Company moves extremely fast with speed in interview and decision-making processes

Operates with startup intensity and casual, flexible approach to work

Know someone who'd be great for this?

Top Benefits

  • Unlimited PTO
  • Work-from-home flexibility