Kernel Engineer (ML Accelerators)
- Location
- Tel Aviv-Yafo, San Francisco
- Workplace
- On-site
- Compensation
- $200k – $240k + equity
- Visa
- Visa Sponsorship Available
About this role
About the role
You'll diagnose and resolve performance problems across Decart's ML systems, spanning research, training, and production inference. The largest share of the work is writing and optimizing kernels for TPU and Trainium. You'll also advise researchers on the performance cost of proposed model changes. We're looking for engineers with a demonstrated record in large-scale systems engineering and low-level optimization.
Minimum requirements
Bachelor's degree in Electrical/Computer Engineering, Computer Science, or a related field, plus 2+ years of relevant experience (or equivalent practical experience)
2+ years developing low-level software in C/C++ (Python proficiency a plus)
Solid grounding in operating systems fundamentals (process/thread scheduling, synchronization, virtual memory), CPU/GPU architecture, and hardware/software co-design
Working knowledge of PyTorch and machine learning algorithms, with a focus on engineering application
Demonstrated ability to profile compute and memory behavior, diagnose bottlenecks, and validate improvements with rigorous measurement
What we're looking for
Production experience squeezing performance out of ML workloads on TPU, Trainium, GPU, or other accelerators
You've authored kernels for an ML accelerator — not just consumed them
Depth in computer architecture: you can reason about systolic arrays, memory hierarchies, and interconnect topologies from first principles
Familiarity with compiler and toolchain internals (e.g., XLA, MLIR, Triton, or vendor stacks)
Experience scaling training or inference workloads across multi-accelerator clusters
You've read — or patched — the internals of an ML framework
Projects you might work on
Cut milliseconds off end-to-end token latency in DOS by restructuring attention and sampling paths for a new accelerator generation
Design communication schedules that overlap compute with network transfer across multi-chip topologies
Build analytical performance models to predict where the next 2x is hiding before writing a line of kernel code
Trace a throughput regression from a framework-level symptom down to instruction scheduling in generated assembly — and fix it
Port DOS's kernel suite to new hardware and close the gap to theoretical peak
What happens next
Skip the application pile. I get you in front of the people who decide.
Confirm the fit
A few questions to make sure this role is the right shape for you. Two minutes.
I pitch you to the company
I write the intro, send it to the founder, and handle the back-and-forth.
A meeting lands on your calendar
When the company wants to meet, I get the call on your calendar. You just show up.
Know someone who'd be great for this?
