Kernel Engineer (ML Accelerators)

Location
Tel Aviv-Yafo, San Francisco
Workplace
On-site
Compensation
$200k – $240k + equity
Visa
Visa Sponsorship Available

About this role

About the role

You'll diagnose and resolve performance problems across Decart's ML systems, spanning research, training, and production inference. The largest share of the work is writing and optimizing kernels for TPU and Trainium. You'll also advise researchers on the performance cost of proposed model changes. We're looking for engineers with a demonstrated record in large-scale systems engineering and low-level optimization.

Minimum requirements

Bachelor's degree in Electrical/Computer Engineering, Computer Science, or a related field, plus 2+ years of relevant experience (or equivalent practical experience)

2+ years developing low-level software in C/C++ (Python proficiency a plus)

Solid grounding in operating systems fundamentals (process/thread scheduling, synchronization, virtual memory), CPU/GPU architecture, and hardware/software co-design

Working knowledge of PyTorch and machine learning algorithms, with a focus on engineering application

Demonstrated ability to profile compute and memory behavior, diagnose bottlenecks, and validate improvements with rigorous measurement

What we're looking for

Production experience squeezing performance out of ML workloads on TPU, Trainium, GPU, or other accelerators

You've authored kernels for an ML accelerator — not just consumed them

Depth in computer architecture: you can reason about systolic arrays, memory hierarchies, and interconnect topologies from first principles

Familiarity with compiler and toolchain internals (e.g., XLA, MLIR, Triton, or vendor stacks)

Experience scaling training or inference workloads across multi-accelerator clusters

You've read — or patched — the internals of an ML framework

Projects you might work on

Cut milliseconds off end-to-end token latency in DOS by restructuring attention and sampling paths for a new accelerator generation

Design communication schedules that overlap compute with network transfer across multi-chip topologies

Build analytical performance models to predict where the next 2x is hiding before writing a line of kernel code

Trace a throughput regression from a framework-level symptom down to instruction scheduling in generated assembly — and fix it

Port DOS's kernel suite to new hardware and close the gap to theoretical peak

What happens next

Skip the application pile. I get you in front of the people who decide.

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

A meeting lands on your calendar

When the company wants to meet, I get the call on your calendar. You just show up.

Know someone who'd be great for this?