Vinny's ListJobs

Runtime Engineer

MatX

<h3><strong>What MatX is Building</strong></h3> <p>MatX is building custom silicon for large-language-model inference and training, with HW/SW co-design across ISA, RTL, simulator, compiler, and kernels so each layer benefits from the others. The runtime owns the host-side stack and the contracts that bind those teams together.</p> <h3><strong>What You'll Do Here</strong></h3> <ul> <li>Build the host-side interface library — device memory management, DMA, streams and events, sync primitives — that every compiler-emitted program runs on top of</li> <li>Own and extend the executable format: the compiler→runtime contract, its versioning, the weight and quantization layouts that let compiler and runtime evolve independently</li> <li><strong>Design the custom-kernel ABI</strong> — calling convention, sync semantics, lifecycle — and the host-side marshaling layer (DLPack, the buffer protocol, numpy) that gets Python tensors to the device</li> <li>Build Python bindings via PyO3, with a C-ABI shim as the alternative integration path for downstream consumers</li> <li>Build the LLM inference serving stack — paged KV cache, continuous batching, request scheduling, token streaming — and the cluster orchestration primitives underneath it</li> <li>Bring up interconnect topology from the host and own the failure-detection and clean-teardown path for stop-restructure-resume recovery across racks</li> <li>Design what the chip exposes to host-side profilers and debuggers — perf counters, traces, and the Python surfaces ML engineers actually use — and hit measurable performance targets on runtime overhead and serving throughput</li> </ul> <h3><strong>Who You Are</strong></h3> <ul> <li>Strong experience in a systems programming language — Rust, C, C++, or Go — including memory management, allocator design, and FFI/ABI work</li> <li>Have built Python interop layers in production (PyO3, ctypes, pybind11, or equivalent C-ABI bridging)</li> <li>Have designed and maintained API or ABI contracts between teams — versioning, evolution, breaking-change discipline — not just consumed someone else's</li> <li>Hands-on with at least one accelerator programming model (CUDA, ROCm, oneAPI Level Zero, TPU, or comparable) — enough to reason about device memory, async execution, and kernel launch</li> <li>ML-systems literate — comfortable with the training and inference loop, what collectives do, what a tensor layout is. Research depth not required.</li> </ul> <h3><strong>Bonus Points If You Have</strong></h3> <ul> <li>LLM inference internals — vLLM, TensorRT-LLM, or SGLang (paged attention, scheduler design)</li> <li>Rust at depth, including proc macros, unsafe with soundness reasoning, and complex lifetime/trait work</li> <li>Custom allocator design (slab, paged, arena) or other low-level memory work</li> <li>ML framework integration experience (PyTorch custom backends, JAX/XLA, ONNX runtime)</li> <li>Profiler or tracing infrastructure work (perfetto, Nsight, or a custom stack)</li> <li>Driver-adjacent or kernel-bypass work, or prior new-silicon bring-up</li> </ul> <h3><strong>Compensation</strong></h3> <p>The US base salary for this full-time position is determined based on a variety of factors including role, experience, location, job related skills, and relevant education and training. Career length is only a guideline for compensation.</p> <ul> <li>Early Career - $120,000 - $250,000 + equity</li> <li>Mid Career - $175,000 - $362,500 + equity</li> <li>Senior Career - $250,000 - $475,000 + equity</li> </ul&gt

Role

Role
Manufacturing Engineer
Workplace
Onsite

Location

City
Mountain View
State
CA

Company

Company
MatX
Industry sector
electronics_electrical

Source

Job board
greenhouse
Posted
2026-06-03
Posted
Last 90 days

Open in the interactive directory →