ML Systems Engineer, Inference

Runpod

Remote regions

Global

Salary range

$150,000–$220,000/yr

Benefits

Unlimited PTO

Similar Jobs

See all

Key Responsibilities:

  • Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build tooling for rigorous, repeatable measurements.
  • Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.
  • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.

Requirements:

  • 5+ years of professional system engineering experience with deep, hands-on experience in vLLM, SGLang, or a comparable serving engine in production.
  • Strong software engineering skills in Python and a solid understanding of what drives LLM inference performance: batching, memory, parallelism, and latency/throughput trade-offs.
  • Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving, plus rigor in benchmarking and GPU profiling.

What You'll Receive:

  • Competitive base pay from $150,000 - $220,000, meaningful equity, and generous medical, dental, and vision plans.
  • Flexible PTO, a $1,200 home office and equipment stipend, and the opportunity to join a passionate team on the cutting edge of AI infrastructure.
  • Most roles are remote-first with inclusive, collaborative teams using Slack for internal communication.

Runpod

Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. We're a small, remote-first team that takes ownership seriously, moves fast, and has processed more than 20 billion inference requests.

Apply for This Position