Similar Jobs

See all

Responsibilities:

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between infrastructure and external providers based on cost, capacity, and performance.

Qualifications:

  • 5+ years in ML systems, inference infrastructure, or performance engineering, with measurable improvements in cost or latency.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.

Benefits:

  • Flexible work with in-person collaboration in the Bay Area and a distributed global-first team.
  • Adaption Passport: annual travel stipend to explore a new country.
  • Lunch stipend and comprehensive medical benefits with generous paid time off.

Adaption

Adaption builds efficient AI that evolves in real-time, making intelligence flexible, personalized, and accessible to everyone. They focus on talent density, bringing together driven individuals to push the boundaries of continual adaptation.

Apply for This Position