Similar Jobs
See allSenior Inference Engineer
vCluster Labs
Global
Python
Golang
VLLM
Senior Inference Engineer
Undisclosed
Brazil
Python
Golang
VLLM
Research Engineer Intern - AI Systems
Yotta Labs
Global
CUDA
Python
C++
AI Inference Engineer
Baseten
US
Python
Machine Learning
Software Engineering
Staff Research Engineer, Model Efficiency
Cohere
Global
Machine Learning
Large Language Models
Software Engineering
Responsibilities:
- Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
- Optimize long-context prefill and decode workloads based on real production traffic.
- Tune routing between infrastructure and external providers based on cost, capacity, and performance.
Qualifications:
- 5+ years in ML systems, inference infrastructure, or performance engineering, with measurable improvements in cost or latency.
- Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
- Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
Benefits:
- Flexible work with in-person collaboration in the Bay Area and a distributed global-first team.
- Adaption Passport: annual travel stipend to explore a new country.
- Lunch stipend and comprehensive medical benefits with generous paid time off.
Adaption
Adaption builds efficient AI that evolves in real-time, making intelligence flexible, personalized, and accessible to everyone. They focus on talent density, bringing together driven individuals to push the boundaries of continual adaptation.