Similar Jobs
See allSenior Inference Engineer
vCluster Labs
Global
Python
Golang
VLLM
AI Inference Engineer
Baseten
US
Python
Machine Learning
Software Engineering
Senior AI Engineer
Pragmatike
US
Python
PyTorch
Transformers
Staff AI Engineer
Blip
Brazil
Python
PyTorch
Machine Learning
Senior Forward Deployed Engineer
Eliza
Global
Python
AWS
GCP
Accountabilities:
- Build and deploy production-grade LLM inference systems across one or more GPU machines, owning the complete pipeline from customer query to served response.
- Design, implement, and operate model-serving infrastructure using technologies such as vLLM, SGLang, and TensorRT-LLM.
- Optimize inference workloads for scale, balancing latency, throughput, reliability, and infrastructure costs.
Requirements:
- Significant professional experience building and operating production software or infrastructure systems, with strong hands-on engineering capabilities.
- Demonstrated experience deploying and serving large language models in production, ideally using vLLM, SGLang, TensorRT-LLM, or comparable inference frameworks.
- Strong programming skills in Python or Golang, with a track record of writing and maintaining production-quality code.
Benefits:
- Competitive compensation package including equity, health, dental, vision, and life insurance.
- Flexible working schedule focused on outcomes rather than fixed hours, with high workplace flexibility supporting remote work.
- Significant ownership over the architecture, implementation, and long-term roadmap of the inference platform, with direct collaboration with senior technical leadership and Product teams.
Undisclosed
The company is an open-source-oriented startup building production-grade inference infrastructure for large language models. It is a remote-first, globally distributed team that values engineering ownership, speed, and customer impact.