Similar Jobs
See allSenior Machine Learning Engineer, ML Efficiency
US
PyTorch
Machine Learning
GPU
LLM Inference Engineer
NEAR AI
US
PyTorch
CUDA
Machine Learning Technical Lead, Artificial Intelligence (AI)
Gina's Tech Jobs
US
Python
PyTorch
JAX
Lead Machine Learning Engineer, Inference & Performance
Egen
Python
Kubernetes
VLLM
Principal Machine Learning Engineer, Artificial Intelligence (AI)
Confidential
United States
PyTorch
Deep Learning
GPU Optimization
Responsibilities:
- Optimize ML inference systems to improve latency, throughput, scalability, and cost.
- Profile and identify bottlenecks across GPU and CPU inference pipelines.
- Implement advanced optimization techniques such as quantization, KV-cache optimization, and speculative decoding.
Requirements:
- Strong experience in ML inference optimization and high-performance systems.
- Deep understanding of neural network architectures, attention mechanisms, and compute graphs.
- Hands-on experience with PyTorch, CUDA, and inference frameworks like TensorRT or vLLM.
Benefits:
- Competitive compensation package with meaningful equity participation.
- Opportunity to work on performance-critical AI systems with direct product impact.
- Flexible remote work environment and a culture that values technical excellence.
Partner Company
The company is an AI-focused organization that develops advanced machine learning systems for production environments. It values technical excellence and experimentation, offering a flexible remote work environment.