Similar Jobs
See allMachine Learning Engineer - Large Language Models
Unknown
US
Large Language Models
PyTorch
VLLM
Machine Learning Engineer — Inference Optimization
Partner Company
Canada
PyTorch
CUDA
Machine Learning
Senior Machine Learning Engineer, ML Efficiency
US
PyTorch
Machine Learning
GPU
Staff Research Engineer, Model Efficiency
Cohere
Global
Machine Learning
Large Language Models
Software Engineering
Senior Inference Engineer
vCluster Labs
Global
Python
Golang
VLLM
Responsibilities:
- Design and implement advanced knowledge distillation pipelines, including teacher-student approaches, self-distillation, and multi-teacher architectures.
- Distill large foundation models into smaller, faster, and more efficient models optimized for production inference.
- Run large-scale machine learning experiments to evaluate model quality, latency, efficiency, and cost tradeoffs.
Requirements:
- Strong background in machine learning, deep learning, and neural network architectures.
- Hands-on experience implementing model distillation techniques for large language models or other neural networks.
- Solid understanding of training dynamics, optimization methods, loss functions, and model evaluation.
Benefits:
- Competitive compensation package with meaningful equity opportunities.
- Opportunity to work on core machine learning systems that directly impact product performance and efficiency.
- Remote-friendly work environment with an async-first culture.
Jobgether
Our partner is an innovative company focused on advancing the efficiency and scalability of next-generation machine learning systems. They offer a remote-friendly work environment with an async-first culture and a small, senior team combining research expertise and engineering excellence.