Similar Jobs
See allInference Infrastructure Architect
Telnyx
China
VLLM
SGLang
Kubernetes
Technical Lead - GPU Infrastructure
Jobgether
Canada
Kubernetes
Linux
Senior Platform Engineer, GitLab Orbit
GitLab
Canada
Rust
Kubernetes
Distributed Systems
ML Systems Engineer, Inference
Runpod
Global
Python
VLLM
SGLang
Python Inference Engineer
Gcore
Global
Python
PyTorch
Kubernetes
About the role:
- Owns container runtime for AI inference and training.
- Works across Linux, storage, GPUs, and distributed systems.
What you'll work on:
- Speed up container startup, checkpoint, and restore.
- Build GPU-aware snapshotting and optimize memory/filesystem paths.
- Extend sandboxed runtime for new GPUs and profiling.
What we're looking for:
- Strong Linux systems and container runtime knowledge.
- Experience with sandboxes, kernel, or low-level infrastructure.
- Proficiency in Rust, Go, or another systems language.
Modal
Modal builds serverless cloud infrastructure for AI/ML inference and training workloads. The company is an engineering-driven team that values deep systems work and open-source contributions.