Similar Jobs
See allSenior Solutions Engineer
TensorWave
US
Kubernetes
GPU Infrastructure
Linux
Technical Support Engineer (GPU Clusters)
Together AI
US
Kubernetes
GPU
Ansible
Forward Deployed Engineer - SRE
Andromeda Cluster
North America
GPU Clusters
InfiniBand
Kubernetes
Infrastructure/GPU Cluster/Platform Operations Lead
ELEKS
Canada
Kubernetes
NVIDIA GPU
CUDA
Site Reliability Engineer
Boson AI
Global
Kubernetes
Linux
Networking
Who You Are:
- You understand how accelerator compute, memory, and networking topology constrain AI workloads.
- You work directly with customers to understand their challenges and provide effective solutions.
- You are comfortable debugging across the full inference stack.
What We Need:
- Strong software engineering skills with 5+ years of relevant technical experience.
- Experience turning ambiguous customer requirements into verifiable acceptance criteria.
- Kubernetes and Helm experience at multi-node, HPC, or AI cluster scale.
What You Will Learn:
- Where co-design of AI hardware and software translates into unique latency and throughput performance.
- How to scale disaggregated inference services on Kubernetes.
- Why customer insights from the field shape the best products.
Tenstorrent
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.