Similar Jobs
See allSoftware Golang Engineer (Slurm)
Gcore
Global
Go
Kubernetes
Slurm
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
Partner Company
Europe
GPU Infrastructure
Linux
Networking
Staff AI Scheduling & Orchestration Engineer
Bitdeer
US
Kubernetes
Distributed Systems
Terraform
Principal Product Manager
Vultr
Global
Kubernetes
Slurm
Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)
Unknown
Global
Linux
Ceph
Networking
About the Role:
- Lead the architecture and delivery of a sophisticated GPU infrastructure platform.
- Manage a distributed team and ensure high technical standards across all engineering disciplines.
Technical Leadership:
- Own end-to-end platform architecture, including design reviews and documentation.
- Make key decisions on distributed systems, HPC, networking, storage, and GPU workloads.
Team and Operations:
- Hire, develop, and mentor team members, fostering a high-performance culture.
- Oversee incident response, on-call model, and sustainable operations for a lean team.
Collaboration and Growth:
- Work with research and product teams to translate workloads into platform requirements.
- Enjoy a 100% remote environment with opportunities for travel and professional growth.
Jobgether
Jobgether is an AI-powered job matching platform that connects candidates with relevant roles, ensuring a fair and objective review process. It operates with a distributed team and partners with companies globally, focusing on efficient and transparent recruitment.