Lead the architecture and delivery of a large-scale GPU infrastructure platform, evolving from managed Kubernetes to bare-metal with Slurm and inference support.
Manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation, setting technical standards and overseeing implementation.
Own GPU infrastructure operations, including Slurm, Kubernetes, NVIDIA hardware, observability, and incident response, while acting as the primary technical interface with partners.
Design and build a managed Slurm service on Kubernetes
Write clean, reliable, and maintainable Go code, developing scheduling and orchestration for GPU workloads
Build observability and automated remediation for GPU, node, network, and control-plane failures
Gcore provides infrastructure and software solutions for AI, cloud, network, and security, powering digital experiences worldwide. With over 550 professionals, they build and support the global digital ecosystem.
Deploy and commission GPU cloud infrastructure across regional and core datacenters, covering hardware, networking, storage, and platform software.
Act as a hands-on technical escalation point, troubleshooting complex issues across physical and software layers and driving incidents to resolution.
Establish deployment standards, validation procedures, documentation, and operational practices for a rapidly evolving AI infrastructure environment.
The company builds and operates cutting-edge GPU cloud and AI infrastructure at significant scale. It is a fast-growing international scale-up with a diverse and flexible working environment.
Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
Collaborate with GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.
Bitdeer provides comprehensive Bitcoin mining solutions and AI computational infrastructure. The company operates globally with data centers in multiple countries and focuses on AI and blockchain technology.
Lead product engagement with strategic AI infrastructure customers to define requirements and drive execution from discovery to production readiness.
Collaborate cross-functionally with engineering, infrastructure, and operations teams to deliver customer-ready solutions.
Translate complex customer needs into clear product priorities, technical specifications, and scalable AI infrastructure offerings.
Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. It is a privately-held company valued at $3.5 billion, trusted by hundreds of thousands of customers across 185 countries, and known for its self-funded growth and inclusive culture.
Design, deploy, and operate large-scale Linux infrastructure including bare metal servers, enterprise storage, and GPU clusters for AI/ML workloads.
Manage and optimize AI Factory environments with NVIDIA GPU technologies such as A100, H100, and H200, including provisioning, monitoring, and performance tuning.
Ensure high availability and reliability through expert-level Linux administration, storage management with Ceph and high-performance platforms, and networking in data centers.
The company is seeking a senior Linux infrastructure engineer with expertise in bare metal, storage, and AI Factory platforms. The size, employees, and culture are not specified.
Build and operate the Kubernetes platform supporting AI test and evaluation frameworks.
Design infrastructure-as-code, GitOps workflows, and automated deployment pipelines.
Own platform reliability, observability, capacity planning, and operational readiness.
OpenTeams helps enterprises and governments build AI they control, govern, and evolve themselves. Founded by the creator of NumPy and SciPy, the company is built by people with deep roots across the open-source ecosystem and maintains a remote-first culture.
Design and build LLM serving infrastructure on Kubernetes, including deployment, GPU scheduling, and model lifecycle management.
Package the platform for enterprise environments with Helm-based installs and support for restricted or offline networks.
Integrate the serving layer with API gateway, identity, and metering services, and build observability for GPU inference in production.
Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build and operate scalable infrastructure for AI and data-intensive applications. It is part of an IREN company and empowers platform engineering teams with open-source innovation and deep expertise in Kubernetes orchestration.
Act as senior technical resource and final escalation point for strategic and VIP customers, owning complex issues across Kubernetes, GPU, and enterprise stack.
Train and mentor Technical Support Engineers in advanced Linux troubleshooting and customer architectures.
Author advanced troubleshooting documentation and drive incident resolution through root cause analysis.
Vultr makes high-performance cloud infrastructure easy to use and affordable for enterprises and AI innovators worldwide. It is the world's largest privately-held cloud infrastructure company with 33 data centers and hundreds of thousands of active customers, committed to growth and employee investment.
Own the platform including GCP, Kubernetes, Temporal, GPU fleet, and deploy/rollback machinery.
Contribute to AI enablement substrate: GPU capacity, training/inference pipelines, and cost optimization.
Strengthen team practices through tooling, standards, tests, observability, and release processes.
Descript is building a simple, intuitive, fully-powered editing tool for video and audio — an editing tool built for the age of AI. They are a team of 150 backed by top investors like OpenAI and Andreessen Horowitz, with a culture that values collaboration and serendipitous discovery.
Define the strategy, roadmap, and feature priorities for k0rdent AI Kubernetes services, empowering Neocloud operators to launch managed Kubernetes on their own GPU infrastructure.
Translate requirements from NeoClouds, GPU clouds, telcos, and enterprise platform teams into clear product direction and partner with engineering to ship secure, scalable cluster lifecycle capabilities.
Manage the Kubernetes backlog, define positioning and competitive differentiation, and create field-facing assets to support strategic accounts.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI. They have a world-class, distributed team committed to openness and technical excellence.
Own end-to-end infrastructure deployment programs for new capacity and site expansions, including hardware dependencies and commissioning gate frameworks.
Deliver crisp, data-driven executive updates and govern cross-organizational dependencies without escalation.
Coach junior TPMs and drive AI tool integration to improve program tracking and risk detection.
Evergrid builds frontier-level AI infrastructure for the fourth industrial revolution, designing, deploying, and operating large-scale AI systems. With over 1GW of capacity in active development and a 24/7 engineering team, they foster a culture of high conviction and high trust.
Build the technical product marketing function for the Provider business, creating collateral like white papers, reference architectures, and demo environments.
Directly support pipeline development by partnering with sales and solution architects through technical storytelling and proof-points.
Develop competitive intelligence and represent Mirantis as a credible technical voice on GPU infrastructure and sovereign AI.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It empowers platform engineering teams across any environment with a strong benefits plan and professional development.
Lead a distributed infrastructure operations team across US and APAC time zones, overseeing bare-metal environments and network engineering.
Own incident response, capacity planning, hardware lifecycle, and operational reliability across four global data centers.
Manage vendor relationships, budgets, and infrastructure strategy while collaborating with security and engineering leadership.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to ensure applications are reviewed quickly and objectively, and the platform is designed to streamline the hiring process.
Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.
Jobgether is a platform that connects job seekers with opportunities through an AI-powered matching process. The company fosters a remote-first culture and emphasizes autonomy and ownership for engineers.
Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
Together AI is a research-driven artificial intelligence company that builds open and transparent AI systems. The company has contributed to leading open-source research like FlashAttention and RedPajama, and aims to lower the cost of modern AI through co-designed software, hardware, algorithms, and models.
Design scalable storage architectures for edge and core GPU deployments using platforms like StorPool, NVMe, VAST Data, and Weka.
Optimize storage for AI workloads including distributed training, fine-tuning, and inference with GPU Direct Storage and RDMA.
Act as primary storage design authority, influencing platform architecture and mentoring engineers across infrastructure domains.
Radian Arc builds scalable storage architectures powering AI and GPU infrastructure across edge and core environments. They are a growing company with a remote-friendly work model and an inclusive environment focused on next-generation AI infrastructure.
Define DevOps strategy and lead infrastructure architecture across multi-environment, multi-region cloud systems.
Architect and own scalable Kubernetes platforms, infrastructure as code, and DevSecOps implementation.
Drive platform reliability, performance SLAs, cost optimization, and lead complex migrations and AI/ML platform infrastructure.
Robots & Pencils is an applied AI engineering firm that designs and ships AI co-workers for enterprise operations. Founded in 2009, the company has delivery centers across Canada, the US, Eastern Europe, and Latin America, with teams averaging over 15 years of experience.
Deploy and integrate high-performance NFS-based storage into Kubernetes clusters via CSI for GPU-accelerated workloads.
Automate storage provisioning and monitoring using infrastructure-as-code tools like Terraform and GitOps pipelines.
Tune Linux and network settings to optimize throughput and latency for demanding AI and machine learning applications.
Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build scalable, secure infrastructure for AI and data-intensive workloads. With deep expertise in open source and Kubernetes orchestration, they enable platform engineering teams across on-premises, cloud, edge, and sovereign environments.
Own the infrastructure layer for AI workloads including inference serving, Kubernetes, and agent-sandboxing platforms.
Manage the serving tier for open-weight models, Kubernetes operators, and stateful data planes.
Oversee the sandbox runtime, control-plane services, and observability tooling.
AZX accelerates positive impact in critical industries through AI transformation, specializing in physics-informed ML and enterprise AI solutions for climate and sustainability. Founded in 2024, the company is a profitable public benefit corporation with a growing team working with category leaders in real estate, energy, logistics, and utilities.
Design and build production AI agent systems for diagnosing and remediating infrastructure issues in large-scale GPU environments.
Develop distributed services, orchestration frameworks, knowledge graphs, and retrieval systems to power AI agents.
Own services end-to-end from architecture through production, collaborating with infrastructure and engineering teams.
The company builds AI agents that operate and automate large-scale GPU infrastructure. The engineering team is highly collaborative and remote, fostering ownership and autonomy.