Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms, owning the storage layer where Kubernetes meets bare metal.
Tune NFS data paths for high-throughput, low-latency GPU/AI workloads, integrating NFS-based storage into clusters via CSI, storage classes, and persistent volumes.
Automate storage provisioning with infrastructure-as-code (Terraform/OpenTofu) and GitOps pipelines, and build monitoring and observability for storage performance and health.
Design and build control-plane services, drivers, and tooling for high-performance storage integration with Kubernetes.
Write production-quality Go software with strong testing and operational rigor for automation.
Deliver storage integration for Kubernetes via Cluster API and K0rdent in hybrid and air-gapped environments.
Mirantis is a Kubernetes-native AI infrastructure company enabling organizations to build scalable infrastructure for AI and data-intensive applications. A Silicon Valley leader with a young, collaborative culture.
Lead investigation and resolution of complex infrastructure, networking, and platform incidents.
Provide technical leadership for Kubernetes platform operations and drive automation initiatives.
Mentor engineers and develop operational standards, runbooks, and best practices.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Serving enterprises like Adobe, PayPal, and Volkswagen, Mirantis is committed to open standards and freedom from lock-in.
Own the infrastructure and platform powering the marketplace, focusing on reliability, observability, security, and automation.
Manage production AWS and EKS clusters, infrastructure as code with Terraform and GitOps, and CI/CD pipelines via GitHub Actions.
Build automation and internal tooling in Python, Bash, Go, and Node.js/TypeScript, and operate PostgreSQL, MongoDB, and Temporal.
Office Hours is an on-demand expert network that connects leading organizations with trusted experts across various knowledge domains. The company is hyper-growth, profitable, and expanding quickly, backed by top marketplace investors.
Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.
Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
Drive reliability, monitoring, automation, and incident response for AI infrastructure.
Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.
Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.
Design, operate, and improve reliable infrastructure for AI training and inference workloads.
Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.
Design, build, and maintain platform infrastructure using IaC principles with tools like Terraform.
Develop memory services, vector storage patterns, and semantic search capabilities.
Operate a Kubernetes-based platform for the Data & AI department, enabling deployment and scaling via GitOps.
Redcare Pharmacy is Europe's No.1 e-pharmacy, striving to improve global health through innovation and collaboration. The company fosters a healthy, inclusive work environment where employees feel valued and inspired.
Lead engineers deploying, operating, and optimizing AI compute clusters at scale.
Oversee cluster reliability, GPU fleet operations, and incident response.
Coordinate with cross-functional teams to ensure cluster capabilities meet AI workloads.
Vultr makes high-performance cloud infrastructure easy to use, affordable, and locally accessible for global enterprises and AI innovators. It is the world’s largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries.
Mature deployment artifacts for enterprise customers, including hardened images, Terraform modules, and Helm charts.
Own reference architectures and improve the upgrade experience for seamless, predictable operations.
Mentor other engineers and contribute to architecture decisions across engineering teams.
Coder is an AI software development company that empowers teams to build software faster and more securely through autonomous coding agents. They are a growing company that values innovation, collaboration, and high standards for code quality and security.
Support and improve production and development infrastructure for multiple teams handling high traffic products.
Help developers debug intricate issues and architect scalable solutions across cloud and on-premise environments.
Promote CICD strategies, document processes, and mentor junior DevOps engineers.
We are a tech pioneer offering world-class adult entertainment and games on safe platforms. With an international team of dynamic innovators, we have offices in Montreal, Austin, and Nicosia and celebrate diversity and inclusion.
End-to-end ownership of internal orchestration platform built on event-driven architecture with Redpanda, including code, architecture, and roadmap.
Own infrastructure-as-code using Terraform Cloud, manage Kubernetes workloads with Helm, and provide self-service tooling for engineering teams.
Set SLOs, handle production on-call, lead incident response, author design docs, and operate AI-natively using tools like Cursor and Notion AI.
Velora unifies Aplos, Raisely, and Keela into one company with a shared mission to help nonprofit organizations thrive by offering fundraising, donor management, financial tracking, and communications tools. We are a financially solid company with a combined team dedicated to making nonprofit work easier, more impactful, and more sustainable.
Lead core infrastructure and SRE teams to ensure the highest reliability and performance for massive GPU computing demands.
Oversee HPC networking and distributed storage engine innovation to support massive multi-node AI workloads.
Build and scale a high-output engineering org while partnering cross-functionally with product and GTM leadership.
Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. As a small, remote-first team that closed a $100M Series A, we move fast and take ownership seriously.
Design, implement, and maintain reliable, scalable infrastructure, applications, and tooling for software-defined manufacturing.
Write clean, maintainable code and perform peer code reviews to ensure high-quality deliverables.
Collaborate with cross-team members to prototype new technology and evaluate technical feasibility.
Bright Machines is an innovator in software-defined manufacturing, using intelligent automation to transform factories. The company is building a team of experts to create a new manufacturing category, offering a culture of innovation and impact.
Build and operate the infrastructure behind AI-powered products, improving reliability, security, scalability, and cost efficiency.
Write code, automate infrastructure, investigate production issues, and design systems that reduce operational complexity.
Take ownership of unfamiliar systems, identify highest-leverage improvements, and balance immediate production needs with long-term platform investments.
Zencoder builds and orchestrates AI agents that ship real work across code, research, and operations. It is a growing platform where people and agents collaborate, with a high-caliber team and a culture that values individual contributors.
Lead deployment and operation of product infrastructure in federal environments within AWS.
Build and maintain scalable, secure cloud-native platforms using Kubernetes, Terraform, and GitLab CI.
Improve development and deployment processes, create tooling for telemetry, and foster documentation culture.
Horizon3.ai is a fast-growing, remote cybersecurity company that helps organizations proactively find and fix exploitable attack vectors. We are a team of former special ops cyber operators and engineers committed to a culture of respect, collaboration, ownership, and results.
Support engineering teams by developing resilient applications on GKE and advising on best practices.
Develop and maintain infrastructure-as-code using Terraform, ArgoCD, and Python.
Manage GKE environments across multiple regions, including IAM and identity management.
Mimica uses AI-powered task mining to observe employee actions and create process maps, helping enterprises improve efficiency. The company is a fast-growing scale-up with a lean, collaborative culture.
Establish and promote modern DevOps culture and practices across engineering teams.
Deploy, manage, and optimize both on-premises and cloud-based Kubernetes clusters.
Automate and continuously improve CI/CD pipelines to ensure seamless delivery.
Kyivstar.Tech is a Ukrainian hybrid IT company and a subsidiary of Kyivstar, one of Ukraine’s largest telecom operators, creating innovative technological solutions that transform lives. With over 500 specialists, we embrace an entrepreneurial culture that fosters continuous growth and challenges conventional thinking.
Lead the design and operation of GPU infrastructure for AI workloads.
Manage Kubernetes-based environments and optimize for AI training and inference.
Define operational standards, implement monitoring, and collaborate with AI engineering teams.
ELEKS is a software engineering company that partners with enterprises to accelerate digital transformation. They have a global team of over 2,000 professionals and foster a culture of innovation and collaboration.
Operate and maintain Linux-based infrastructure, deploy and scale Kubernetes clusters, and implement automation with Ansible and GitOps.
Design networking architecture, build observability stacks, and lead incident response across the platform.
Manage virtualization layers and collaborate with development teams to optimize resource utilization and system availability.
Pragmatike develops cutting-edge solutions in Cloud Computing, focusing on ambitious projects with a culture of collaboration and innovation. The team is passionate and collaborative, working in a dynamic and flexible environment to shape tomorrow's technologies.
Architect and evolve a Kubernetes-native platform for NBC's broadcast production environments, leading technical strategy and automation.
Design and build production-grade Go services, Kubernetes operators, and custom Crossplane providers for multi-account AWS and hybrid cloud.
Drive GitOps-based continuous delivery, observability, and mentoring while bridging broadcast hardware and automated infrastructure.
NBCUniversal is a leading media and entertainment company creating world-class content across film, television, streaming, and theme parks. As a subsidiary of Comcast, it employs thousands and fosters an inclusive culture with a focus on community and diversity.