Remote Devops Jobs · Kubernetes

Job listings

$170,000–$218,000/yr

  • Lead compliance engineering for FedRAMP High-aligned environments, translating obligations into practical infrastructure changes.
  • Design and operate secure cloud infrastructure using AWS, Terraform, and Kubernetes across commercial and GovCloud environments.
  • Improve platform reliability, release operations, and incident response while mentoring engineers on secure operations.

Arch Systems empowers discrete manufacturing facilities with data insights for optimal efficiency and proactive decision-making. It is a rapidly scaling, remote-first company that values collaboration, innovation, and continuous learning.

$6,500–$9,000/mo

  • Architect, deliver, and maintain critical cloud platform components on AWS EKS, focusing on production reliability and observability.
  • Establish SRE standards including SLOs, error budgets, and automated tooling to reduce operational friction.
  • Provide technical advisory through code reviews and architecture recommendations to maintain high platform standards.

Inflect is a US-based advisory and marketplace revolutionizing digital infrastructure procurement. They are a small team focused on reducing friction in buying datacenter, cloud, and network services through automation and better deal terms.

UK 5w PTO

  • Lead the development and implementation of secure cloud environments using CI/CD and Infrastructure-as-Code.
  • Build automated deployments, perform maintenance, and assist in production support and incident resolution.
  • Mentor team members and drive delivery to meet customer requirements in a collaborative, no-blame culture.

Aker Systems builds and operates secure, high-performance cloud-based data infrastructure for enterprises. The company prioritizes an inclusive culture and has been recognized for growth and innovation, winning awards in 2020 and 2024.

  • Provide solutions to customers to make them successful using our products.
  • Troubleshoot customer environments and engage in active triaging with customers.
  • Participate in on-call rotation for weekend coverage.

Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life with its unified DataOps platform, Astro, powered by Apache Airflow. Trusted by more than 800 enterprises, the company fosters a diverse and inclusive culture as an equal opportunity employer.

  • Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
  • Define and drive SRE platform strategy, incident management, and observability engineering.
  • Mentor team members, foster collaboration, and ensure operational excellence.

XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.

  • Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
  • Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
  • Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.

Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.

$0–$150,000/yr

  • Help design, build, and operate the Kubernetes platform used across PulsePoint.
  • Own reliability, observability, and incident response across platform services.
  • Build infrastructure automation and GitOps workflows to reduce operational toil.

PulsePoint sits at the intersection of healthcare and adtech, helping brands interpret health signals using real-world data. With over 300 employees, the company is a post-acquisition profitable leader in the US healthcare ad market, known for a flat hierarchy and high engineering bar.

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

  • Lead the SRE strategy and execution for a high-growth AI company.
  • Build and scale a high-performing SRE team while defining reliability standards.
  • Architect secure, scalable cloud infrastructure and implement observability practices.

This company develops advanced AI products and agentic technology. It operates in a high-growth, international environment with a focus on operational excellence and innovation.

  • Build and configure GKE clusters and Google Cloud project environments using Terraform.
  • Implement and maintain CI/CD pipelines and environment promotion workflows.
  • Configure and maintain Google Cloud-native monitoring, alerting, and logging.

We design, build, and scale AI-powered solutions that create real business impact. We are building a high-performance culture grounded in five values: Empowering Excellence, Collaborative Teamwork, Unsolicited Respect, Consistent Transparency, and Efficient Communication.