Source Job

US Unlimited PTO

  • Building and coaching a high-performing distributed team with a shared operating model.
  • Owning platform capabilities for provisioning, deployment, and operations of infrastructure.
  • Leading infrastructure migration towards a modern SaaS model with incremental delivery.

AWS Terraform Site Reliability Engineering FinOps Leadership

20 jobs similar to Head of Cloud Platform Engineering

Jobs ranked by similarity.

Europe 4w PTO

  • Take full ownership of the company's platform and infrastructure function, defining and executing the platform engineering strategy in partnership with the CTO.
  • Lead, mentor, and develop an established team of three DevOps Engineers and one SRE, setting clear responsibilities and development plans.
  • Own platform-related budgeting, cloud spend optimization, and technology decisions to improve reliability, security, and developer productivity.

Our client is a remote-first digital product company building and scaling SaaS products, AI-powered solutions, and web and mobile applications for global markets. The company operates across more than 20 countries with approximately 40–50 engineers supporting a portfolio of around 15 digital products.

$100,000–$145,000/yr
US

  • Own and operate a production agentic AI platform on AWS, ensuring reliability and scaling.
  • Lead infrastructure automation and release management, driving best practices in security and compliance.
  • Collaborate with the platform team on agile ceremonies and proactively communicate status to stakeholders.

Inizio Evoke is a healthcare communications company dedicated to making health more human. As part of the larger Inizio network, it emphasizes a collaborative, inclusive culture where employees are encouraged to be their authentic selves.

$140,000–$175,000/yr
US

  • Lead and grow a team of platform engineers, coaching them on infrastructure and cloud challenges.
  • Drive the platform roadmap, balancing reliability, cost, security, and developer experience with AWS and Kubernetes.
  • Partner cross-functionally to align platform priorities with business goals and ensure system reliability.

PerfectServe is a leading provider of clinical communication and physician scheduling solutions in the health IT space. The company has 400+ employees and 30,000+ customers, with over $100 million in annual revenue, and has received multiple Best in KLAS awards.

UK

  • Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
  • Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
  • Drive AI-specific observability, FinOps, and security practices across the platform.

We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.

US

  • Architect and implement infrastructure for seamless migrations of critical systems, keeping uptime and reliability front and center.
  • Design, build, and maintain the tooling and processes that let our SaaS products scale and run themselves on AWS.
  • Partner closely with developers to strengthen the reliability, performance, security, and scalability of our application architecture.

Duetto is the hospitality industry's leading revenue management platform, providing a suite of tools for hotels, resorts, and casinos. Backed by GrowthCurve Capital since 2024, the company has been named the #1 Best Place to Work in Hotel Tech in 2025 and is passionate about the industry it serves.

Global

  • Design and maintain AWS infrastructure using Terraform, with a focus on scalability cost and PCI-scoped network segmentation
  • Build and evolve the observability stack and CI/CD pipelines to ensure smooth production operations and rapid deployment
  • Lead incident response define SLOs and run performance tests to optimize payment-critical services

Xplor Technologies provides vertical software, embedded payments, and AI tools for membership-based and service-based industries. With over 130,000 businesses in 72+ countries and processing $47 billion in payments annually, the company values diversity, collaboration, and a people-first culture.

$150,000–$250,000/yr
US Europe Singapore

  • Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
  • Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
  • Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.

A fast-growing AI/ML platform startup building infrastructure for training, evaluating, and aligning AI models within reinforcement learning environments. The engineering team of ~15 includes competitive programming medalists, serial AI startup founders, and researchers published at top venues.

North America Canada Latin America

  • Design and operate scalable cloud infrastructure across AWS and GCP.
  • Build and improve Kubernetes, Linux, and cloud networking environments.
  • Strengthen security, disaster recovery, and platform resilience.

Hubstaff provides workforce analytics and time tracking for remote teams, serving over 200,000 global users. The company is a product-led organization with a winning culture and a fully remote team of experienced engineers.

Argentina

  • Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
  • Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
  • Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.

Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.

Europe 4w PTO

  • Lead the platform engineering function, defining strategy and long-term vision for cloud infrastructure and developer experience.
  • Partner with senior engineering leadership to establish a platform roadmap focused on reliability, scalability, security, and automation.
  • Mentor a team of platform engineers, support hiring, and drive operational maturity across global engineering teams.

The company is a fast-growing digital product organization that develops and operates a portfolio of SaaS products. It employs dozens of engineers and fosters a collaborative, remote-first culture focused on ownership and continuous improvement.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

US

  • Design, implement, and maintain AWS cloud infrastructure using Terraform.
  • Build and optimize CI/CD pipelines to enable rapid, safe deployments across multiple environments.
  • Own observability strategy with comprehensive monitoring, logging, and alerting systems using Datadog.

Avantos is building an AI-native operating system for financial services, transforming fragmented data into a single intelligent system. They are a product-led, fast-moving team at the intersection of AI, fintech, and modern infrastructure.

Brazil

  • Architect, deliver, and maintain critical cloud platform components on AWS EKS, focusing on production reliability and observability.
  • Establish SRE standards including SLOs, error budgets, and automated tooling to reduce operational friction.
  • Provide technical advisory through code reviews and architecture recommendations to maintain high platform standards.

Inflect is a US-based advisory and marketplace revolutionizing digital infrastructure procurement. They are a small team focused on reducing friction in buying datacenter, cloud, and network services through automation and better deal terms.

US Unlimited PTO

  • Provide hands-on engineering leadership for cloud migration and modernization initiatives in high-transaction payments environments.
  • Troubleshoot complex networking and infrastructure issues across legacy and AWS environments.
  • Define technical deliverables, communicate risks, and drive work through completion autonomously.

EverOps is an Embedded Service Provider that partners with customer engineering teams to address mission-critical infrastructure, cloud, and delivery challenges. The company has been remote since day one and offers equity, 401k, and healthcare.

$120,000–$155,000/yr
Global

  • Own infrastructure as code across development, staging, and production environments
  • Build, maintain, and improve CI/CD pipelines for reliable and efficient deployments
  • Manage cloud infrastructure, establish scalable engineering practices, and lead incident response

CelebriOS is a software company building B2B SaaS products that help businesses make better decisions and streamline operations. The company has a remote-first working environment and a benefits package designed to support their team.

$190,000–$225,000/yr
US

  • Design, build, and operate core cloud infrastructure on AWS, including compute, networking, and container orchestration.
  • Own the CI/CD platform used across engineering teams, including build pipelines, environment promotion, and progressive rollout.
  • Build and maintain the observability stack across the organization, including logging, metrics, distributed tracing, and alerting.

RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company is an Equal Opportunity and Affirmative Action employer committed to diversity and collaboration.

Global

  • Architect, build, and operate secure, multi-account AWS environments using modern Infrastructure as Code.
  • Design and optimize container orchestration with AWS ECS (Fargate) and Kubernetes (EKS) for specialized workloads.
  • Build unified CI/CD pipelines with GitHub Actions, embed security by design, and establish observability with Prometheus, Grafana, and CloudWatch.

Berlitz is a global language education company with a history of nearly 150 years, currently undergoing a digital transformation. The company fosters a remote-first, AI-native culture with a focus on autonomy and greenfield development, though employee count is not specified.

US

  • Design, maintain, and support secure AWS environments across compute, storage, networking, account structure, and operational practices.
  • Administer Azure-hosted applications, data services, and supporting infrastructure including tenant configuration and access controls.
  • Implement and maintain infrastructure as code using tools such as Terraform or CloudFormation to improve consistency and reliability.

JerseySTEM is a mission-driven professional network of pro-bono contributors dedicated to improving access to STEM education and career pathways for underserved middle school girls in New Jersey. Members contribute their professional skills and leverage their networks in service of the organization's gender-equity agenda, operating remotely with a small volunteer base.

Europe US

  • Design, build, and operate multi-region AWS infrastructure on Kuberneties with Terraform and Helm at 15PB+ scale.
  • Own high-availablity, event-driven architectures and cost optimization across the stack.
  • Drive developer experience, security, and incident response as the second platform team member.

ScorePlay is the AI-powered media infrastructure for sports, automating content operations for the world's biggest sports organizations. We are a 50-person remote-first team based in New York and Paris, growing 2x year over year with 98% retention.

US

  • Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
  • Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
  • Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.

ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.