Source Job

UK

  • Lead Cloud Platform and SRE teams, driving infrastructure strategy and ownership including Kubernetes, GCP, and Terraform.
  • Champion SRE culture, define SLOs, SLAs, and enhance observability and incident management.
  • Own the developer-facing platform as a product, ensuring self-service infrastructure and security compliance (SOC2, ISO-27001).

GCP Kubernetes Terraform Python SRE

20 jobs similar to Platform Engineering Manager

Jobs ranked by similarity.

Global

  • Lead a high-leverage remote team of four infrastructure engineers, driving the evolution toward a scalable zero-toil platform.
  • Guide the team through an AI-driven engineering approach to reduce manual work and achieve zero-touch, scalable infrastructure.
  • Prepare and execute the strategy for CI/CD and artifact distribution systems to scale during a quality surge without increasing engineering toil.

Camunda is the enterprise platform for agentic orchestration, enabling organizations to coordinate AI agents, people, and systems across complex business processes. Trusted by over 700 organizations worldwide, including 9 of top 10 US banks, Camunda is a fully remote and global company with 150+ engineers across 20+ teams, and is transforming into an AI-first organization.

$100,000–$122,222/yr
Canada 8w PTO

  • End-to-end ownership of internal orchestration platform built on event-driven architecture with Redpanda, including code, architecture, and roadmap.
  • Own infrastructure-as-code using Terraform Cloud, manage Kubernetes workloads with Helm, and provide self-service tooling for engineering teams.
  • Set SLOs, handle production on-call, lead incident response, author design docs, and operate AI-natively using tools like Cursor and Notion AI.

Velora unifies Aplos, Raisely, and Keela into one company with a shared mission to help nonprofit organizations thrive by offering fundraising, donor management, financial tracking, and communications tools. We are a financially solid company with a combined team dedicated to making nonprofit work easier, more impactful, and more sustainable.

US

  • Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
  • Use AI agents as force multipliers to automate manual processes and improve developer experience.
  • Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.

Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.

Global

  • Support engineering teams by developing resilient applications on GKE and advising on best practices.
  • Develop and maintain infrastructure-as-code using Terraform, ArgoCD, and Python.
  • Manage GKE environments across multiple regions, including IAM and identity management.

Mimica uses AI-powered task mining to observe employee actions and create process maps, helping enterprises improve efficiency. The company is a fast-growing scale-up with a lean, collaborative culture.

$185,000–$280,000/yr
US 4w PTO

  • Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
  • Scale single-tenant deployments and build observability, incident response, and compliance practices.
  • Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.

Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.

$160,000–$190,000/yr
US

  • Design, implement, and maintain reliable, scalable infrastructure, applications, and tooling for software-defined manufacturing.
  • Write clean, maintainable code and perform peer code reviews to ensure high-quality deliverables.
  • Collaborate with cross-team members to prototype new technology and evaluate technical feasibility.

Bright Machines is an innovator in software-defined manufacturing, using intelligent automation to transform factories. The company is building a team of experts to create a new manufacturing category, offering a culture of innovation and impact.

26w maternity 26w paternity

  • Lead and grow a high-performing engineering team through hiring, coaching, and career development.
  • Drive the team's direction and execution rhythm, balancing roadmap delivery and technical debt.
  • Partner with product and engineering stakeholders to identify leverage points for platform improvements.

Bloomreach is building the world's premier agentic platform for personalization. They power personalization for more than 1,400 global brands, including American Eagle, Sonepar, and Pandora.

Europe 6w PTO

  • Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
  • Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
  • Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.

Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.

$120,000–$165,000/yr
US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

Global Unlimited PTO

  • Manage the lifecycle of internal Platform-as-a-Service and Data-as-a-Service products, overseeing highly technical engineering teams.
  • Define error budgets for 99.99% availability, lead cloud scaling strategies, and own the GCP cloud computing budget.
  • Partner with Data Engineers to optimize data pipelines and architecture, and ensure compliance with privacy regulations.

Sardine is the leading agentic risk platform for fighting financial crime, with an integrated solution unifying data across risk teams. The company has hubs in multiple locations but maintains a remote-first culture, hiring self-motivated individuals with extreme ownership and a high growth orientation.

US

  • Lead a Dedicated Tenant Site Reliability Engineering organization, driving complex initiatives and operational excellence across multiple teams.
  • Oversee delivery and operation of PingOne Advanced Identity Cloud and Advanced Services, improving consistency and reliability.
  • Partner with SRE, Security, and Development teams to manage dependencies and evolve software delivery strategies.

Ping Identity provides an intelligent cloud identity platform that secures and streamlines digital experiences. Headquartered in Denver, Colorado, the company serves more than half of the Fortune 100 and fosters a culture that champions individuality and digital freedom.

Japan APAC

  • Design and evolve cloud architecture on GCP using Terraform and GitOps.
  • Build CI/CD pipelines for IaC with Policy-as-Code and drift detection.
  • Strengthen observability stack with Prometheus, Grafana, Loki, and Tempo.

Alpaca is a US-headquartered self-clearing broker-dealer and brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, and 24/5 trading. With over $320 million in total investment and a diverse global team of 380+ members spanning multiple countries, we are committed to opening financial services to everyone on the planet.

US

  • Collaborate with cross-functional teams to design, implement, and maintain scalable and reliable infrastructure.
  • Utilize Infrastructure as Code (IaC) principles to automate provisioning, configuration, and deployment processes.
  • Troubleshoot and resolve complex technical issues related to infrastructure and deployment.

Pano AI is the leader in early wildfire detection and intelligence, helping fire professionals respond to fires faster and more safely using a combination of advanced hardware, software, and AI. The company is a 175+ person growth-stage hybrid-remote startup headquartered in San Francisco, recognized as one of the most innovative AI companies by Fast Company and TIME.

Costa Rica

  • Ensure reliability, performance, and scalability of Backcountry's multi-cloud platform.
  • Drive incident resolution, postmortems, and automation to reduce operational toil.
  • Leverage AI-assisted engineering tools and collaborate with teams to build and maintain observability and SLI/SLO instrumentation.

Backcountry is an online retailer of outdoor gear and apparel, rooted in adventure and the outdoor lifestyle. The company fosters a culture of recognition, wellbeing, and connection, with a lean, fast-paced engineering team.

Global Unlimited PTO

  • Lead a high-impact infrastructure team, evolving internal platforms and CI/CD systems to support large-scale engineering operations.
  • Drive automation initiatives and AI-driven practices to reduce operational complexity and improve developer experience.
  • Define and execute strategies for scalable infrastructure, cloud environments, and platform engineering.

The partner company is a technology organization focused on building infrastructure platforms that enable engineering teams to deliver software faster. It is a remote-first company with a collaborative culture and a focus on innovation and scalability.

Europe

  • Build and operate the infrastructure behind AI-powered products, improving reliability, security, scalability, and cost efficiency.
  • Write code, automate infrastructure, investigate production issues, and design systems that reduce operational complexity.
  • Take ownership of unfamiliar systems, identify highest-leverage improvements, and balance immediate production needs with long-term platform investments.

Zencoder builds and orchestrates AI agents that ship real work across code, research, and operations. It is a growing platform where people and agents collaborate, with a high-caliber team and a culture that values individual contributors.

Global

  • Embed with product and platform teams from early stages to ensure reliability is designed in from the start.
  • Define production-readiness standards and measurable SLIs/SLOs to guide operational excellence.
  • Build tooling and infrastructure across AWS, GCP, and Azure using Terraform, and share on-call rotation.

We build WebContainers and Bolt.new, an AI-powered app builder that lets you create, edit, and deploy full-stack apps instantly in your browser. We are a fully remote, globally distributed team of passionate engineers serving over 1 million developers monthly.

$160,000–$180,000/yr
US

  • Own the infrastructure and platform powering the marketplace, focusing on reliability, observability, security, and automation.
  • Manage production AWS and EKS clusters, infrastructure as code with Terraform and GitOps, and CI/CD pipelines via GitHub Actions.
  • Build automation and internal tooling in Python, Bash, Go, and Node.js/TypeScript, and operate PostgreSQL, MongoDB, and Temporal.

Office Hours is an on-demand expert network that connects leading organizations with trusted experts across various knowledge domains. The company is hyper-growth, profitable, and expanding quickly, backed by top marketplace investors.

UK

  • Work with clients to understand requirements and develop cloud-native solutions using Google Cloud technologies.
  • Provide technical oversight and end-to-end delivery of projects as a lead architect.
  • Mentor team members and act as a trusted advisor to clients.

Zencore is a cloud consulting company founded by former Google Cloud leaders. They are a fast-growing company with a collaborative and inclusive culture.