Source Job

$56,924–$61,801/yr
Europe

  • Acts as the strategic bridge between Cloud Operations, Product Management, Engineering, and other teams to drive service quality and operational excellence.
  • Drives large-scale transformation programs, promotes operational best practices, and ensures lessons learned translate into portfolio-wide improvements.
  • Champions automation, observability, and reliability standards while influencing engineering practices and product roadmaps.

Cloud Operations Site Reliability Engineering DevOps Program Management

20 jobs similar to Cloud Operations Service Strategy Lead

Jobs ranked by similarity.

EMEA

  • Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
  • Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
  • Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.

ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.

UK

  • Lead Cloud Platform and SRE teams to scale securely and efficiently.
  • Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
  • Champion SRE culture with SLOs, error budgets, and observability.

Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.

US Unlimited PTO

  • Building and coaching a high-performing distributed team with a shared operating model.
  • Owning platform capabilities for provisioning, deployment, and operations of infrastructure.
  • Leading infrastructure migration towards a modern SaaS model with incremental delivery.

Totara is a global learning platform trusted by more than 1,500 organisations and 21 million users worldwide, offering flexible learning, compliance, and talent development solutions. With a distributed team across New Zealand, Australia, the UK, and the US, the company values diverse perspectives and offers flexible, hybrid working.

$114,700–$195,000/yr
North America

  • Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
  • Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
  • Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.

Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.

$180,000–$220,000/yr
North America

  • Lead and modernize Sectigo's global infrastructure organization with a focus on reliability and operational maturity.
  • Develop a measurable operating model using SLAs, SLOs, and key metrics to drive improvement.
  • Drive automation, AI-enabled operations, and closer collaboration with engineering teams.

Sectigo is an innovative provider of certificate lifecycle management (CLM) solutions, helping large brands simplify digital trust. With over 700,000 customers including 65% of the Fortune 500, they emphasize a culture of support, excellence, and teamwork.

India

  • Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
  • Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
  • Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.

Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.

US

  • Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
  • Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
  • Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.

Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.

Poland

  • Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
  • Define and drive SRE platform strategy, incident management, and observability engineering.
  • Mentor team members, foster collaboration, and ensure operational excellence.

XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.

US

  • Lead the transformation of a diverse operations-heavy organization into a modern, AI-first Production Engineering function.
  • Own end-to-end reliability, performance, scalability, and security of NICE's global cloud, telecom, and datacenter platforms.
  • Drive adoption of software-first operational practices including automated recovery, infrastructure as code, and observability.

NICE provides software products used by 25,000+ global businesses to deliver extraordinary customer experiences, fight financial crime, and ensure public safety. With over 8,500 employees across 30+ countries, the company fosters a culture of ambition, game-changing innovation, and high standards.

$150,000–$185,000/yr
US Unlimited PTO

  • You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
  • You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
  • You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.

Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.

$120,000–$155,000/yr
Global

  • Own infrastructure as code across development, staging, and production environments
  • Build, maintain, and improve CI/CD pipelines for reliable and efficient deployments
  • Manage cloud infrastructure, establish scalable engineering practices, and lead incident response

CelebriOS is a software company building B2B SaaS products that help businesses make better decisions and streamline operations. The company has a remote-first working environment and a benefits package designed to support their team.

$185,940–$227,260/yr
US 12w maternity 12w paternity

  • Own cloud infrastructure and Kubernetes environment, keeping it reliable, secure, and cost efficient.
  • Lead the team in using AI-assisted engineering to design, build, and operate platform infrastructure.
  • Manage and mentor a team of engineers while staying hands-on to contribute directly to the work.

Doma Technology provides solutions for lenders, real estate professionals, title agents, and homeowners that make closings simpler and more efficient. The company values an entrepreneurial, people-first culture with a focus on diversity, equity, and inclusion.

India

  • Own every engagement end-to-end as Engagement Lead, from kickoff to executive readout, serving as primary customer contact.
  • Lead a small pod of engineers, setting goals and coaching their growth to ensure strong team performance.
  • Drive technical discovery and analysis across multi-cloud platforms, platform engineering, and AI-driven solutions.

AHEAD builds platforms for digital business, helping enterprises deliver on digital transformation through cloud infrastructure, automation, and analytics. The company prioritizes a culture of belonging and is an equal opportunity employer.

Europe

  • Lead the SRE strategy and execution for a high-growth AI company.
  • Build and scale a high-performing SRE team while defining reliability standards.
  • Architect secure, scalable cloud infrastructure and implement observability practices.

This company develops advanced AI products and agentic technology. It operates in a high-growth, international environment with a focus on operational excellence and innovation.

Global

  • Design, provision, and maintain Azure cloud infrastructure using infrastructure-as-code.
  • Build and improve CI/CD pipelines to enable fast, reliable software delivery.
  • Implement observability practices and support incident response to ensure system reliability.

INNERGY transforms the woodworking industry with cloud-based ERP software for custom manufacturers. Founded in 2016, we are a globally distributed team of 200+ professionals united by deep expertise and a passion for solving real-world problems.

$145,000–$177,000/yr
US

  • Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
  • Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, and networking.
  • Design and implement solutions that improve availability, scalability, performance, and resilience of the platform.

Everbridge empowers enterprises and government organizations to anticipate, mitigate, respond to, and recover from critical events. The company focuses on building resilient systems and fostering a culture of ownership, continuous improvement, and operational excellence.

Spain 5w PTO

  • Define SLIs, SLOs, and reliability targets for the platform.
  • Improve observability, alerting, and production readiness across services.
  • Automate operational work and support cloud/Kubernetes infrastructure.

Lodgify is a fast-growing scale-up in vacation rental technology, backed by $30M in funding. Headquartered in Barcelona, the 380+ person team of 60+ nationalities is passionate about transforming short-term rentals.

$59,400–$65,880/yr
Europe

  • Lead the design, implementation, and ongoing improvement of reliable, scalable, and secure production platforms and services.
  • Work closely with cross-functional teams to build and maintain resilient infrastructure and deployment patterns.
  • Provide technical leadership and mentorship, promoting strong engineering standards and operational best practices.

Cision is a global leader in PR, marketing and social media management technology and intelligence, helping brands connect with customers and stakeholders. They have offices in 24 countries, a network of over 1.1 billion influencers, and a culture that champions diversity, equity, and inclusion.

$191,000–$226,000/yr
US Unlimited PTO

  • Own the reliability, performance, and resilience of cloud environments (AWS, Kubernetes) and define SLOs across critical services.
  • Lead incident response, on-call rotation, and drive root cause analysis to ensure high production quality.
  • Build and maintain observability systems and automate operational toil using AI tools.

Garner partners with employers to redesign healthcare by using clinical metrics to identify top doctors and incentivize members to better care. The company has helped over 2.5 million people, saved $1B in costs, and doubled annually for five years, fostering a mission-driven, high-performance culture.

Latin America

  • Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.

We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.