Source Job

$185,000–$200,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Coordinate with technical and non-technical staff across departments, including workflow automation that bridges infrastructure and business processes.
  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure in GCP.
  • Maintain incident response process and tooling, and build automation that reduces toil and enables self-healing infrastructure.

GCP Kubernetes Docker Python Terraform

20 jobs similar to Staff Cloud Operations Engineer

Jobs ranked by similarity.

France

  • Evolve an Internal Developer Platform enabling development teams to deploy and operate applications securely with high self-service.
  • Play a crucial role in infrastructure architecture in a multi-cloud environment with a focus on GCP.
  • Build and maintain the platform used by over 50 internal clients, then support their concrete use by teams.

Lifen believes medical data can transform healthcare by reducing administrative burden, improving care coordination, and accelerating scientific discovery. Since 2015, the company has connected 800 hospitals and 150,000 healthcare professionals, with over 150 employees working remotely and from offices to unlock the potential of health data.

Global

  • Support management and maintenance of Google Cloud Platform infrastructure.
  • Assist in ensuring reliability, scalability, and performance of cloud services.
  • Contribute to GitLab CI/CD pipelines, work with Kubernetes, and write automation scripts.

Miratech is a global IT services and consulting company that helps visionaries change the world through digital transformation. With nearly 1,000 professionals across 25 countries, the company maintains a culture of Relentless Performance with a 99% project success rate and over 30% year-over-year growth.

$190,000–$230,000/yr
US

  • Design and implement scalable cloud infrastructure using Kubernetes, Pub/Sub, and distributed systems technologies.
  • Collaborate with our AI team to optimize data pipelines and integrate AI to remove performance bottlenecks.
  • Drive platform reliability initiatives including alerting, health checking, and incident management.

Syllo is building a unified litigation platform that helps lawyers and paralegals use AI throughout the litigation life cycle. We are a quickly expanding company with enterprise customers including major law firms and corporations.

$175,000–$185,000/yr
US

  • Consolidate Terraform and establish conventions for state management, modules, and CI checks.
  • Improve monitoring, observability, and automation in Datadog and Cloud Monitoring.
  • Right-size workloads, evaluate Kubernetes architecture, and retire legacy tooling.

Vida is a virtual, personalized obesity care provider that combines evidence-based treatment with advanced technology to help patients improve their health. Trusted by Fortune 100 companies and growing for years, Vida takes a whole-person approach to care and celebrates diversity across its team.

$135,000–$150,000/yr
US

  • Operate, scale, and troubleshoot Bitsight's SaaS cloud infrastructure with focus on reliability, efficiency, and security.
  • Tackle complex system-level designs and proactively anticipate performance and scalability issues.
  • Pioneer self-optimizing infrastructure systems using AI, ensuring manual and staging validation before production deployment.

Bitsight is a cyber risk management leader transforming how companies manage exposure, performance, and risk. Over 3,500 customers and 600 teammates work across Boston, Raleigh, New York, Lisbon, Singapore, and remote locations.

$143,200–$243,400/yr
North America

  • Contribute to infrastructure automation and operational resilience across hybrid cloud and data center operations.
  • Implement closed-loop auto-remediation systems and SRE tooling to reduce manual intervention and incident resolution time.
  • Develop and maintain SLO frameworks, alerting policies, and Infrastructure-as-Code pipelines for reproducible deployments.

ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter, faster, and better. They foster an AI-native culture where technology and talent are unstoppable together.

India

  • Lead the architecture and implementation of complex cloud solutions across AWS and GCP.
  • Drive cloud automation and optimization initiatives to improve scalability and reliability.
  • Provide technical leadership and mentorship to engineers while collaborating with global teams.

The company focuses on cloud infrastructure and platform engineering. They operate with global teams and emphasize automation, security, and reliability.

$165,000–$200,000/yr
US

  • Design, build, and maintain cloud infrastructure on GCP and AWS using Terraform, optimizing CI/CD pipelines for rapid deployments.
  • Implement comprehensive observability including monitoring, logging, alerting, and distributed tracing to ensure platform health.
  • Establish and enforce security best practices, support AI/ML infrastructure, and build developer experience tooling.

Re:Build operates an advanced, end-to-end manufacturing platform that partners with industrial companies to bring products from concept to full-scale production. The company is guided by The Re:Build Way principles and aims to revitalize America's manufacturing base, creating meaningful jobs across the country.

Canada

  • Design and evolve highly available cloud architecture on Google Cloud using Terraform and GitOps.
  • Build and maintain secure CI/CD pipelines for Infrastructure-as-Code and develop self-service developer platforms.
  • Strengthen platform observability and apply SRE principles to improve reliability and operational maturity.

They are a financial technology company that provides production-critical infrastructure. They have a globally distributed team and a culture of autonomy and async-first collaboration.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

Brazil

  • Support implementation, automation, migration, and management of cloud environments, primarily on GCP.
  • Work with Infrastructure as Code, Kubernetes, Docker, monitoring, backup, and cloud-management technologies.
  • Collaborate with internal specialists and customers throughout the project lifecycle, from implementation to ongoing support.

The company specializes in cloud implementation, infrastructure modernization, and DevOps services, with a primary focus on Google Cloud Platform. It fosters a collaborative environment focused on technology, professional development, and continuous improvement.

$145,000–$145,000/yr
US

  • Design, deploy, and maintain GCP infrastructure, including Compute Engine, GKE, Cloud Storage, IAM, and Cloud Interconnect, following well-architected principles.
  • Automate provisioning using Terraform, implement IAM best practices, and partner on security audits and hardening.
  • Monitor performance, availability, and cost, and support cross-cloud and on-premises networking as needed.

Dragos is the global leader in xOT cybersecurity, protecting critical infrastructure systems that deliver water, power, and healthcare. The remote-first team spans North America, Europe, the Middle East, and APAC, built on authenticity, transparency, and trust.

$220,000–$292,000/yr
US Unlimited PTO

  • Own the platform including GCP, Kubernetes, Temporal, GPU fleet, and deploy/rollback machinery.
  • Contribute to AI enablement substrate: GPU capacity, training/inference pipelines, and cost optimization.
  • Strengthen team practices through tooling, standards, tests, observability, and release processes.

Descript is building a simple, intuitive, fully-powered editing tool for video and audio — an editing tool built for the age of AI. They are a team of 150 backed by top investors like OpenAI and Andreessen Horowitz, with a culture that values collaboration and serendipitous discovery.

Brazil 4w PTO

  • Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
  • Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
  • Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.

Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.

  • Own the reliability, performance, and scalability of Runlayer's infrastructure across AWS and GCP.
  • Manage Kubernetes clusters, database reliability, and CI/CD pipelines for rapid deployments.
  • Lead incident response and partner with product engineers to design resilient systems for enterprise customers.

Runlayer builds a unified platform for MCPs, Skills, and AI Agents, providing enterprises with security, governance, and observability to deploy AI safely and at scale. Founded by engineers who built AI Actions for OpenAI and Zapier Agents, the team has raised $42M from Felicis and Khosla Ventures, serving companies like Gusto, Instacart, and Opendoor.

$59,400–$65,880/yr
Europe

  • Lead the design, implementation, and ongoing improvement of reliable, scalable, and secure production platforms and services.
  • Work closely with cross-functional teams to build and maintain resilient infrastructure and deployment patterns.
  • Provide technical leadership and mentorship, promoting strong engineering standards and operational best practices.

Cision is a global leader in PR, marketing and social media management technology and intelligence, helping brands connect with customers and stakeholders. They have offices in 24 countries, a network of over 1.1 billion influencers, and a culture that champions diversity, equity, and inclusion.

$191,000–$226,000/yr
US Unlimited PTO

  • Own the reliability, performance, and resilience of cloud environments (AWS, Kubernetes) and define SLOs across critical services.
  • Lead incident response, on-call rotation, and drive root cause analysis to ensure high production quality.
  • Build and maintain observability systems and automate operational toil using AI tools.

Garner partners with employers to redesign healthcare by using clinical metrics to identify top doctors and incentivize members to better care. The company has helped over 2.5 million people, saved $1B in costs, and doubled annually for five years, fostering a mission-driven, high-performance culture.

US Unlimited PTO

  • Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
  • Design and maintain infrastructure as code across multiple cloud providers.
  • Provide technical leadership and mentorship across the Systems Engineering team.

Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.

US

  • Architect and build a FedRAMP High-compliant GCP landing zone supporting a multi-department enterprise migration.
  • Establish repeatable, automated IaC pipelines and enforce least-privilege IAM models integrated with external identity providers.
  • Deliver end-to-end network security, centralized API gateway, and audit-ready logging integration with client SIEM platforms.

OnTrac is a cloud infrastructure consulting firm specializing in Google Cloud Platform solutions and FedRAMP-compliant enterprise migrations. The company operates with a focus on security and compliance, working with a team of experienced engineers on high-stakes government and enterprise projects.

$180,000–$220,000/yr
US

  • Design, implement, and maintain reliable, scalable, and secure infrastructure to support applications and automation systems.
  • Automate infrastructure provisioning, configuration management, and deployment pipelines using tools like Terraform and ArgoCD.
  • Implement observability solutions and enforce security best practices to ensure uptime and system performance.

Bright Machines is a next-generation, AI-enabled manufacturer focused on data center infrastructure production, using proprietary AI-based robotics and software to assemble hardware products for hyperscalers and OEMs. The company is headquartered in San Francisco, California, with an integration center in Guadalajara, Mexico, and has been recognized by Forbes' AI 50 and other leading organizations.