Source Job

Canada

  • Support production systems across Azure, GCP, and datacenters to meet SLA targets.
  • Automate with Terraform, GitHub Actions, ArgoCD, and Python/Bash/PowerShell.
  • Participate in on-call rotation and collaborate to resolve incidents and improve reliability.

Terraform GitHub Actions Python Kubernetes Azure

20 jobs similar to Cloud Engineer 2, Site Reliability Engineering

Jobs ranked by similarity.

$175,000–$185,000/yr
US

  • Consolidate Terraform and establish conventions for state management, modules, and CI checks.
  • Improve monitoring, observability, and automation in Datadog and Cloud Monitoring.
  • Right-size workloads, evaluate Kubernetes architecture, and retire legacy tooling.

Vida is a virtual, personalized obesity care provider that combines evidence-based treatment with advanced technology to help patients improve their health. Trusted by Fortune 100 companies and growing for years, Vida takes a whole-person approach to care and celebrates diversity across its team.

UK

  • Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
  • Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.

India

  • Design, deploy, and maintain the reliability, availability, and performance of critical systems and APIs across AWS and GCP.
  • Build observability frameworks, define SLIs/SLOs, and implement monitoring using Datadog and Kubernetes.
  • Participate in on-call rotations, incident response, and blameless post-incident reviews to drive systemic improvements.

JumpCloud is an AI-powered unified IT management platform that secures the modern workforce by consolidating identity, device, and access management. The company is remote-first with teams in over 15 countries and values building connections, thinking big, and continuous improvement.

$81,000–$88,000/yr
Canada Unlimited PTO

  • Design and implement scalable infrastructure using IaC and automation.
  • Build and maintain robust CI/CD pipelines for frictionless releases.
  • Monitor system performance and lead incident response efforts.

Practice Better is an all-in-one platform for health and wellness practitioners to manage their businesses and client care. Founded in 2016, the remote-first company now serves tens of thousands of practitioners across 70+ countries and has a team of curious, driven, and empathetic people.

$180,000–$220,000/yr
US

  • Design, implement, and maintain reliable, scalable, and secure infrastructure to support applications and automation systems.
  • Automate infrastructure provisioning, configuration management, and deployment pipelines using tools like Terraform and ArgoCD.
  • Implement observability solutions and enforce security best practices to ensure uptime and system performance.

Bright Machines is a next-generation, AI-enabled manufacturer focused on data center infrastructure production, using proprietary AI-based robotics and software to assemble hardware products for hyperscalers and OEMs. The company is headquartered in San Francisco, California, with an integration center in Guadalajara, Mexico, and has been recognized by Forbes' AI 50 and other leading organizations.

India

  • Architect and scale multi-region microservices, APIs, and authentication infrastructure on AWS/GCP.
  • Lead SLOs, observability, incident management, and disaster recovery automation to maintain 99.99% availability.
  • Manage Kubernetes clusters and Terraform IaC while eliminating toil with Python/Go tooling.

JumpCloud is an AI-powered unified IT management platform that secures the modern workforce through identity, device, and access management. The company is remote-first with teams in 15+ countries and values building connections, thinking big, and continuous improvement.

Canada

  • Design, develop, and maintain cloud infrastructure using Terraform, Kubernetes, and CI/CD pipelines in an IaC manner.
  • Improve monitoring and logging systems like Prometheus, Grafana, Fluentbit, and Loki to ensure high availability and performance.
  • Collaborate closely with cross-functional teams across APAC and EMEA to automate operations and implement disaster recovery solutions.

Crypto.com is a global cryptocurrency platform founded in 2016, serving more than 150 million customers with the vision of Cryptocurrency in Every Wallet. The company offers an empowered, transformational work environment with a talented, ambitious, and supportive team.

US

  • Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil and improve platform reliability.
  • Build observability and monitoring solutions with Grafana and AWS CloudWatch, and implement Infrastructure as Code using Terraform.
  • Develop CI/CD pipelines, define SRE standards and metrics, and participate in incident management and on-call rotation.

Peraton is a next-generation national security company that delivers mission-critical solutions and transformative IT services to government agencies and the U.S. armed forces. The company operates across land, sea, space, air, and cyberspace, with employees solving the most daunting challenges facing customers worldwide.

Canada USA Unlimited PTO

  • Own the observability, logging and alerting for Kubernetes clusters and critical workloads.
  • Build and maintain automation for lifecycle management of Kubernetes clusters.
  • Identify and root-fix reliability bottlenecks before they become incidents.

Wrapbook is an AI platform for production finance, built for feature films and TV, trusted by Netflix and Paramount. Backed by top investors, our team of over 350 employees uses AI to transform how finance teams work.

$185,000–$200,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Coordinate with technical and non-technical staff across departments, including workflow automation that bridges infrastructure and business processes.
  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure in GCP.
  • Maintain incident response process and tooling, and build automation that reduces toil and enables self-healing infrastructure.

Branch empowers workers with financial freedom by helping companies accelerate payments and providing accessible, free financial services. It is a remote-first, award-winning FinTech with employees across the U.S., fostering a culture of transparency, accountability, and trust.

  • Architect and automate scalable cloud environments across AWS and Azure using Terraform, Ansible, Helm, and CDK.
  • Serve as Linux subject matter expert, managing system builds, core services, and performance from kernel up.
  • Lead CI/CD pipelines, observability, security, and incident response to ensure platform reliability.

Fueled is a leading digital strategy, design, and engineering agency. The 300+ person team has designed and built hundreds of digital products for major brands like Google, Apple, and The New York Times, and thrives in a culture that values flexibility, creativity, and cutting-edge technology.

$135,000–$150,000/yr
US

  • Build, deploy, and maintain secure GCP environments across Compute Engine, GKE, and Cloud Run.
  • Write reusable Terraform modules and CI/CD pipelines for automated testing and deployment.
  • Administer IAM roles, Workload Identity Federation, and GCP monitoring and alerting tools.

Latitude is a Service-Disabled Veteran Owned federal IT company that helps government agencies modernize technology. It promotes a remote work-from-home culture and focuses on automating mundane tasks and investing in employee growth.

Global

  • Contribute to platform and harness engineering, including CI/CD and developer tooling.
  • Build systems to reduce toil and maintain production infrastructure under conversational AI traffic.
  • Participate in on-call rotation and incident management to ensure platform uptime.

Replicant builds an AI-powered customer service platform that helps contact centers resolve requests and improve agent performance. The company is distributed, with a focus on ownership and collaboration, and serves Fortune 500 companies.

$168,675–$229,900/yr
US

  • Establish and employ continuous integration and delivery (CI/CD) patterns for successful software solutions.
  • Design secure, operationally sound solutions across AWS, Azure, OpenShift, and IBM Cloud.
  • Implement observability stacks, manage Kubernetes clusters, and automate infrastructure with Terraform.

Conga unifies commercial operations by aligning pricing, quoting, contracting, rebates, and communications so companies run as connected, smarter businesses. With more than 10,000 customers worldwide, including over 50% of the Fortune 100, Conga fosters a collaborative culture where every voice is heard.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

France

  • Evolve an Internal Developer Platform enabling development teams to deploy and operate applications securely with high self-service.
  • Play a crucial role in infrastructure architecture in a multi-cloud environment with a focus on GCP.
  • Build and maintain the platform used by over 50 internal clients, then support their concrete use by teams.

Lifen believes medical data can transform healthcare by reducing administrative burden, improving care coordination, and accelerating scientific discovery. Since 2015, the company has connected 800 hospitals and 150,000 healthcare professionals, with over 150 employees working remotely and from offices to unlock the potential of health data.

India

  • Build and maintain scalable cloud infrastructure for high availability.
  • Enhance observability and monitoring frameworks for accurate alerts.
  • Support on-call rotations and incident response with post-mortems.

GoGuardian is an award-winning learning solutions company purpose-built for K-12, trusted by educators to promote effective teaching and keep students safe. They are a remote, diverse, and committed team of mission-driven employees focused on improving learning environments.

$54,000–$64,800/yr
Europe

  • Own identity, access, and IT operations across the SaaS and cloud stack, automating repetitive tasks.
  • Manage user lifecycle, handle support requests, and administer tools like Google Workspace, GitHub, and 1Password.
  • Automate processes using Terraform, Python, and Slack APIs, and support security operations and deployments.

Centrifuge is building the open infrastructure for tokenized real-world assets, having partnered with S&P Dow Jones Indices and crossed $1.7B in TVL. The company is a well-funded, remote-first team of over 20 people, backed by leading investors and live across more than ten blockchains.

  • Own the reliability, performance, and scalability of Runlayer's infrastructure across AWS and GCP.
  • Manage Kubernetes clusters, database reliability, and CI/CD pipelines for rapid deployments.
  • Lead incident response and partner with product engineers to design resilient systems for enterprise customers.

Runlayer builds a unified platform for MCPs, Skills, and AI Agents, providing enterprises with security, governance, and observability to deploy AI safely and at scale. Founded by engineers who built AI Actions for OpenAI and Zapier Agents, the team has raised $42M from Felicis and Khosla Ventures, serving companies like Gusto, Instacart, and Opendoor.

UK Ireland Estonia Netherlands Sweden Israel Eastern Europe Portugal Unlimited PTO

  • Develop and evolve foundational software and services enabling product and development teams.\n- Architect, design, and implement Infrastructure as Code using Terraform.\n- Deploy, manage, and optimize Kubernetes clusters on GCP (GKE) and AWS (EKS).

DoiT is a global technology company that helps cloud-driven organizations leverage cloud for business growth and innovation through data, technology, and human expertise. They work with over 4,000 customers worldwide and foster a remote-first, entrepreneurial culture.