Source Job

US

  • Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil and improve platform reliability.
  • Build observability and monitoring solutions with Grafana and AWS CloudWatch, and implement Infrastructure as Code using Terraform.
  • Develop CI/CD pipelines, define SRE standards and metrics, and participate in incident management and on-call rotation.

Python AWS Terraform CI/CD SRE

20 jobs similar to Senior Cloud Site Reliability Engineer (SRE)

Jobs ranked by similarity.

India

  • Design, deploy, and maintain the reliability, availability, and performance of critical systems and APIs across AWS and GCP.
  • Build observability frameworks, define SLIs/SLOs, and implement monitoring using Datadog and Kubernetes.
  • Participate in on-call rotations, incident response, and blameless post-incident reviews to drive systemic improvements.

JumpCloud is an AI-powered unified IT management platform that secures the modern workforce by consolidating identity, device, and access management. The company is remote-first with teams in over 15 countries and values building connections, thinking big, and continuous improvement.

India

  • Architect and scale multi-region microservices, APIs, and authentication infrastructure on AWS/GCP.
  • Lead SLOs, observability, incident management, and disaster recovery automation to maintain 99.99% availability.
  • Manage Kubernetes clusters and Terraform IaC while eliminating toil with Python/Go tooling.

JumpCloud is an AI-powered unified IT management platform that secures the modern workforce through identity, device, and access management. The company is remote-first with teams in 15+ countries and values building connections, thinking big, and continuous improvement.

$150,000–$185,000/yr
US Unlimited PTO

  • You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
  • You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
  • You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.

Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.

India

  • Build and maintain scalable cloud infrastructure for high availability.
  • Enhance observability and monitoring frameworks for accurate alerts.
  • Support on-call rotations and incident response with post-mortems.

GoGuardian is an award-winning learning solutions company purpose-built for K-12, trusted by educators to promote effective teaching and keep students safe. They are a remote, diverse, and committed team of mission-driven employees focused on improving learning environments.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

Brazil Unlimited PTO

  • Build and maintain the company's internal platform, driving operational excellence.
  • Collaborate with engineering squads to ensure applications are safe and reliable.
  • Take ownership of software infrastructure projects and provide off-hours support.

Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.

$165,000–$165,000/yr
US Unlimited PTO

  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure.
  • Build and manage CI/CD pipelines, automate operational tasks, and improve deployment processes.
  • Monitor production systems, participate in incident response, and champion DevOps best practices.

First Due provides transformative end-to-end software solutions for fire and EMS agencies, helping them run safer, smarter, and more effective operations. The company offers a fully remote workplace, comprehensive benefits, and opportunities for advancement, with a culture focused on respect, inclusivity, and equal opportunity.

$175,000–$185,000/yr
US

  • Consolidate Terraform and establish conventions for state management, modules, and CI checks.
  • Improve monitoring, observability, and automation in Datadog and Cloud Monitoring.
  • Right-size workloads, evaluate Kubernetes architecture, and retire legacy tooling.

Vida is a virtual, personalized obesity care provider that combines evidence-based treatment with advanced technology to help patients improve their health. Trusted by Fortune 100 companies and growing for years, Vida takes a whole-person approach to care and celebrates diversity across its team.

US

  • Design, build, and optimize secure CI/CD pipelines within GitLab, enforcing DevSecOps principles through automated vulnerability scanning and compliance gates.
  • Manage and scale AWS infrastructure, optimizing workloads on EKS, EC2, S3, and RDS.
  • Establish comprehensive monitoring, logging, and alerting systems across EKS and EC2 nodes to ensure maximum uptime.

LMI is a digital solutions provider that accelerates government impact with innovation and speed, offering commercial-grade platforms and mission-ready AI to federal agencies. Headquartered in Tysons, Virginia, the company serves defense, space, healthcare, and energy sectors, focusing on agility and collaboration.

Brazil

  • Design, implement, and evolve cloud platforms with focus on reliability, scalability, and security.
  • Build and maintain CI/CD pipelines, automate infrastructure using Terraform, Kubernetes, and Docker.
  • Implement observability, define SLIs/SLOs, and lead incident investigation and root-cause analysis.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through a fair, objective review process. The platform ensures applications are quickly evaluated and shortlists are shared with employers, who manage interviews and final decisions.

$180,000–$220,000/yr
US

  • Design, implement, and maintain reliable, scalable, and secure infrastructure to support applications and automation systems.
  • Automate infrastructure provisioning, configuration management, and deployment pipelines using tools like Terraform and ArgoCD.
  • Implement observability solutions and enforce security best practices to ensure uptime and system performance.

Bright Machines is a next-generation, AI-enabled manufacturer focused on data center infrastructure production, using proprietary AI-based robotics and software to assemble hardware products for hyperscalers and OEMs. The company is headquartered in San Francisco, California, with an integration center in Guadalajara, Mexico, and has been recognized by Forbes' AI 50 and other leading organizations.

$250,000–$280,000/yr
United States

  • Lead both Site Reliability Engineering and Corporate IT, shaping technology operations strategy for a global SaaS environment.
  • Oversee platform reliability, incident response, compliance, and automation to reduce manual effort and improve efficiency.
  • Partner with Security and Compliance teams to support FedRAMP authorization and NIST 800-53 compliance frameworks.

Our partner is a high-growth, global SaaS environment seeking a Vice President of Enterprise Technology. They offer a fully remote executive role with approximately 10% travel and a focus on operational excellence.

US

  • Own the CI/CD foundation for AI-assisted development, including hardened pipeline gates and automated rollback.
  • Build centralized logging, monitoring, and alerting across AWS, and lead security engineering.
  • Contribute to backend development for products while building paved-path self-service tooling.

Samsara is a public company pioneering the Connected Operations Cloud, using IoT data to improve safety and efficiency in industries like transportation and manufacturing. With a focus on innovation and long-term growth, it fosters a culture of autonomy and support for its employees.

  • Own the reliability, performance, and scalability of Runlayer's infrastructure across AWS and GCP.
  • Manage Kubernetes clusters, database reliability, and CI/CD pipelines for rapid deployments.
  • Lead incident response and partner with product engineers to design resilient systems for enterprise customers.

Runlayer builds a unified platform for MCPs, Skills, and AI Agents, providing enterprises with security, governance, and observability to deploy AI safely and at scale. Founded by engineers who built AI Actions for OpenAI and Zapier Agents, the team has raised $42M from Felicis and Khosla Ventures, serving companies like Gusto, Instacart, and Opendoor.

$100,000–$166,000/yr
US

  • Design, build, and maintain CI/CD pipelines for secure application delivery.
  • Automate infrastructure provisioning and deploy cloud services using AWS and Terraform.
  • Integrate security controls and enforce cybersecurity standards in development and deployment.

LMI is a digital solutions provider dedicated to accelerating government impact with innovation and speed. Headquartered in Tysons, Virginia, LMI serves the defense, space, healthcare, and energy sectors, helping agencies navigate complexity.

Global 4w PTO

  • Champion SRE culture and best practices to improve production reliability and system resilience.
  • Communicate with stakeholders at all stages and bring fresh ideas to the table.
  • Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.

Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.

$133,000–$209,000/yr
US

  • Design and continuously improve detection and alerting controls to reduce noise and enable rapid response.
  • Build, test, and automate incident response playbooks to increase efficiency across the incident lifecycle.
  • Drive prioritization of alerts using a data-driven triage framework aligned with business impact and threat context.

Sword builds AI to heal billions, pioneering AI Care for healthcare delivery across physical therapy, women’s health, and more. With over 700,000 members and 1,000+ enterprise clients, they have raised $500 million and emphasize a culture of proactive security and AI proficiency.

$124,750–$178,215/yr
US

  • Design, build, and maintain automation pipelines for workload packaging, validation, testing, deployment, and monitoring across AWS environments.
  • Collaborate with data engineers to operationalize workloads within a data lakehouse ecosystem and develop reusable infrastructure-as-code constructs using Terraform, AWS CDK, or CloudFormation.
  • Ensure pipelines meet enterprise security, compliance, and scalability standards while mentoring junior engineers and contributing to DevSecOps practices.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. The company uses AI to review applications and ensure fair, objective candidate evaluation, and operates in a distributed, fully remote environment.

Europe

  • Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
  • Establish and manage SLOs and SLAs for ClickHouse Cloud, ensuring monitoring and alerting are in place for all infrastructure.
  • Lead incident response, blameless postmortems, and chaos initiatives to continuously improve reliability and performance.

ClickHouse develops an open-source column-oriented database management system and offers a cloud database service. The company is a rapidly scaling, globally distributed startup with employees in over 25 countries, offering a flexible and collaborative culture.

Canada

  • Design and operate scalable AWS infrastructure with containerization and orchestration tools.
  • Implement monitoring, logging, and infrastructure as code using Terraform.
  • Improve CI/CD pipelines and troubleshoot production issues in complex SDLC environments.

Sureify builds systems that support millions of users. It is a high-growth, engineering-driven SaaS company with a remote-first culture across the Americas.