Source Job

US

  • Develop and maintain infrastructure automation solutions using Ansible.
  • Design, implement, and enhance CI/CD pipelines and operational tooling.
  • Troubleshoot complex Linux-based infrastructure and distributed systems issues to maintain high availability.

Ansible Ruby Linux CI/CD AWS

20 jobs similar to Site Reliability Engineer

Jobs ranked by similarity.

Canada

  • Design, implement, maintain, and optimize highly available infrastructure supporting mission-critical applications and services.
  • Monitor production environments, analyze system performance, and proactively identify opportunities to improve stability, scalability, and operational efficiency.
  • Respond to technical escalations, troubleshoot infrastructure, networking, hardware, and software issues, and lead resolution of critical incidents.

Our partner is a technology company focused on high-availability platforms and mission-critical infrastructure. The team is collaborative and works with modern cloud technologies.

Brazil

  • Ensure system architecture meets technical requirements by collaborating with IT teams (Architecture, Security, Infrastructure).
  • Maintain and evolve the microservices environment on AWS with a focus on information security.
  • Implement DevOps practices, automation, and monitoring tools to ensure system reliability and scalability.

Experian is a global data and technology company that powers opportunities for people and businesses worldwide. With 25,200 employees across 32 countries, it fosters a people-centric, inclusive culture recognized as a World's Best Workplace.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.

US

  • Design and implement scalable cloud infrastructure to support growth.
  • Develop monitoring, alerting, and incident response for system reliability.
  • Automate deployment pipelines and ensure high availability and security.

Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.

United States

  • Design and build scalable, secure cloud infrastructure and deployment pipelines.
  • Own infrastructure end-to-end, including architecture, provisioning, deployment, and operation.
  • Lead technical design discussions and contribute to infrastructure and platform architecture decisions.

VulnCheck is the Exploit Intelligence Company, delivering structured exploit intelligence for cybersecurity. Founded in 2021, the company has a transparent, collaborative, and supportive culture with a team of experts.

Global

  • Ensure availability, performance, scalability, and resilience of production services in AWS.
  • Automate infrastructure provisioning and management using Infrastructure as Code (IaC).
  • Collaborate with development, architecture, security, and product teams to promote reliability best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide, operating in markets such as financial services, healthcare, automotive, and insurance. The company has over 25,200 employees across 32 countries and is recognized as a Top 25 global workplace by Fortune.

$75,600–$124,200/yr
Europe

  • Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
  • Drive automation and eliminate operational toil through self-service tooling and process improvements.
  • Lead incident response and mentor engineers to improve reliability practices.

Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.

US

  • Participate in DevOps initiatives and drive automation and cloud-native solutions to enhance system efficiency and reliability.
  • Collaborate within an Agile DevOps team and implement Infrastructure as Code using Terraform/Ansible across AWS services.
  • Design and optimize CI/CD pipelines using Jenkins, GitHub, and container-based platforms while ensuring security and compliance.

Sparksoft provides innovative IT solutions to government clients, serving as a catalyst for change. They are a collaborative team of problem-solvers and innovators committed to excellence, recognized as a Great Place to Work.

Canada

  • Design, implement, and maintain highly available and scalable infrastructure solutions.
  • Monitor system performance, identify bottlenecks, and resolve reliability issues proactively.
  • Automate infrastructure deployment, configuration management, and operational workflows.

The company is a technology firm that provides critical authorization solutions to organizations worldwide. It is a remote-first organization with a collaborative culture, offering equity opportunities and a focus on team building.

$120,000–$150,000/yr
US Unlimited PTO

  • Optimize new and existing systems by increasing reliability, performance, and scalability.
  • Automate routine operational tasks to reduce toil and improve efficiency.
  • Ensure infrastructure security compliance and implement least-privilege access controls.

Prove provides phone-centric identity tokenization and passive cryptographic authentication solutions to reduce friction and enhance security across digital channels. With over 1,000 enterprise customers processing 20 billion requests annually, they foster a fast-paced, collaborative culture focused on impact and tenacity.

$120,000–$165,000/yr
US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

$118,000–$151,000/yr
US 4w PTO

  • Build and improve platform services, including CI/CD pipelines and cloud infrastructure.
  • Collaborate with senior engineers to design scalable solutions and enhance developer experience.
  • Participate in incident response and retrospectives to drive continuous improvement.

Octopus Energy is a tech-powered energy company focused on renewable energy and customer experience. The company culture emphasizes ownership, collaboration, and making a tangible impact across teams.

$101,260–$101,260/yr
Canada

  • Act as technical point of contact for customer cloud environments, leading reliability initiatives and incident response.
  • Troubleshoot complex infrastructure issues and drive automation using scripting and infrastructure-as-code.
  • Participate in 24/7 on-call rotations to ensure timely resolution of production incidents.

Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. They process applications for partner companies and manage initial candidate screening.

Ireland

  • Design, build, and deploy production systems with focus on scalability, reliability, and security.
  • Develop and maintain automation to streamline operations and eliminate toil.
  • Proactively monitor systems and implement automated incident response to minimize downtime.

Arista Networks is an industry leader in data-driven networking for large data centers, campus, and routing. With over $8 billion in revenue and a culture valuing diversity, Arista is a Great Place to Work for Best Engineering Team and Best Company for Diversity.

North America Canada Latin America

  • Design and operate scalable cloud infrastructure across AWS and GCP.
  • Build and improve Kubernetes, Linux, and cloud networking environments.
  • Strengthen security, disaster recovery, and platform resilience.

Hubstaff provides workforce analytics and time tracking for remote teams, serving over 200,000 global users. The company is a product-led organization with a winning culture and a fully remote team of experienced engineers.

Global

  • Manage and optimize multi-cloud infrastructure (AWS required, GCP optional) with Kubernetes and CI/CD pipelines.
  • Improve observability through monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Coralogix).
  • Drive automation and Infrastructure as Code (IaC) using Terraform and Helm, and provide architectural guidance.

NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior. In 2023, NIQ combined with GfK, bringing together two industry leaders with operations in 100+ markets and covering more than 90% of the world's population.

$152,000–$195,000/yr
US Unlimited PTO

  • Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
  • Build and operate AI tooling infrastructure, including MCP servers and secure AI access.
  • Optimize CI/CD pipelines, implement progressive delivery, and advance Infrastructure as Code.

SecurityScorecard is the global leader in cybersecurity ratings, rating over 12 million companies across 64 countries. Headquartered in New York, it is recognized as a best workplace and funded by top investors.

$200,000–$235,000/yr
US 2w PTO

  • Handle infrastructure work across a growing scope — from hand-picked, low-urgency familiarization tasks toward independently managing significant components.
  • Troubleshoot and resolve infrastructure issues with autonomy, seeking guidance on necessary changes and architecture decisions.
  • Engage in infrastructure-as-code projects and play an active role in CI/CD pipeline management.

Planning Center exists to help churches build stronger, more connected communities. Since 2006, over 80,000 churches have trusted our tools, and we are an independent, profitable company with no outside investors, operating with a fully remote team.

$185,000–$280,000/yr
US 4w PTO

  • Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
  • Scale single-tenant deployments and build observability, incident response, and compliance practices.
  • Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.

Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.

$160,000–$208,000/yr
US

  • Build systems for declarative application and infrastructure lifecycle management, including CI/CD, Kubernetes, and service inventory.
  • Prioritize and troubleshoot infrastructure issues to minimize downtime and respond to alerts efficiently.
  • Contribute to setting the SRE team's direction and streamline automation of infrastructure processes.

Counterpart Health develops Counterpart Assistant, an AI-enabled primary care tool that supports physicians in chronic disease management. It is a subsidiary of Clover Health, with a remote-first culture and a focus on value-based care through technology.