Source Job

Brazil 9w maternity 9w paternity

  • Monitor application environments to maintain high availability and detect production incidents rapidly.
  • Apply Site Reliability Engineering principles through automation, monitoring, and incident response on Google Cloud Platform.
  • Collaborate with development and operations teams to improve system resilience and operational efficiency.

Google Cloud Platform .NET Site Reliability Engineering Agile DevOps

20 jobs similar to Site Reliability Engineer [SRE]

Jobs ranked by similarity.

US

  • Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
  • Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
  • Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.

They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.

Europe UK North America

  • Design and evolve cloud infrastructure on GCP for scale and resilience.
  • Build internal tooling and automation that promote team autonomy and developer productivity.
  • Advance observability platform with metrics, logging, tracing, and alerting to reduce recovery time.

The company is a well-funded AI/ML company at the intersection of geospatial intelligence and climate technology, building products on scalable cloud infrastructure. The engineering team fosters a culture of reliability and continuous improvement, operating with a focus on SLOs, error budgets, and DORA metrics.

Latin America

  • Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.

We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.

Brazil

  • Maintain stability, reliability, and performance of cloud and core platform environments supporting enterprise customers.
  • Investigate complex technical issues across infrastructure, APIs, integrations, and distributed services.
  • Act as a technical bridge between customers, technical account teams, and Engineering to turn findings into actionable solutions.

Jobgether uses an AI-powered matching process to streamline job applications. It is an inclusive equal-opportunity workplace that processes applications quickly and fairly.

$192,000–$192,000/yr
US Unlimited PTO

  • Design, implement, and maintain scalable and reliable systems.
  • Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
  • Develop and maintain automation tools for deployment, monitoring, and system health checks.

LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.

US

  • Design and implement scalable cloud infrastructure to support growth.
  • Develop monitoring, alerting, and incident response for system reliability.
  • Automate deployment pipelines and ensure high availability and security.

Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.

US

  • Apply SRE principles to improve reliability, scalability, and performance of production systems.
  • Design and implement automation to reduce operational toil and improve engineering efficiency.
  • Lead incident response and develop sustainable solutions for complex production issues.

The hiring company is a technology organization focused on reliability and operational excellence. They offer a fully remote, collaborative environment with opportunities for technical leadership and career growth.

Europe

  • Own daily IT and platform operations, resolving access requests, deployments, and infrastructure tasks.
  • Manage cloud infrastructure on GCP and Cloudflare, CI/CD pipelines, and monitoring.
  • Collaborate with DevOps & Security lead to harden systems and scale the platform.

Centrifuge is building open infrastructure for real-world assets on blockchain, partnering with major financial institutions. We are a well-funded, small, high-trust team backed by leading investors, with over $1.7B in TVL.

Spain

  • Define and implement reliability strategy including SLOs, SLIs, error budgets, and incident practices.
  • Manage cloud infrastructure on AWS using Infrastructure as Code and ensure Kubernetes scalability.
  • Lead incident response and establish chaos engineering practices to strengthen platform resilience.

This partner company builds a globally scaled, AI-native platform with a focus on reliability and event-driven systems. They offer a collaborative international culture with significant technical ownership and continuous improvement.

US Unlimited PTO

  • Lead and mentor a US-based team of Site Reliability Engineers, driving operational excellence across production platforms.
  • Serve as an escalation point and incident commander for major production incidents, ensuring timely triage and resolution.
  • Drive automation, reliability, and observability improvements using Datadog, Kubernetes, and CI/CD pipelines.

NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions that partners with managed care organizations. Recognized as one of the fastest-growing companies in America with multiple locations in the US, South America, and India, they offer a fulfilling work environment that encourages associates to contribute to delivering premier service.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

Brazil

  • Lead and mentor software engineering teams to drive technical excellence and operational efficiency.
  • Collaborate with product and engineering leaders to align technical initiatives with business goals and long-term strategy.
  • Optimize team processes, manage capacity, and foster an inclusive environment for continuous improvement.

This company is a fast-growing technology organization focused on building scalable digital products. It fosters a culture of collaboration, inclusion, and continuous improvement, with significant career growth opportunities.

$240,000–$240,000/yr
US

  • Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
  • Drive SLOs, observability, alerting, and on-call processes across teams.
  • Build the platform engineering function from the ground up and influence cross-cutting architecture.

First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.

US

  • Provide production support and maintenance for enterprise applications on Google App Engine, ensuring availability and performance.
  • Lead incident response, troubleshooting, and root-cause analysis for user-impacting issues while meeting SLAs.
  • Own deployments, feature enhancements, and operational improvements across microservices environments on GCP.

Innodata is a global data engineering company enabling responsible AI advancement through data, evaluation frameworks, and human expertise. With over 36 years of legacy, it delivers high-quality data solutions and services to AI builders and adopters.

Brazil

  • Operate and evolve AWS infrastructure for Data/AI platforms, ensuring security, scalability, and high availability.
  • Build CI/CD pipelines, automate provisioning with Terraform, and implement observability.
  • Collaborate with Data, AI, and Infrastructure teams, document standards, and drive platform improvements.

The partner company is building a modern Data Platform team focused on secure, scalable, and highly available infrastructure for Data and AI workloads. The culture emphasizes DevOps, automation, and continuous improvement, with close collaboration across Data, AI, and Infrastructure teams.

$175,000–$195,000/yr

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
  • Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
  • Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.

Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.

Europe

  • Lead the SRE strategy and execution for a high-growth AI company.
  • Build and scale a high-performing SRE team while defining reliability standards.
  • Architect secure, scalable cloud infrastructure and implement observability practices.

This company develops advanced AI products and agentic technology. It operates in a high-growth, international environment with a focus on operational excellence and innovation.

Canada

  • Design, implement, maintain, and optimize highly available infrastructure supporting mission-critical applications and services.
  • Monitor production environments, analyze system performance, and proactively identify opportunities to improve stability, scalability, and operational efficiency.
  • Respond to technical escalations, troubleshoot infrastructure, networking, hardware, and software issues, and lead resolution of critical incidents.

Our partner is a technology company focused on high-availability platforms and mission-critical infrastructure. The team is collaborative and works with modern cloud technologies.

Brazil

  • Drive application modernization initiatives and design cloud-native solutions.
  • Apply DevOps, CI/CD, and automation to improve software delivery.
  • Leverage GCP and AI-driven approaches to enhance the software development lifecycle.

The partner company is a global technology consulting firm that helps organizations transform their technology landscape through modern application strategies. The company emphasizes collaboration, innovation, and continuous learning in a multicultural environment, though its size is not specified.

Global

  • Support the deployment, operation, and maintenance of the Karuna service running on Kubernetes.
  • Monitor production environments to ensure high availability, reliability, and performance.
  • Investigate, troubleshoot, and resolve production incidents, performing root cause analysis.

Software Mind develops solutions that make an impact for companies around the globe. They build cross-functional engineering teams with a culture of openness, respect, grit, and enjoyment.