Source Job

Brazil Unlimited PTO

  • Build and maintain the company's internal platform, driving operational excellence.
  • Collaborate with engineering squads to ensure applications are safe and reliable.
  • Take ownership of software infrastructure projects and provide off-hours support.

AWS Kubernetes Terraform Python CI/CD

20 jobs similar to Site Reliability Engineer, Tech Lead

Jobs ranked by similarity.

US

  • Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
  • Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
  • Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.

Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.

US Unlimited PTO

  • Design, build, and operate shared cloud infrastructure using AWS, Kubernetes, Terraform, Databricks, and Cloudflare.
  • Deliver SRE and DevOps initiatives to improve reliability, scalability, observability, and deployment safety.
  • Build reusable infrastructure modules, automation, and self-service workflows to reduce manual work and improve developer experience.

YipitData is the leading market research and analytics firm for the disruptive economy, recently raising up to $475M from The Carlyle Group at a valuation over $1B. We analyze billions of alternative data points daily and have been recognized as one of Inc’s Best Workplaces, cultivating a people-centric culture focused on mastery, ownership, and transparency.

Brazil

  • Drive the performance, stability, security, and reliability of production environments with a focus on automation and proactive improvements.
  • Design and maintain infrastructure using Infrastructure as Code tools like Terraform, and manage Kubernetes and cloud environments.
  • Lead vulnerability management, incident response, and secure CI/CD practices to ensure resilience and operational excellence.

Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. It processes applications and shares shortlists with employers, offering a remote-first and inclusive work environment.

Latin America

  • Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.

We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.

US Unlimited PTO

  • Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
  • Design and maintain infrastructure as code across multiple cloud providers.
  • Provide technical leadership and mentorship across the Systems Engineering team.

Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.

$241,000–$270,000/yr
US Unlimited PTO

  • Architect the end-to-end reliability, performance, and resilience of cloud environments, including the SLO framework for critical services.
  • Lead incident response, on-call rotation, root cause analysis, and build a culture of corrective actions.
  • Build observability platforms to detect issues proactively and mentor engineers on reliability standards.

Garner is on a mission to transform the U.S. healthcare system by partnering with employers to steer members to better-performing doctors, resulting in better care and lower costs. With 550+ proprietary clinical metrics, they have helped over 2.5 million people and saved $1B in healthcare costs, recently raising a Series E and doubling five years running.

$53,300–$119,850/yr
Global Unlimited PTO 16w maternity 16w paternity

  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.

Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.

Brazil

  • Define and monitor reliability metrics such as SLI, SLO, SLA, MTTR, and MTTD.
  • Implement observability solutions including monitoring, alerting, dashboards, and APM.
  • Collaborate with multidisciplinary teams to embed reliability and observability into solutions.

The company is a technology organization focused on building and maintaining reliable digital environments. It fosters a culture of engineering excellence, collaboration, and data-driven decision-making.

$191,000–$226,000/yr
US Unlimited PTO

  • Own the reliability, performance, and resilience of cloud environments (AWS, Kubernetes) and define SLOs across critical services.
  • Lead incident response, on-call rotation, and drive root cause analysis to ensure high production quality.
  • Build and maintain observability systems and automate operational toil using AI tools.

Garner partners with employers to redesign healthcare by using clinical metrics to identify top doctors and incentivize members to better care. The company has helped over 2.5 million people, saved $1B in costs, and doubled annually for five years, fostering a mission-driven, high-performance culture.

US

  • Improve system availability, scalability, and resilience across Flowcode's platforms.
  • Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
  • Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.

Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.

$114,700–$195,000/yr
North America

  • Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
  • Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
  • Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.

Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.

Brazil

  • Ensure reliability, scalability, and performance of cloud-based systems using Kubernetes and observability tools.
  • Define and monitor reliability metrics (SLIs, SLOs, MTTR) to continuously improve operational performance.
  • Automate operational tasks and implement Infrastructure as Code to reduce manual work and enhance efficiency.

Our partner is a technology company focused on building and maintaining reliable, scalable digital environments. They promote a culture of continuous improvement, collaboration, and proactive engineering.

Global 4w PTO

  • Champion SRE culture and best practices to improve production reliability and system resilience.
  • Communicate with stakeholders at all stages and bring fresh ideas to the table.
  • Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.

Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.

$160,000–$185,000/yr
US Unlimited PTO

  • Actively identify, plan and implement developer tooling and automation.
  • Participate in on-call rotations and assist with diagnostics and troubleshooting of platform and infrastructure.
  • Set technical direction for the team's infrastructure decisions and define overall DevOps strategy.

Bluesight creates groundbreaking solutions that increase efficiency, safety and visibility for health systems, hospital pharmacy, and pharmaceutical manufacturers. They are a high-growth healthcare information technology company with over 3,000 customers and a startup culture.

US

  • Apply SRE principles to improve reliability, scalability, and performance of production systems.
  • Design and implement automation to reduce operational toil and improve engineering efficiency.
  • Lead incident response and develop sustainable solutions for complex production issues.

The hiring company is a technology organization focused on reliability and operational excellence. They offer a fully remote, collaborative environment with opportunities for technical leadership and career growth.

$180,000–$250,000/yr
US Unlimited PTO

  • Own the technical direction and architecture of critical infrastructure domains, establishing scalable patterns and standards.
  • Lead complex, multi-team infrastructure initiatives from design through implementation and production operation.
  • Design and evolve AWS and Kubernetes infrastructure to enable teams to build and deploy systems reliably at scale.

We provide innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. Our company is backed by world-class investors including Craft Ventures and Andreessen Horowitz, with offices across the US and India, and we are growing extremely quickly.

India

  • Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
  • Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
  • Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.

Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.

$75,450–$169,700/yr
Global Unlimited PTO

  • Lead Remote's SRE team owning Kubernetes, AWS, PostgreSQL, CI, and observability.
  • Balance 60% hands-on technical work with 40% people leadership and career growth.
  • Drive a maturing reliability practice including SLOs, incident response, and on-call.

Remote is a global employment platform that helps companies recruit, pay, and manage international teams. The company is fully remote with a future-focused, async culture and employees across six continents.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

Brazil 4w PTO

  • Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
  • Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
  • Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.

Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.