Jobs Similar to Staff Site Reliability Engineer | TangerineFeed

Staff Site Reliability Engineer

StackBlitz 1 day ago

Global

Embed with product and platform teams from early stages to ensure reliability is designed in from the start.
Define production-readiness standards and measurable SLIs/SLOs to guide operational excellence.
Build tooling and infrastructure across AWS, GCP, and Azure using Terraform, and share on-call rotation.

AWS GCP Azure Terraform TypeScript

20 jobs similar to Staff Site Reliability Engineer

Jobs ranked by similarity.

Site Reliability Engineer (SRE)

Supabase 14 days ago

Global

Collaborate with service teams to define SLIs and SLOs based on customer experience and build error budget policies that influence engineering decisions.
Own the Operational Readiness Review process, conducting reviews for new services and major changes across observability, alerting, runbooks, capacity, and graceful degradation.
Act as a reliability expert for architecture reviews, failure mode analysis, dependency mapping, and resilience design.

Supabase provides the Postgres development platform with a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. With 280+ team members across 55+ countries, they are an open-source-first company that values async work and has raised $500M.

View details Similar jobs

Sr. Site Reliability Engineer

Versant 15 days ago

US

Lead design and operation of internal developer platforms and self-service infrastructure.
Build and optimize CI/CD pipelines, deployment workflows, and automation across GitHub Actions, Jenkins, ArgoCD.
Apply SRE principles to improve developer-facing systems and software delivery performance.

Versant is a media company owning iconic brands in news, sports, and entertainment, including USA Network, Fandango, and Rotten Tomatoes. It is an independent, publicly traded company with a collaborative, inclusive culture and a remote-first work environment.

View details Similar jobs

Senior Software Engineer- Site Reliability Engineering (SRE)

Noctua Technology, LLC 6 days ago

US

Drive the definition and adoption of SLIs and SLOs across services, reducing toil through automation and incident response.
Design and architect Infrastructure as Code solutions for large-scale environments using Docker, Kubernetes, and cloud-native services.
Serve as primary SRE liaison for development teams, influencing architecture and conducting training for clients.

Noctua Technology, LLC is a company that drives digital transformation by treating operations as a software engineering challenge, focusing on cloud native systems. They are a dynamic team seeking a Senior SRE to define strategy and bridge development and operations for clients.

View details Similar jobs

Senior Site Reliability Engineer (Remote Build)

Remote 18 days ago

Global Unlimited PTO 16w maternity 16w paternity

Own the operational excellence and infrastructure strategy for Remote Build's platform, ensuring reliability, performance, and security.
Lead incident response, build observability systems, and drive continuous improvement in system reliability.
Embed security into infrastructure, optimize costs, and automate operational toil to scale efficiently.

Remote solves modern organizations' biggest challenge of navigating global employment compliantly. With a fully distributed team across 6 continents, the company fosters a future-focused culture with core values of innovation and async work.

View details Similar jobs

Staff Software Engineer - Databases SRE

Grafana Labs 23 hours ago

UK Sweden Spain Germany Ireland 6w PTO

Partner with product engineering squads to own production reliability for high-SLA customer environments, designing automation and defining per-tenant SLOs.
Serve as a primary escalation point for incidents, leading response, post-incident reviews, and reducing SLO burn to prevent repeats.
Influence feature design for scalability and operability, improve alert quality, and eliminate toil through automation.

Grafana Labs is the company behind the open observability cloud, providing a fully managed observability platform for organizations to see, understand, and act on their data. With over 35 million users, 7,000+ customers, and 1,600+ team members across 40+ countries, we foster a remote, collaborative culture rooted in open-source values.

View details Similar jobs

Staff Engineer, Site Reliability

Babylist 22 days ago

US Canada

Own and evolve AWS infrastructure using Terraform, managing EKS clusters, databases, and core services.
Maintain CI/CD reliability and developer tooling across the full engineering org.
Lead incident response, drive post-incident reviews, and improve monitoring and alerting standards.

Babylist is the leading platform for expecting and new families, helping parents feel confident, connected, and cared for at every step. As a modern, AI-forward tech company with over 10 million yearly shoppers, Babylist has expanded into a full ecosystem and generated $750M in revenue in 2025, reshaping the $235B kids and baby market.

View details Similar jobs

Site Reliability Engineer (E3)

Vynca 17 days ago

US

Design, provision, and manage AWS infrastructure using Terraform and Kubernetes.
Build, operate, and improve observability, monitoring, and incident response processes.
Collaborate with engineering teams on capacity planning, performance optimization, and resilient system design.

Vynca provides comprehensive care for individuals with complex needs, focusing on quality days at home. The company is a close-knit community guided by core values of Excellence, Compassion, Curiosity, and Integrity.

View details Similar jobs

Head of Site Reliability Engineering

Titan 12 days ago

US

Build the SRE practice from scratch: define SLO frameworks, on-call rotation, and incident command for live bank customers.
Define severity tiers, SLA commitments, and escalation paths for production support, acting as the technical owner during incidents.
Set engineering operations across sprint discipline, release rituals, code review standards, and compliance artifacts for bank examiners.

Titan builds AI software for banks, specializing in purpose-built small language models and AI bankers that financial institutions trust. The company is a backed fintech startup scaling from a handful to hundreds of customers, with a hands-on, build-first culture under strict compliance standards.

View details Similar jobs

Site Reliability Engineer (SRE)

Synthesia 14 days ago

US

Take ownership of incident management and operational excellence across cloud infrastructure.
Automate high-risk manual processes and drive reliability gains through engineering.
Own a platform domain such as Temporal, observability, or Kubernetes operations.

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London with offices across Europe and the US, and has over $530 million in funding from premier investors like Accel and Nvidia's VC arm.

View details Similar jobs

Senior Site Reliability Engineer (SRE)

Oowlish 10 days ago

Latin America

Design, implement, and improve Site Reliability Engineering practices across production environments with a focus on SLOs, SLIs, and error budgets.
Lead incident response processes and build observability strategies including monitoring, logging, alerting, and distributed tracing.
Partner with engineering teams to enhance system reliability, availability, scalability, and operational efficiency.

Oowlish is a rapidly expanding software development company in Latin America that collaborates with premier clients from the United States and Europe to create pioneering digital solutions. Certified as a Great Place to Work, it offers a nurturing environment with opportunities for professional growth and international impact.

View details Similar jobs

Senior Site Reliability Engineer

CertifyOS 15 days ago

US Unlimited PTO

Design and build cloud-native infrastructure for reliability, observability, and automation across GCP, GKE, and Cloud Run.
Own incident response, root cause analysis, escalation workflows, and runbooks to prevent hard problems from recurring.
Develop Infrastructure as Code, CI/CD pipelines, and operational tooling to improve developer velocity and platform efficiency.

CertifyOS is building the data infrastructure that powers modern healthcare, automating provider licensing, enrollment, credentialing, and network monitoring through an API-first platform. The company is backed by leading investors with a team of deep experience in provider data systems, valuing authenticity, accountability, collaboration, results, and openness to feedback.

View details Similar jobs

SRE Engineer

IPSY 3 days ago

Mexico Colombia

Build and maintain observability across the platform in Datadog, including dashboards, monitors, APM, and log pipelines.
Participate in on-call rotation and incident response, driving blameless post-incident reviews and automating toil.
Leverage AI tools to accelerate debugging, generate runbooks, and build automation for operational efficiency.

IPSY is a beauty subscription platform that connects brands and consumers through curated beauty products. It is a remote-first company with a focus on community and engagement.

View details Similar jobs

Principal Cloud Engineer - DevOps/Infrastructure

Resultant 6 days ago

United States

Act as the design authority for multi-cloud infrastructure across AWS, Google Cloud, and Azure, owning the hardest architecture decisions.
Define firm-wide standards, patterns, and reference architectures for landing zones, networking, identity, and workload platforms.
Build reusable, modular Terraform and Kubernetes standards, and drive CI/CD pipelines, observability, security, and cost optimization.

Resultant is a modern consulting firm that partners with clients to solve complex challenges through data analytics, technology solutions, and digital transformation. Founded in Indianapolis in 2008, it employs over 400 team members across offices in the United States, fostering a culture of collaboration and ownership.

View details Similar jobs

Platform Engineer

Terzo 11 days ago

US

Owning cloud infrastructure on Azure, data pipeline orchestration, CI/CD, and observability to ensure production-grade reliability.
Building and maintaining foundational infrastructure that enables fast engineering velocity without breaking things.
Applying SRE principles such as SLOs, capacity planning, incident response, and eliminating toil through automation.

Terzo's platform processes enterprise-scale document corpora, powers real-time AI agents, and serves the Financial Intelligence Graph to Fortune 500 customers. As a small, senior team with strong ownership and minimal bureaucracy, we foster a culture of collaboration, mentorship, and continuous improvement.

View details Similar jobs

Senior / Staff Site Reliability Engineer (Infrastructure), APAC

Tilt 21 days ago

APAC

Define and own the APAC infrastructure architecture end-to-end on Azure, including compute, networking, and containerisation.
Lead incident response for the region with calm, methodical root cause analysis and durable fixes.
Drive infrastructure migrations and PCI-DSS hardening programs across clouds safely.

Tilt is a mobile-first fintech company that uses machine learning to provide credit beyond traditional credit scores. With millions of customers worldwide, they value ownership, excellence, and mutual respect.

View details Similar jobs

Sr. Site Reliability Engineer

Filevine 9 days ago

United States

Own and evolve observability strategy including monitoring, alerting, dashboards, logging, and distributed tracing.
Define and manage SLIs, SLOs, and reliability metrics, improving MTTD and MTTR through automation.
Build and maintain reliable cloud infrastructure on AWS and Kubernetes while mentoring engineers on SRE best practices.

Filevine is a Legal AI company delivering Legal Operating Intelligence for legal work. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.

View details Similar jobs

Senior Site Reliability Engineer

MZLA Technologies Corporation 25 days ago

US 5w PTO

Design and develop CI/CD systems for websites, services, and release workflows, and operate an EKS-based Kubernetes platform.
Diagnose debug production incidents, drive root-cause analysis, and implement improvements to enhance system reliability.
Write and maintain infrastructure as code using Pulumi or Terraform/OpenTofu across multiple AWS accounts with security-conscious practices.

Thunderbird is one of the world’s most trusted open-source email applications, empowering more than 20 million people globally. Our small but growing distributed team includes 65+ people across seven countries, and we build privacy-respecting communication tools with a collaborative, inclusive, and user-first spirit.

View details Similar jobs

Sr. Cloud Platform Engineer

Applied 16 days ago

North America

Design, build, and maintain cloud infrastructure across Azure, GCP, and AWS, including landing zones, Kubernetes, and CI/CD pipelines.
Implement monitoring, security, and hybrid connectivity for enterprise-scale cloud environments.
Collaborate cross-functionally, mentor engineers, and leverage AI tools to accelerate infrastructure development.

Applied is an Insurtech company that builds technology solutions for insurance professionals. With over 40 years of experience, they foster a culture of trust, inclusion, and growth.

View details Similar jobs

Senior Site Reliability Engineer (SRE)

Oowlish 3 days ago

Latin America

Define and implement SLOs, SLIs, and Error Budgets to ensure production system reliability.
Lead incident command during major outages and drive blameless postmortems.
Develop observability strategies, including monitoring, logging, tracing, and alerting.

Oowlish is a rapidly expanding software development company in Latin America. It is certified as a Great Place to Work and offers a nurturing environment with professional development opportunities.

View details Similar jobs

Senior Site Reliability Engineer II - Infrastructure (AI Native)

Life360 18 days ago

Canada

Build and maintain infrastructure platforms for over 200 backend services running on Kubernetes clusters with 40,000+ cores.
Lead and mentor other engineers, own complex infrastructure failures, and participate in a shared on-call rotation.
Drive cloud cost efficiency, estimate schedules, and use AI tools as a first-class collaborator in daily workflows.

Life360's mission is to keep people close to the ones they love through location sharing, safe driver reports, and crash detection. The company serves approximately 97.8 million monthly active users across more than 180 countries and has more than 500 remote-first employees.

View details Similar jobs