Design, implement, and improve Site Reliability Engineering practices across production environments with a focus on SLOs, SLIs, and error budgets.
Lead incident response processes and build observability strategies including monitoring, logging, alerting, and distributed tracing.
Partner with engineering teams to enhance system reliability, availability, scalability, and operational efficiency.
Oowlish is a rapidly expanding software development company in Latin America that collaborates with premier clients from the United States and Europe to create pioneering digital solutions. Certified as a Great Place to Work, it offers a nurturing environment with opportunities for professional growth and international impact.
Lead a Dedicated Tenant Site Reliability Engineering organization, driving complex initiatives and operational excellence across multiple teams.
Oversee delivery and operation of PingOne Advanced Identity Cloud and Advanced Services, improving consistency and reliability.
Partner with SRE, Security, and Development teams to manage dependencies and evolve software delivery strategies.
Ping Identity provides an intelligent cloud identity platform that secures and streamlines digital experiences. Headquartered in Denver, Colorado, the company serves more than half of the Fortune 100 and fosters a culture that champions individuality and digital freedom.
Define and implement SLOs, SLIs, and Error Budgets to ensure production system reliability.
Lead incident command during major outages and drive blameless postmortems.
Develop observability strategies, including monitoring, logging, tracing, and alerting.
Oowlish is a rapidly expanding software development company in Latin America. It is certified as a Great Place to Work and offers a nurturing environment with professional development opportunities.
Own reliability and operational stability of BJAK’s production systems.
Design and improve monitoring, alerting, logging and observability across services.
Lead incident response, troubleshooting and structured root cause analysis.
BJAK's automation systems power end-to-end insurance journeys across quote generation, policy issuance, claims, and more. They are a global engineering team with modern engineering culture, offering fully remote work and a high-ownership environment.
Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.
Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.
Own and evolve observability strategy including monitoring, alerting, dashboards, logging, and distributed tracing.
Define and manage SLIs, SLOs, and reliability metrics, improving MTTD and MTTR through automation.
Build and maintain reliable cloud infrastructure on AWS and Kubernetes while mentoring engineers on SRE best practices.
Filevine is a Legal AI company delivering Legal Operating Intelligence for legal work. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.
Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.
MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.
Design the human side of reliability: on-call rotations, incident roles, blameless postmortems, and change management.
Build and operate the core reliability toolkit: observability, CI/CD, infrastructure as code, and incident response.
Embed with product engineers to raise operational literacy and define SLOs grounded in patient needs.
Novellia is a health tech startup that provides patients with access to their health records and enables biopharma innovation through real-world data. They are growing 5x year over year, have raised close to $30M in funding, and are backed by tier-1 investors, with a focus on building a diverse and inclusive workplace.
Establish an SRE function, building a team of 4-6 and professionalizing incident management with clear processes and tooling.
Drive engineering excellence through design reviews, code reviews, and blameless retrospectives to foster a quality culture.
Lead people and projects, balancing incident response with a roadmap of observability and reliability engineering initiatives.
We are a fast-growing, remote cybersecurity company dedicated to enabling organizations to proactively find and fix exploitable attack vectors. Our team is a fusion of former special operations cyber operators and engineers, committed to a culture of respect, collaboration, and ownership.
Co-own the architecture of cloud infrastructure on Azure and Kubernetes clusters for high throughput and availability.
Drive resilience strategy for global scaling, zero-downtime deployments, and disaster recovery.
Evolve observability stack with LGTM (Loki, Grafana, Tempo, Mimir) and lead incident response.
Flip is an AI-powered employee experience platform for frontline workers in retail, manufacturing, and logistics. The company is a young, rapidly growing tech company with a remote-first culture and offices in Berlin and Stuttgart.
Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
Partnering with development teams to establish production readiness and operational readiness.
Building tooling to automate observability and operational workflows, eliminating manual toil.
Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.
Drive the definition and adoption of SLIs and SLOs across services, reducing toil through automation and incident response.
Design and architect Infrastructure as Code solutions for large-scale environments using Docker, Kubernetes, and cloud-native services.
Serve as primary SRE liaison for development teams, influencing architecture and conducting training for clients.
Noctua Technology, LLC is a company that drives digital transformation by treating operations as a software engineering challenge, focusing on cloud native systems. They are a dynamic team seeking a Senior SRE to define strategy and bridge development and operations for clients.
Collaborate with engineering teams to design scalable, secure systems.
Establish SLOs, manage incident response, and drive reliability improvements.
Leverage expertise in Go, Python, Kubernetes, and cloud platforms.
ClickHouse is a leading real-time analytics company recognized on the 2025 Forbes Cloud 100 list. With over 3,000 customers and rapid growth, the company offers a remote-friendly, globally distributed culture.
Partner with product engineering squads to own production reliability for high-SLA customer environments, designing automation and defining per-tenant SLOs.
Serve as a primary escalation point for incidents, leading response, post-incident reviews, and reducing SLO burn to prevent repeats.
Influence feature design for scalability and operability, improve alert quality, and eliminate toil through automation.
Grafana Labs is the company behind the open observability cloud, providing a fully managed observability platform for organizations to see, understand, and act on their data. With over 35 million users, 7,000+ customers, and 1,600+ team members across 40+ countries, we foster a remote, collaborative culture rooted in open-source values.
Design, implement, and manage scalable, highly available systems on Azure Cloud, optimizing Kubernetes and containerized workloads.
Build and maintain robust CI/CD pipelines with GitHub Actions and implement infrastructure as code using Helm Charts.
Monitor system performance, troubleshoot issues, ensure uptime, and perform root cause analysis to improve reliability.
Resilinc is pioneering intelligent, autonomous systems that redefine supply chain risk management using agentic AI, trusted by top companies in life sciences, aerospace, high tech, and automotive. We are a fully remote, mission-led team with a collaborative culture focused on high-impact work.
Design, build, and run distributed cloud architectures and large-scale production systems.
Ensure reliability, observability, performance, and cost efficiency of the platform.
Collaborate with product and backend teams to design system architecture and optimize resource use.
Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.
Ensure availability, performance, scalability, and resilience of production services in AWS.
Automate infrastructure provisioning and management using Infrastructure as Code (IaC).
Collaborate with development, architecture, security, and product teams to promote reliability best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide, operating in markets such as financial services, healthcare, automotive, and insurance. The company has over 25,200 employees across 32 countries and is recognized as a Top 25 global workplace by Fortune.
Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.