Lead, mentor, and grow a team of SRE/DevOps engineers while partnering with engineering leadership to assess team needs and develop talent.
Oversee the incident management process end to end, including on-call rotations, escalation paths, incident command, postmortems, and root cause analysis.
Define and drive SRE principles like SLIs, SLOs, error budgets, capacity planning, and observability standards, championing a culture of reliability and operational excellence.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
Drive automation and eliminate operational toil through self-service tooling and process improvements.
Lead incident response and mentor engineers to improve reliability practices.
Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.
Lead and mentor a US-based team of Site Reliability Engineers, driving operational excellence across production platforms.
Serve as an escalation point and incident commander for major production incidents, ensuring timely triage and resolution.
Drive automation, reliability, and observability improvements using Datadog, Kubernetes, and CI/CD pipelines.
NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions that partners with managed care organizations. Recognized as one of the fastest-growing companies in America with multiple locations in the US, South America, and India, they offer a fulfilling work environment that encourages associates to contribute to delivering premier service.
Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
Drive SLOs, observability, alerting, and on-call processes across teams.
Build the platform engineering function from the ground up and influence cross-cutting architecture.
First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.
Ensure production system reliability, scalability, and performance through automation and AI-driven operations.
Design and implement AIOps workflows for event ingestion, anomaly detection, and incident response.
Collaborate with engineering teams to embed reliability into the software lifecycle and improve system resilience.
Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. The company is a partner-focused recruitment service with a remote-first, globally distributed team.
Lead Cloud Platform and SRE teams, driving infrastructure strategy and ownership including Kubernetes, GCP, and Terraform.
Champion SRE culture, define SLOs, SLAs, and enhance observability and incident management.
Own the developer-facing platform as a product, ensuring self-service infrastructure and security compliance (SOC2, ISO-27001).
Prolific builds human data infrastructure for AI development, connecting researchers with a global participant pool. They foster a culture of high performance, ownership, and cross-functional collaboration.
Lead the SRE strategy and execution for a high-growth AI company.
Build and scale a high-performing SRE team while defining reliability standards.
Architect secure, scalable cloud infrastructure and implement observability practices.
This company develops advanced AI products and agentic technology. It operates in a high-growth, international environment with a focus on operational excellence and innovation.
Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.
Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.
Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
Scale single-tenant deployments and build observability, incident response, and compliance practices.
Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.
Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.
Own the reliability of Supabase's deployment and release systems against clear SLOs and error budgets.
Standardize and instrument pre-production deployment workflows for trustworthy signal.
Drive disaster-recovery readiness and reduce mean-time-to-detect and recover for deploy-related incidents.
Supabase is the Postgres development platform providing a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. They are a globally distributed team of ~400 members across 60+ countries, with over $1B raised and 540,000+ community members.
Work with teams to define SLIs and SLOs, and create systems for observability.
Analyze failure scenarios, create runbooks, and reduce work that does not add value.
Participate in incident management and facilitate on-call duty to ensure reliable production environments.
Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.
Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
Lead a team of five DevOps engineers, driving self-service infrastructure and progressive delivery across product and platform teams.
Champion the evolution of deployment patterns on Kubernetes, including blue/green, canary, and feature-flag releases to minimize risk.
Mentor and develop engineers while staying hands-on with Terraform, Kubernetes, and production incident response.
Loop builds a connected commerce operations suite that helps merchants manage returns, exchanges, order tracking, and fraud prevention. Trusted by over 5,000 brands, the company fosters a culture of high empathy and high standards, where employees grow quickly and shape the future of commerce.
You will own and deliver quarterly goals for your team, leading engineers through ambiguity to solve open-ended problems.
You will proactively identify technical solutions and operational processes that strengthen incident readiness and response.
You will foster a culture of quality and ownership by setting or improving code review and design standards.
Affirm is reinventing credit to make it more honest and friendly, offering consumers the flexibility to buy now and pay later. The company has a strong engineering culture focused on reliability and ownership.
Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.
Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.
Design, build, and optimize cloud infrastructure (AWS/Kubernetes/EKS) and CI/CD pipelines across multiple teams.
Troubleshoot and resolve production incidents of varying scope, ensuring reliability and performance.
Drive infrastructure projects end-to-end, mentor engineers, and establish standards that improve developer productivity.
Pacvue is a leading Commerce Media OS powering over $12B in advertising spend across 100+ global retail media networks. It enables over 70,000 brands and agencies with an inclusive global community that fosters innovation and career growth.
You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.
Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.
Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
Drive AI-specific observability, FinOps, and security practices across the platform.
We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.
Build and operate the self-service infrastructure platform where developers and agents can validate changes in minutes.
Build golden paths for CI/CD, GitOps, and IaC to enable self-service provisioning and shipping.
Own reliability and observability, carrying on-call and turning recurring toil into automation.
Luxury Presence is building the AI growth platform for real estate. Backed by Bessemer Venture Partners, the company is a Series C firm with over 90,000 real estate professionals and has been ranked on the Inc. 5000 fastest-growing companies list three years in a row.