Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Design, implement, and maintain highly available and scalable infrastructure solutions.
Monitor system performance, identify bottlenecks, and resolve reliability issues proactively.
Automate infrastructure deployment, configuration management, and operational workflows.
The company is a technology firm that provides critical authorization solutions to organizations worldwide. It is a remote-first organization with a collaborative culture, offering equity opportunities and a focus on team building.
Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.
They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.
Design, implement, and maintain scalable and reliable systems.
Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
Develop and maintain automation tools for deployment, monitoring, and system health checks.
LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.
Design, build, and run distributed cloud architectures and large-scale production systems.
Ensure reliability, observability, performance, and cost efficiency of the platform.
Collaborate with product and backend teams to design system architecture and optimize resource use.
Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.
You will own and deliver quarterly goals for your team, leading engineers through ambiguity to solve open-ended problems.
You will proactively identify technical solutions and operational processes that strengthen incident readiness and response.
You will foster a culture of quality and ownership by setting or improving code review and design standards.
Affirm is reinventing credit to make it more honest and friendly, offering consumers the flexibility to buy now and pay later. The company has a strong engineering culture focused on reliability and ownership.
Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.
MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.
Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.
The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.
Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
Partnering with development teams to establish production readiness and operational readiness.
Building tooling to automate observability and operational workflows, eliminating manual toil.
Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.
Empower engineers on other teams by maintaining monitoring tooling and collaborating on observability best practices.
Enhance reliability of Kubernetes applications through resource optimization, streamlined upgrades, and scalability.
Participate in on-call and incident response processes, occasionally diving into application code to debug production issues.
Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. It serves over 2 million users worldwide across 190 countries, with tens of thousands of projects launched each month, and fosters a culture of grit, speed, and craft.
Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Lead and mentor a US-based team of Site Reliability Engineers, driving operational excellence across production platforms.
Serve as an escalation point and incident commander for major production incidents, ensuring timely triage and resolution.
Drive automation, reliability, and observability improvements using Datadog, Kubernetes, and CI/CD pipelines.
NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions that partners with managed care organizations. Recognized as one of the fastest-growing companies in America with multiple locations in the US, South America, and India, they offer a fulfilling work environment that encourages associates to contribute to delivering premier service.
Design, implement, maintain, and optimize highly available infrastructure supporting mission-critical applications and services.
Monitor production environments, analyze system performance, and proactively identify opportunities to improve stability, scalability, and operational efficiency.
Respond to technical escalations, troubleshoot infrastructure, networking, hardware, and software issues, and lead resolution of critical incidents.
Our partner is a technology company focused on high-availability platforms and mission-critical infrastructure. The team is collaborative and works with modern cloud technologies.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
Establish an SRE function, building a team of 4-6 and professionalizing incident management with clear processes and tooling.
Drive engineering excellence through design reviews, code reviews, and blameless retrospectives to foster a quality culture.
Lead people and projects, balancing incident response with a roadmap of observability and reliability engineering initiatives.
We are a fast-growing, remote cybersecurity company dedicated to enabling organizations to proactively find and fix exploitable attack vectors. Our team is a fusion of former special operations cyber operators and engineers, committed to a culture of respect, collaboration, and ownership.
Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
Drive automation and eliminate operational toil through self-service tooling and process improvements.
Lead incident response and mentor engineers to improve reliability practices.
Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.
Build systems for declarative application and infrastructure lifecycle management, including CI/CD, Kubernetes, and service inventory.
Prioritize and troubleshoot infrastructure issues to minimize downtime and respond to alerts efficiently.
Contribute to setting the SRE team's direction and streamline automation of infrastructure processes.
Counterpart Health develops Counterpart Assistant, an AI-enabled primary care tool that supports physicians in chronic disease management. It is a subsidiary of Clover Health, with a remote-first culture and a focus on value-based care through technology.
Design and implement scalable cloud infrastructure to support growth.
Develop monitoring, alerting, and incident response for system reliability.
Automate deployment pipelines and ensure high availability and security.
Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Work with teams to define SLIs and SLOs, and create systems for observability.
Analyze failure scenarios, create runbooks, and reduce work that does not add value.
Participate in incident management and facilitate on-call duty to ensure reliable production environments.
Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.