Own the end-to-end major incident lifecycle, acting as Incident Commander for critical events.
Drive reliability metrics and improve MTTR for Yuno's 99.99% uptime target.
Run blameless postmortems and translate findings into actionable reliability improvements.
Yuno builds payment infrastructure for global market participation, enabling companies to integrate over 1,000 payment methods via a single API. They empower high-performing teams and use advanced AI for smart routing and fraud prevention across 80+ countries.
Build and maintain observability across the platform in Datadog, including dashboards, monitors, APM, and log pipelines.
Participate in on-call rotation and incident response, driving blameless post-incident reviews and automating toil.
Leverage AI tools to accelerate debugging, generate runbooks, and build automation for operational efficiency.
IPSY is a beauty subscription platform that connects brands and consumers through curated beauty products. It is a remote-first company with a focus on community and engagement.
Own and evolve observability strategy including monitoring, alerting, dashboards, logging, and distributed tracing.
Define and manage SLIs, SLOs, and reliability metrics, improving MTTD and MTTR through automation.
Build and maintain reliable cloud infrastructure on AWS and Kubernetes while mentoring engineers on SRE best practices.
Filevine is a Legal AI company delivering Legal Operating Intelligence for legal work. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Oversee IT incidents from detection to resolution, coordinating cross-functional teams to minimize disruption.
Maintain accurate documentation, analyze trends, and ensure compliance with SLAs.
Facilitate post-incident reviews and drive process improvements to strengthen operational resilience.
Jobgether uses AI-powered matching to connect candidates with hiring companies, focusing on efficiency and fairness. They operate as a platform that processes applications and shares top-fitting candidates with employers, emphasizing innovation and continuous improvement.
Design and build automated reliability and self-healing systems to protect production at scale.
Own and improve incident management tooling and on-call health, reducing alert noise and empowering teams.
Develop and evolve observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection.
Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for improved safety, efficiency, and sustainability. As a public company with over 2.3 million IoT devices deployed, we foster a culture of growth mindset and collaboration.
Lead end-to-end management of major production incidents, coordinating cross-functional teams from detection to resolution.
Drive improvements in operational reliability by reducing MTTD, MTTR, and optimizing on-call programs.
Facilitate blameless postmortems and establish incident management standards including severity frameworks and runbooks.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They focus on innovation and operational excellence within a globally distributed engineering organization.
Design, develop, and enhance proactive monitoring capabilities for the AWS Connect CCaaS platform to ensure system reliability and operational excellence.
Troubleshoot production issues, perform root cause analysis, and implement corrective actions to minimize system downtime and service disruptions.
Collaborate closely with developers, architects, and platform owners to enforce logging standards and monitoring best practices that enable effective troubleshooting and observability.
Miratech is a global IT services and consulting company that helps visionaries change the world by supporting digital transformation for large enterprises. With nearly 1000 full-time professionals and a culture of Relentless Performance, they achieve over 99% project success rate and operate in over 25 countries.
Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.
MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.
Lead, mentor, and develop a team of engineers and administrators managing enterprise productivity platforms.
Oversee 24x7x365 operational support, including on-call rotations and escalation management.
Act as senior technical escalation point for complex incidents and drive platform enhancements.
Cision is a global leader in consumer and media intelligence, engagement, and communication solutions. The company serves over 75,000 organizations, including 84% of the Fortune 500, and is committed to fostering an inclusive environment.
Establish an SRE function, building a team of 4-6 and professionalizing incident management with clear processes and tooling.
Drive engineering excellence through design reviews, code reviews, and blameless retrospectives to foster a quality culture.
Lead people and projects, balancing incident response with a roadmap of observability and reliability engineering initiatives.
We are a fast-growing, remote cybersecurity company dedicated to enabling organizations to proactively find and fix exploitable attack vectors. Our team is a fusion of former special operations cyber operators and engineers, committed to a culture of respect, collaboration, and ownership.
Lead and coordinate major incident war rooms from initiation through service restoration, driving cross-functional collaboration.
Provide timely executive-level updates and maintain accurate incident documentation, leveraging AI tools for efficiency.
Participate in a 24x7 on-call rotational support model for enterprise production incidents.
Miratech is a global IT services and consulting company that helps visionaries change the world by supporting digital transformation for large enterprises. With nearly 1,000 full-time professionals and a 99% project success rate, they maintain a culture of relentless performance and operate in 25 countries worldwide.
Monitor and support 24x7 production operations for enterprise API gateway environments, including incident triage and service restoration.
Troubleshoot and document production issues, perform root cause analysis, and update runbooks in ServiceNow.
Validate monitoring, patching, and post-change readiness to ensure reliable gateway performance.
Acuity, Inc. is a management and technology consulting firm that supports federal agencies with IT modernization, data enablement, and hyperautomation. The company has been recognized as a Best Place to Work for over 9 years and fosters a people-first culture with diverse and inclusive environment.
Respond to inbound requests and alerts for cloud hosted solutions, developing case management skills with oversight.
Interact with customers, partners, and internal resources to understand needs and ensure professional communication.
Troubleshoot and resolve basic cloud issues using company systems, contributing to documentation and maintaining high availability.
Hyland is the pioneer of the Content Innovation Cloud, delivering enterprise intelligence and automation. With nearly 4,000 employees, the company fosters an inclusive, employee-centric culture focused on growth and community impact.
Support and maintain AWS-based CCaaS contact center environments to ensure high availability and system reliability.
Monitor, troubleshoot, and resolve production issues using observability tools like Splunk and Zabbix.
Collaborate with developers and architects to enforce logging standards and enhance monitoring capabilities.
Miratech is a global IT services and consulting company that helps visionaries change the world by supporting digital transformation for large enterprises and startups. With nearly 1000 full-time professionals and a culture of Relentless Performance, the company has a 99% project success rate and operates in 25 countries across 5 continents.
Design and deploy a modern observability stack including logging, metrics, and distributed tracing across hundreds of services.
Build alerting policies and incident response workflows to reduce manual escalations and improve mean time to detect and resolve.
Automate toil and set SRE standards while mentoring engineers on observability tooling.
WeightWatchers is a global digital health company and the world's #1 doctor-recommended behavioral weight health program. They have a cloud architecture supporting hundreds of services and are expanding their platform engineering team.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Architect, build, and evolve secure, scalable cloud infrastructure on AWS and AWS GovCloud to power the Hypori SaaS platform.
Independently own ambiguous, high-impact infrastructure problems, guide technical direction, and act as a senior escalation point during production incidents.
Drive strategy and execution of Infrastructure as Code, observability, CI/CD, and operational frameworks while mentoring engineers and raising the technical bar across the organization.
Hypori Inc. is a high-growth cybersecurity SaaS company transforming secure mobility through a virtual workspace platform that enables users to access enterprise apps and data from any mobile device with zero data on the endpoint and total personal privacy. Backed by $55M in funding from investors including UBS, AE Industrial Partners, Hale Capital Partners, and GreatPoint Ventures, the company is expanding into new commercial and regulated markets.
Configure and maintain Dynatrace dashboards, reports, and alerts to support Salesforce application monitoring.
Develop visibility into user actions and key application performance events to improve operational insight.
Set up, configure, and maintain synthetic monitoring for critical Salesforce transactions and business workflows.
VetsEZ is a company that provides technology solutions, supporting a large federal government healthcare modernization project. They are an equal opportunity employer with a focus on mission-critical systems and offer training opportunities.
Drive the definition and adoption of SLIs and SLOs across services, reducing toil through automation and incident response.
Design and architect Infrastructure as Code solutions for large-scale environments using Docker, Kubernetes, and cloud-native services.
Serve as primary SRE liaison for development teams, influencing architecture and conducting training for clients.
Noctua Technology, LLC is a company that drives digital transformation by treating operations as a software engineering challenge, focusing on cloud native systems. They are a dynamic team seeking a Senior SRE to define strategy and bridge development and operations for clients.
Lead corporate IT and enterprise security operations, including system architecture and budget management.
Partner with DevOps to ensure security requirements extend to the product platform.
Manage external SOC providers and MDR platforms across the corporate fleet.
VIA is a digital infrastructure company that provides mission-critical organizations with speed and security, specializing in agentic AI, zero-trust identity, quantum-resistant data, and offline stablecoin payments. The company is fast-paced and technology-focused, with a flat, non-traditional hierarchy emphasizing continuous coaching and cross-functional collaboration.