Manage the ticket queue, prioritize and resolve requests, and identify recurring categories for automation.
Participate in rotating on-call and incident response, troubleshooting and documenting issues in real time.
Build and maintain monitoring dashboards (Tableau, Superset, Grafana) to track service health and data quality.
Airbnb is a global community marketplace that connects hosts with guests for unique stays and experiences. With over 5 million hosts and 2 billion guest arrivals, the company fosters a culture of inclusion and belonging, emphasizing innovation and engagement.
Design and build automated reliability and self-healing systems to protect production at scale.
Own and improve incident management tooling and on-call health, reducing alert noise and empowering teams.
Develop and evolve observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection.
Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for improved safety, efficiency, and sustainability. As a public company with over 2.3 million IoT devices deployed, we foster a culture of growth mindset and collaboration.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Design the human side of reliability: on-call rotations, incident roles, blameless postmortems, and change management.
Build and operate the core reliability toolkit: observability, CI/CD, infrastructure as code, and incident response.
Embed with product engineers to raise operational literacy and define SLOs grounded in patient needs.
Novellia is a health tech startup that provides patients with access to their health records and enables biopharma innovation through real-world data. They are growing 5x year over year, have raised close to $30M in funding, and are backed by tier-1 investors, with a focus on building a diverse and inclusive workplace.
Work with teams to define SLIs and SLOs, and create systems for observability.
Analyze failure scenarios, create runbooks, and reduce work that does not add value.
Participate in incident management and facilitate on-call duty to ensure reliable production environments.
Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.
Own the reliability of Supabase's deployment and release systems against clear SLOs and error budgets.
Standardize and instrument pre-production deployment workflows for trustworthy signal.
Drive disaster-recovery readiness and reduce mean-time-to-detect and recover for deploy-related incidents.
Supabase is the Postgres development platform providing a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. They are a globally distributed team of ~400 members across 60+ countries, with over $1B raised and 540,000+ community members.
Partner with Technical Support leadership to develop strategic narratives and executive presentations that communicate business performance and operational outcomes.
Build and maintain reporting frameworks for core metrics, synthesize data into actionable insights for senior leadership.
Drive cross-functional collaboration across teams to align on priorities and execute key initiatives, while leveraging AI tools for efficiency.
Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data to improve safety, efficiency, and sustainability. As a recently public company, they foster a high-growth environment with a culture of rapid career development and innovation.
Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.
MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.
Own production operability by debugging complex issues, improving system visibility, and eliminating recurring problems at the source.
Improve mean time to detect and resolve issues, and reduce recurrence rates through code fixes and architecture improvements.
Partner with product teams to feed production learnings back into design and development, reducing support and incident load.
One Identity helps organizations secure, manage, and analyze information and infrastructure to drive innovation. With a global team, we offer a collaborative environment where employees build products at scale and grow their careers.
Own outcomes for Git & Gitaly Operations, Nonlinear Productivity, and Platform Staff.
Set the technical bar across all three functions by reviewing distributed systems designs and crisis management plans.
Build operating rhythm and metrics to ensure predictable outcomes without relying on one-person heroics.
GitLab is a DevSecOps platform that enables organizations to improve developer productivity and security. With over 50 million users and trust from the Fortune 100, it has a high-performance culture driven by values.
Redesign payment processes, onboard AI vendors, and build predictive models to improve operations.
Spearhead quality initiatives and manage teams within operations workflows.
Solve high-impact problems across customer experience, product, and business with curiosity and persistence.
Clipboard operates an app-based marketplace connecting healthcare professionals with workplaces needing workers. Founded in 2016, it is a remote-first team of over 1,000 people, profitable since 2022, and is a top Y-Combinator company.
Manage and resolve the most challenging issues for the SRE team, focusing on instance performance, reliability, and availability.
Use software development and systems engineering experience to proactively prevent issues and drive improvements in infrastructure reliability.
Drive a culture of automation and scalable solutions, collaborating with partner teams to enhance system design.
ServiceNow provides an AI control tower for business reinvention, integrating any AI, data, and workflow to help 85% of the Fortune 500 work smarter, faster, and better. The company fosters an AI-native culture, combining technology and talent to drive innovation and growth.
Own observability end to end and define how we measure reliability.
Own CI/CD pipelines and make shipping fast and safe.
Footprint builds Percy, an AI agent that runs financial crime investigations end to end. The company is backed by QED, Index, and other investors, and its small, senior team ships fast and grew revenue 5x in the past year.
Frame the problem: clarify ambiguous product and system problems with product.
Design and ship: the extension and capture pipeline across Meet, Zoom, and Teams.
Build the tooling: turn work you repeat into internal tools and checks.
Tactiq is an AI meeting intelligence platform that captures and analyzes meeting transcripts from Google Meet, Zoom, and Teams. It is a small, fast-growing company with a product-led growth motion, aiming for $100M ARR.
Design and deploy a modern observability stack including logging, metrics, and distributed tracing across hundreds of services.
Build alerting policies and incident response workflows to reduce manual escalations and improve mean time to detect and resolve.
Automate toil and set SRE standards while mentoring engineers on observability tooling.
WeightWatchers is a global digital health company and the world's #1 doctor-recommended behavioral weight health program. They have a cloud architecture supporting hundreds of services and are expanding their platform engineering team.
Own the end-to-end major incident lifecycle, acting as Incident Commander for critical events.
Drive reliability metrics and improve MTTR for Yuno's 99.99% uptime target.
Run blameless postmortems and translate findings into actionable reliability improvements.
Yuno builds payment infrastructure for global market participation, enabling companies to integrate over 1,000 payment methods via a single API. They empower high-performing teams and use advanced AI for smart routing and fraud prevention across 80+ countries.
Design and implement monitoring and alerting systems using tools like Prometheus, Grafana, and DataDog to ensure high availability and reliability.
Optimize performance and reliability of healthcare payment applications, lead incident response, and develop SLOs/SLIs.
Automate CI/CD pipelines, infrastructure provisioning with Terraform, and manage cloud infrastructure on AWS with Kubernetes.
LMI is a digital solutions provider accelerating government impact with innovation and speed, bringing commercial-grade platforms and mission-ready AI to federal agencies. Headquartered in Tysons, Virginia, LMI serves the defense, space, healthcare, and energy sectors, focusing on agility and collaboration to drive impactful results.
Investigate and solve complex business problems by translating ambiguous questions into structured analyses and actionable recommendations.
Lead cross-functional initiatives from problem definition through implementation, measurement, and refinement, partnering with Operations, Clinical, Product, and Data teams.
Develop reporting infrastructure and dashboards to transform recurring analyses into scalable, repeatable insights.
Form Health is a virtual obesity medicine clinic delivering evidence-based care for obesity and cardiometabolic conditions. Founded in 2019, it is a venture-backed company with a patient-first culture and a focus on scaling access to care.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Architect the Operations Intelligence Warehouse, leading the cleanup and design of operational data in BigQuery.
Set data modeling standards with dbt, building dimensional models with testing, lineage, and version control.
Instrument the operation with dashboards and metrics for real-time visibility into throughput, cycle time, and yield.
Collective is on a mission to redefine the way businesses-of-one work, providing an integrated platform for business incorporation, accounting, bookkeeping, and tax services. Backed by General Catalyst, Sound Ventures, and other prominent investors, the company empowers self-employed individuals to achieve tax savings and financial independence.