Lead and execute on complex technical troubleshooting and incident resolution from investigation to delivery of permanent solutions.
Design and implement comprehensive monitoring processes, including creating detailed playbooks and runbooks for common scenarios, as well as detailed documentation including post-incident reviews and knowledge base articles.
Leverage AI-powered tools and workflows to automate issue detection, diagnosis, and resolution processes.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Design and maintain AWS infrastructure using Terraform, with a focus on scalability cost and PCI-scoped network segmentation
Build and evolve the observability stack and CI/CD pipelines to ensure smooth production operations and rapid deployment
Lead incident response define SLOs and run performance tests to optimize payment-critical services
Xplor Technologies provides vertical software, embedded payments, and AI tools for membership-based and service-based industries. With over 130,000 businesses in 72+ countries and processing $47 billion in payments annually, the company values diversity, collaboration, and a people-first culture.
Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
Partnering with development teams to establish production readiness and operational readiness.
Building tooling to automate observability and operational workflows, eliminating manual toil.
Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.
Maintain and improve production and staging infrastructure for high availability and scalability. - Design, build, and optimize CI/CD pipelines ensuring reliable deployments. - Automate infrastructure provisioning using Infrastructure as Code and improve monitoring and alerting.
They are an AI-driven payment platform processing millions of transactions across 80+ payment methods including cryptocurrency. Their team of 80+ professionals works in a hybrid format across multiple offices and remotely globally.
Build and operate the self-service infrastructure platform where developers and agents can validate changes in minutes.
Build golden paths for CI/CD, GitOps, and IaC to enable self-service provisioning and shipping.
Own reliability and observability, carrying on-call and turning recurring toil into automation.
Luxury Presence is building the AI growth platform for real estate. Backed by Bessemer Venture Partners, the company is a Series C firm with over 90,000 real estate professionals and has been ranked on the Inc. 5000 fastest-growing companies list three years in a row.
Ensure availability, performance, scalability, and resilience of production services in AWS.
Automate infrastructure provisioning and management using Infrastructure as Code (IaC).
Collaborate with development, architecture, security, and product teams to promote reliability best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide, operating in markets such as financial services, healthcare, automotive, and insurance. The company has over 25,200 employees across 32 countries and is recognized as a Top 25 global workplace by Fortune.
Design and operate the infrastructure for a high-throughput messaging platform operating at 500K+ events/sec.
Build guardrails, runbooks, and validation gates that enable AI agents to safely execute deployments and operations.
Lead incident response and encode every fix as a new runbook and regression test.
Postscript is an AI messaging platform trusted by 20,000+ Shopify brands to drive revenue through SMS. The company is fully remote, backed by Greylock and Y Combinator, and has a culture of ownership and innovation.
Design and implement monitoring and alerting systems using tools like Prometheus, Grafana, and DataDog to ensure high availability and reliability.
Optimize performance and reliability of healthcare payment applications, lead incident response, and develop SLOs/SLIs.
Automate CI/CD pipelines, infrastructure provisioning with Terraform, and manage cloud infrastructure on AWS with Kubernetes.
LMI is a digital solutions provider accelerating government impact with innovation and speed, bringing commercial-grade platforms and mission-ready AI to federal agencies. Headquartered in Tysons, Virginia, LMI serves the defense, space, healthcare, and energy sectors, focusing on agility and collaboration to drive impactful results.
Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.
MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.
Own production operability by debugging complex issues, improving system visibility, and eliminating recurring problems at the source.
Improve mean time to detect and resolve issues, and reduce recurrence rates through code fixes and architecture improvements.
Partner with product teams to feed production learnings back into design and development, reducing support and incident load.
One Identity helps organizations secure, manage, and analyze information and infrastructure to drive innovation. With a global team, we offer a collaborative environment where employees build products at scale and grow their careers.
Ensure production system reliability, scalability, and performance through automation and AI-driven operations.
Design and implement AIOps workflows for event ingestion, anomaly detection, and incident response.
Collaborate with engineering teams to embed reliability into the software lifecycle and improve system resilience.
Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. The company is a partner-focused recruitment service with a remote-first, globally distributed team.
Own the end-to-end major incident lifecycle, acting as Incident Commander for critical events.
Drive reliability metrics and improve MTTR for Yuno's 99.99% uptime target.
Run blameless postmortems and translate findings into actionable reliability improvements.
Yuno builds payment infrastructure for global market participation, enabling companies to integrate over 1,000 payment methods via a single API. They empower high-performing teams and use advanced AI for smart routing and fraud prevention across 80+ countries.
Ensure system architecture meets technical requirements by collaborating with IT teams (Architecture, Security, Infrastructure).
Maintain and evolve the microservices environment on AWS with a focus on information security.
Implement DevOps practices, automation, and monitoring tools to ensure system reliability and scalability.
Experian is a global data and technology company that powers opportunities for people and businesses worldwide. With 25,200 employees across 32 countries, it fosters a people-centric, inclusive culture recognized as a World's Best Workplace.
Work with teams to define SLIs and SLOs, and create systems for observability.
Analyze failure scenarios, create runbooks, and reduce work that does not add value.
Participate in incident management and facilitate on-call duty to ensure reliable production environments.
Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.
Manage the ticket queue, prioritize and resolve requests, and identify recurring categories for automation.
Participate in rotating on-call and incident response, troubleshooting and documenting issues in real time.
Build and maintain monitoring dashboards (Tableau, Superset, Grafana) to track service health and data quality.
Airbnb is a global community marketplace that connects hosts with guests for unique stays and experiences. With over 5 million hosts and 2 billion guest arrivals, the company fosters a culture of inclusion and belonging, emphasizing innovation and engagement.
Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.
Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.
Report to the Manager of Customer Reliability and provide senior level technical support on customer issues.
Work with development engineering to track escalations, bugs, and feature requests.
Develop technical troubleshooting sessions, trainings, and maintain internal knowledge base articles.
Sysdig is a cloud security company that stops attacks in real-time using runtime insights and open source Falco. It is a well-funded startup with a large enterprise customer base, recognized as a "Best Places to Work" fostering an inclusive and diverse culture.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.