Lead and execute on complex technical troubleshooting and incident resolution from investigation to delivery of permanent solutions.
Design and implement comprehensive monitoring processes, including creating detailed playbooks and runbooks for common scenarios, as well as detailed documentation including post-incident reviews and knowledge base articles.
Leverage AI-powered tools and workflows to automate issue detection, diagnosis, and resolution processes.
Alternative Payments is building the financial operating system for SMBs, consolidating the disconnected tech stack that holds service-based businesses back. We're growing fast, thinking big, and building a global team that wants to be part of something that lasts.
Maintain observability platform and introduce observability on new projects.
Implement automated management features and configure solutions per security processes.
Manage CI systems and pipelines, and design and implement infrastructure.
Lingaro is a global technology company providing data, cloud, and DevOps solutions. With over 1,500 employees across 7 sites, they foster a diverse and inclusive culture.
Lead high-impact incident response in a complex cloud environment, shaping how Atlassian responds to security events at scale.
Manage a team of 5-8 US-based incident responders, coaching them in investigations, decision-making, and executive communication.
Drive adoption of investigative AI tools and automation to improve triage speed, signal relevance, and operational consistency.
Atlassian creates software products that help teams collaborate globally. The company values diversity and inclusion, with a culture focused on unleashing team potential.
Lead observability and monitoring operations integration and workflow support, including event-to-incident patterns and dashboard visualization.
Manage events and incidents tied to monitoring platforms, and support OpenTelemetry implementation and automation using Splunk SOAR/Ansible.
Provide operational reporting views and ensure hands-on experience with enterprise monitoring, observability, or APM engineering for technical and non-technical stakeholders.
Makpar is a comprehensive professional and technical solutions provider for the Federal government, combining cloud engineering, data management, cybersecurity, and emerging technologies. They are an Equal Opportunity Employer with a connected and engaged workforce dedicated to delivering mission success for government clients.
Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
Drive SLOs, observability, alerting, and on-call processes across teams.
Build the platform engineering function from the ground up and influence cross-cutting architecture.
First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.
Drive rapid delivery and iteration for the Core Services team of DoiT Cloud Intelligence.
Translate VP-level strategy into a sequenced backlog of problems, user stories, and acceptance criteria.
Leverage hands-on DevOps experience to understand cloud workflows and prioritize enhancements that improve speed, reliability, and simplicity.
DoiT is a global technology company that helps cloud-driven organizations leverage the cloud for business growth and innovation. They are an award-winning strategic partner of AWS, Google Cloud, and Microsoft Azure, working with over 4,000 customers worldwide, and pride themselves on a remote-first culture with flexibility and professional development.
Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
Manage distributed systems, observability, incident response, and automation with a security-first mindset.
Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.
Lead the design and operation of LivePerson's observability platforms across logs, metrics, traces, alerting, and synthetic monitoring.
Own large-scale observability pipelines using technologies like Elastic Cloud, Grafana, Prometheus, and Kafka.
Provide technical leadership and mentorship while driving best practices in DevOps, cloud engineering, and observability.
LivePerson is a leader in trusted enterprise conversational AI and digital transformation, powering nearly a billion conversational interactions every month. The company is recognized as the #1 Most Innovative AI Company by Fast Company and fosters a diverse, inclusive culture that empowers employees globally.
Evolve DevOps and platform engineering, including build pipelines, monitoring, infrastructure as code, and cost optimization.
Design and maintain secure, scalable CI/CD pipelines that support rapid iterations without sacrificing stability.
Improve observability via Datadog and GCP, implement FinOps governance, and enable developer productivity through internal tooling.
Sardine is the leading agentic risk platform for fighting financial crime, unifying data across risk teams to stop fraud in real time and prevent AI-driven attacks. It is a remote-first company with hubs in multiple locations, hiring talented self-motivated individuals who value performance over hours worked.
Manage the ticket queue, prioritize and resolve requests, and identify recurring categories for automation.
Participate in rotating on-call and incident response, troubleshooting and documenting issues in real time.
Build and maintain monitoring dashboards (Tableau, Superset, Grafana) to track service health and data quality.
Airbnb is a global community marketplace that connects hosts with guests for unique stays and experiences. With over 5 million hosts and 2 billion guest arrivals, the company fosters a culture of inclusion and belonging, emphasizing innovation and engagement.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
Define and drive SRE platform strategy, incident management, and observability engineering.
Mentor team members, foster collaboration, and ensure operational excellence.
XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.
Own infrastructure as code across development, staging, and production environments
Build, maintain, and improve CI/CD pipelines for reliable and efficient deployments
Manage cloud infrastructure, establish scalable engineering practices, and lead incident response
CelebriOS is a software company building B2B SaaS products that help businesses make better decisions and streamline operations. The company has a remote-first working environment and a benefits package designed to support their team.
Lead, mentor, and grow a team of SRE/DevOps engineers while partnering with engineering leadership to assess team needs and develop talent.
Oversee the incident management process end to end, including on-call rotations, escalation paths, incident command, postmortems, and root cause analysis.
Define and drive SRE principles like SLIs, SLOs, error budgets, capacity planning, and observability standards, championing a culture of reliability and operational excellence.
Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services, enabling community financial institutions to digitally engage in a secure and compliant way through a world-class digital communications platform. Their platform integrates Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology, bolstered by AI and contact center capabilities, and they value integrity, transparency, and ownership.
Operate the Monad node fleet, including health, sync, upgrades, and incident response for validators, full nodes, and archive nodes.
Own infrastructure-as-code with Ansible, Terraform, and Kubernetes, and build observability with Prometheus, Grafana, and Loki.
Design and build AI agent tooling for automated operations, including runbooks-as-code and deterministic guardrails.
Category Labs designs and builds decentralized technology, including the Monad blockchain, a high-performance EVM-compatible Layer 1. The team raised $225M in series A funding and is a lean, collaborative group of engineers and researchers with a culture of low ego and high-quality output.
Build and operate internal platform services and APIs in Go.
Codify infrastructure with Terraform and GitOps practices.
Operate and scale multi-tenant EKS clusters and traffic systems.
Docker builds tools for developers to build, share, and run applications, trusted by over 20 million monthly users. They are a globally distributed, remote-first team with offices in Seattle and Paris, focused on innovation and inclusion.
Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.
Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.
Own observability for critical product journeys, defining SLIs/SLOs and building metrics, dashboards, and alerts.
Act as first responder for production incidents, investigating signals and mitigating issues independently.
Work within a cross-functional squad of 6-8 engineers to improve reliability, monitoring, and incident response processes.
Feeld is a dating app creating a safer and more inclusive space for exploring relationships and sexuality. They have a distributed engineering team of around 50 people across Europe and the US, working in small autonomous squads.
Implement monitoring use cases under senior direction and support data onboarding, alert configuration, and incident workflow alignment.
Participate in troubleshooting and operational support activities, and assist in configuring and tuning dashboards, alerts, and monitoring rules.
Collaborate with team members to validate monitoring coverage across supported systems and escalate unresolved issues through ITSM channels.
Makpar is a comprehensive professional and technical solutions provider for the Federal government, specializing in cloud engineering, data management, cybersecurity, and emerging technologies. They have a connected and engaged workforce dedicated to delivering success for clients and the American people.