Lead the design and operation of LivePerson's observability platforms across logs, metrics, traces, alerting, and synthetic monitoring.
Own large-scale observability pipelines using technologies like Elastic Cloud, Grafana, Prometheus, and Kafka.
Provide technical leadership and mentorship while driving best practices in DevOps, cloud engineering, and observability.
LivePerson is a leader in trusted enterprise conversational AI and digital transformation, powering nearly a billion conversational interactions every month. The company is recognized as the #1 Most Innovative AI Company by Fast Company and fosters a diverse, inclusive culture that empowers employees globally.
Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, networking, storage, and workload schedulers.
Translate requirements from diverse customers into clear product direction and partner with engineering to define requirements.
Manage the observability backlog using feedback from production deployments and design partners to refine priorities.
Mirantis is a Kubernetes-native AI infrastructure company that enables organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI workloads. It is a distributed team committed to openness and technical excellence.
Empower engineers on other teams by maintaining monitoring tooling and collaborating on observability best practices.
Enhance reliability of Kubernetes applications through resource optimization, streamlined upgrades, and scalability.
Participate in on-call and incident response processes, occasionally diving into application code to debug production issues.
Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. It serves over 2 million users worldwide across 190 countries, with tens of thousands of projects launched each month, and fosters a culture of grit, speed, and craft.
Work collaboratively with a team to create and maintain the foundational platform for Reddit's infrastructure.
Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
Contribute upstream changes to open source projects and share on-call responsibilities.
Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information, employing a flexible-first workforce that values open-source contributions.
Design and operate scalable telemetry pipelines for metrics, logs, and traces across distributed GPU and edge infrastructure.
Architect and maintain telemetry storage systems optimized for large-scale time-series and event data.
Build comprehensive observability across compute, storage, networking, GPU clusters, and inference workloads.
Radian Arc builds and operates large-scale GPU cloud and edge infrastructure for AI workloads. The company is a fast-growing international scale-up with a focus on innovative infrastructure solutions, offering an inclusive and diverse working environment.
Build and maintain telemetry pipelines across Dynatrace, Datadog, and Splunk for metrics, logs, and traces.
Design observability for distributed systems in AWS/GovCloud, including dashboards and golden-signal monitoring.
Build and tune alert definitions, support on-call rotations, and integrate with ServiceNow.
Peraton is a next-generation national security company that drives missions of consequence globally. They are a leading mission capability integrator and enterprise IT provider, serving essential government agencies and supporting the U.S. armed forces.
Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.
They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.
Drive technical strategy and roadmap for adaptive telemetry databases.
Lead end-to-end delivery of large, cross-functional projects.
Own architecture, reliability, performance, and cost for critical systems.
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations. We are a 100% remote company with team members across 40+ countries, backed by leading investors, and we foster a collaborative, open-source culture.
Develop tooling for cloud integrations and observability apps across the full stack.
Ship features end to end, from UI to backend, and contribute to open source.
Own projects from problem statement to production and join on-call.
Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by 10,000+ organizations. We are a 100% remote company with 40+ countries, backed by leading investors and open-source values.
You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.
Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.
Define SLIs, SLOs, and reliability targets for the platform.
Improve observability, alerting, and production readiness across services.
Automate operational work and support cloud/Kubernetes infrastructure.
Lodgify is a fast-growing scale-up in vacation rental technology, backed by $30M in funding. Headquartered in Barcelona, the 380+ person team of 60+ nationalities is passionate about transforming short-term rentals.
Lead escalation triage and incident response for complex customer support cases, coordinating with engineering and product teams to drive resolution.
Build and maintain application observability using Datadog monitors, dashboards, and alerts to ensure platform health and proactive response.
Translate field experience into documentation, authoring troubleshooting guides and runbooks, and mentor Tier 1/2 support staff.
Dragos is the global leader in xOT cybersecurity, protecting critical infrastructure systems that deliver water, power, and keep hospitals running. The company is a remote-first mission-driven team across North America, Europe, Middle East, and APAC, built on authenticity, transparency, and trust.
Apply SRE principles to improve reliability, scalability, and performance of production systems.
Design and implement automation to reduce operational toil and improve engineering efficiency.
Lead incident response and develop sustainable solutions for complex production issues.
The hiring company is a technology organization focused on reliability and operational excellence. They offer a fully remote, collaborative environment with opportunities for technical leadership and career growth.
Design and implement a scalable observability platform for Whatnot's growing infrastructure.
Work with core infrastructure, platform, and developer tools teams to redesign data collection to visualization.
Utilize AI agents and open standards to ensure visibility into software stack performance and reliability.
Whatnot is the largest live shopping platform in North America and Europe, enabling sellers to build businesses across hundreds of categories. They are a remote co-located team anchored in hubs across the US, UK, Ireland, Poland, Germany, and Australia, and were recently named the #1 Best Startup Employer in America by Forbes.