Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, networking, storage, and workload schedulers.
Translate requirements from diverse customers into clear product direction and partner with engineering to define requirements.
Manage the observability backlog using feedback from production deployments and design partners to refine priorities.
Mirantis is a Kubernetes-native AI infrastructure company that enables organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI workloads. It is a distributed team committed to openness and technical excellence.
Lead a globally distributed observability team building and operating metrics, logging, alerting, and capacity planning platforms.
Set priorities with Site Reliability Engineering, Product Engineering, and GitLab Dedicated teams while owning reliability, scalability, and cost.
Participate in incident response and on-call rotations, using SLOs, error budgets, and AI tools to improve alerting and sustain operational load.
GitLab is an intelligent orchestration platform for DevSecOps that helps organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. With more than 50 million registered users and a high-performance culture driven by shared values and AI adoption, GitLab's globally distributed team collaborates to solve complex problems.
Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, and networking.
Design and implement solutions that improve availability, scalability, performance, and resilience of the platform.
Everbridge empowers enterprises and government organizations to anticipate, mitigate, respond to, and recover from critical events. The company focuses on building resilient systems and fostering a culture of ownership, continuous improvement, and operational excellence.
Own infrastructure as a product, defining platform strategy, backlog, and multi-quarter roadmap.
Drive self-service golden paths for engineers to build, deploy, and scale services.
Partner with engineering leadership to align priorities and improve CI/CD adoption.
Jobgether is an AI-powered job matching platform connecting candidates to partner companies. They use technology to streamline hiring, though company size and culture are not detailed in the posting.
Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.
Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.
Lead Cloud Platform and SRE teams to scale securely and efficiently.
Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
Champion SRE culture with SLOs, error budgets, and observability.
Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.
Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Set the vision and roadmap for the Tailscale dataplane, including WireGuard behavior, DERP relay expansion, subnet routers, exit nodes, and client-side networking.
Define performance and reliability as customer-facing product requirements and own connectivity product strategy end to end.
Partner closely with GTM, Customer Success, and Solutions Engineering on production traffic commitments and strategic enterprise deals.
Tailscale is making safe connection effortless by delivering software that securely interconnects people and devices. Founded in 2019 and fully distributed, the company is backed by Accel, CRV, Insight, Heavybit, and Uncork Capital.
Manage assigned technical projects, guiding teams on Agile/Scrum practices to delight clients.
Remove impediments and build a trusting environment for problem-solving without blame or retribution.
Plan and coordinate project activities, schedules, and budgets to ensure delivery on time and within scope.
AHEAD builds platforms for digital business by weaving together cloud infrastructure, automation, analytics, and software delivery. They prioritize creating a culture of belonging, are an equal opportunity employer, and value diverse perspectives.
Architect and build the observability platform for metrics, logs, traces, and events across global infrastructure.
Drive instrumentation with OpenTelemetry, building shared libraries and collector deployments for correlated signals.
Run observability as an internal product with published interfaces, versioned clients, and SLOs to ensure adoption.
Smartsheet empowers teams to manage work and scale solutions, uniting human teams with AI agents to automate tasks and uncover insights. With over 20 years of experience, the company fosters a collaborative, innovative culture focused on employee well-being and professional growth.
Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.
ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.
Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.
Jobgether is a platform that connects job seekers with opportunities through an AI-powered matching process. The company fosters a remote-first culture and emphasizes autonomy and ownership for engineers.
Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
Define SLOs and SLIs to drive architectural decisions and error budget policies.
Conduct blameless post-incident reviews and implement long-term preventive measures.
Airalo is the world's first eSIM store, helping travelers access affordable mobile data in 200+ countries. They are a fully remote team of 400+ people across 60+ countries, with a culture of trust, ownership, and freedom.