Manage and develop a distributed team of backend and frontend engineers, providing regular feedback and supporting career growth.
Collaborate closely with go-to-market and engineering leadership to ensure seamless integration of Fleet Management, Kubernetes Helm Chart, and Instrumentation Hub.
Foster a psychologically safe environment that encourages learning, experimentation, and continuous improvement.
Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations for reliability and telemetry optimization. As a 100% remote company with team members across 40+ countries, Grafana Labs fosters a global, collaborative culture that values transparency, autonomy, and meaningful work.
Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.
Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.
Lead the Data & Storage Reliability Engineering org to improve reliability, scalability, and customer experience.
Build and scale teams focused on observability, performance, diagnostics, automation, and prevention.
Partner with product, database, and operations teams to drive systemic improvements and reduce customer impact.
ServiceNow is an AI platform company that helps businesses reinvent workflows by combining AI, data, and automation. It serves 85% of the Fortune 500 and fosters an AI-native culture focused on innovation and talent.
Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
Define SLOs and SLIs to drive architectural decisions and error budget policies.
Conduct blameless post-incident reviews and implement long-term preventive measures.
Airalo is the world's first eSIM store, helping travelers access affordable mobile data in 200+ countries. They are a fully remote team of 400+ people across 60+ countries, with a culture of trust, ownership, and freedom.
Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.
ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.
Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, networking, storage, and workload schedulers.
Translate requirements from diverse customers into clear product direction and partner with engineering to define requirements.
Manage the observability backlog using feedback from production deployments and design partners to refine priorities.
Mirantis is a Kubernetes-native AI infrastructure company that enables organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI workloads. It is a distributed team committed to openness and technical excellence.
Attract, recruit, ramp, and mentor a team of great Presales SE's from diverse backgrounds and experiences.
Executive sponsorship of key accounts and thought leadership in the region.
Own and report on quarter-over-quarter team cadence, helping prospects evaluate our software through discovery, demos, and technical validation.
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. We are a 100% remote company with team members across 40+ countries, backed by leading investors.
Lead the transformation of a diverse operations-heavy organization into a modern, AI-first Production Engineering function.
Own end-to-end reliability, performance, scalability, and security of NICE's global cloud, telecom, and datacenter platforms.
Drive adoption of software-first operational practices including automated recovery, infrastructure as code, and observability.
NICE provides software products used by 25,000+ global businesses to deliver extraordinary customer experiences, fight financial crime, and ensure public safety. With over 8,500 employees across 30+ countries, the company fosters a culture of ambition, game-changing innovation, and high standards.
Lead Cloud Platform and SRE teams to scale securely and efficiently.
Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
Champion SRE culture with SLOs, error budgets, and observability.
Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.
Architect and build the observability platform for metrics, logs, traces, and events across global infrastructure.
Drive instrumentation with OpenTelemetry, building shared libraries and collector deployments for correlated signals.
Run observability as an internal product with published interfaces, versioned clients, and SLOs to ensure adoption.
Smartsheet empowers teams to manage work and scale solutions, uniting human teams with AI agents to automate tasks and uncover insights. With over 20 years of experience, the company fosters a collaborative, innovative culture focused on employee well-being and professional growth.
Be accountable for the performance of two distributed engineering squads, setting clear expectations, coaching, and developing each team member.
Own hiring and team planning across both squads, aligning capacity with product priorities and long-term needs.
Lead discovery with squads, engage with customers and community, and shape roadmaps using feedback and evidence.
Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by more than 10,000 organizations. They are a 100% remote company with team members across 40+ countries, backed by leading investors.
Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Lead and manage a team of high-performing Solution Engineers across the East region.
Drive key opportunities alongside VP of Sales and develop scalable processes and playbooks.
Act as a cultural flag carrier, embodying collaboration and partnership across AE and SE roles.
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations. We are a 100% remote company with team members across 40+ countries, fostering a culture of transparency, autonomy, and trust.
Act as a subject matter expert for service management tools and practices, creating content across various mediums.
Partner with product engineering teams to build demos and coach internal teams on communication.
Interface with open source communities and contribute to the product through feedback, documentation, or code.
Datadog is the leading observability and security platform for the AI era, providing unified visibility across the technology stack. Trusted by Fortune 500 companies and high-growth AI leaders, Datadog fosters a culture of innovation and collaboration.
Lead and develop a globally distributed Core Infrastructure team to evolve compute, networking, storage, and cloud systems.
Own technical strategy and improve reliability, capacity management, and cloud efficiency across the platform.
Build automation, self-service infrastructure, and leverage AI tools to accelerate engineering execution and reduce toil.
Samsara is the pioneer of the Connected Operations Cloud, helping organizations harness IoT data to improve safety, efficiency, and sustainability. As a public company processing over 25 trillion data points annually, it fosters a growth-minded, customer-focused culture with a globally distributed team.
Manage assigned technical projects, guiding teams on Agile/Scrum practices to delight clients.
Remove impediments and build a trusting environment for problem-solving without blame or retribution.
Plan and coordinate project activities, schedules, and budgets to ensure delivery on time and within scope.
AHEAD builds platforms for digital business by weaving together cloud infrastructure, automation, analytics, and software delivery. They prioritize creating a culture of belonging, are an equal opportunity employer, and value diverse perspectives.
Set the technical direction and roadmap for Platform engineering, owning the reliability, performance, and cost of Traild's GCP infrastructure.
Improve developer experience by enhancing CI/CD pipelines, build tooling, observability, and reducing friction for other engineers.
Lead a small team of platform engineers, staying hands-on while growing and supporting your team members.
Traild is a high-growth SaaS company redefining how finance teams operate by combining AI, automation, and payments infrastructure for B2B finance. They are a rapidly growing global team with a strong culture, as evidenced by an eNPS score of 78.
Lead the Agent Tools, Agent Observability, and Runner Execution teams across frontend, backend, and AI engineering.
Partner with product and UX to turn ambiguous problems into shipped milestones and make prioritization tradeoffs explicit.
Contribute substantively to technical design for running agents safely and effectively at scale.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and accelerate digital transformation. With over 50 million registered users and more than 50% of the Fortune 100 trusting them, GitLab fosters a high-performance culture driven by values, continuous knowledge exchange, and AI integration.
You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.
Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.
Design and operate scalable telemetry pipelines for metrics, logs, and traces across distributed GPU and edge infrastructure.
Architect and maintain telemetry storage systems optimized for large-scale time-series and event data.
Build comprehensive observability across compute, storage, networking, GPU clusters, and inference workloads.
Radian Arc builds and operates large-scale GPU cloud and edge infrastructure for AI workloads. The company is a fast-growing international scale-up with a focus on innovative infrastructure solutions, offering an inclusive and diverse working environment.