Take an active role in influencing our roadmap and your own career objectives.
Design, build, operate, and maintain critical systems, owning reliability, performance, and availability.
Collaborate with your team to deliver new features and iterate based on results.
Grafana Labs is the company behind the open-source observability platform Grafana, providing a fully managed observability cloud. With over 1,600 team members across 40+ countries and 35 million users, the company thrives on a transparent, collaborative, and open-source culture.
Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.
The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.
Lead the design and operation of LivePerson's observability platforms across logs, metrics, traces, alerting, and synthetic monitoring.
Own large-scale observability pipelines using technologies like Elastic Cloud, Grafana, Prometheus, and Kafka.
Provide technical leadership and mentorship while driving best practices in DevOps, cloud engineering, and observability.
LivePerson is a leader in trusted enterprise conversational AI and digital transformation, powering nearly a billion conversational interactions every month. The company is recognized as the #1 Most Innovative AI Company by Fast Company and fosters a diverse, inclusive culture that empowers employees globally.
Design, write, and deliver software (primarily in Go and Python) to improve availability, scalability, latency, and efficiency of Reddit's products.
Dive deep into the codebase of Go services and the Python monolith legacy stack to implement complex system-level improvements.
Collaborate with cross-functional teams and share on-call responsibilities to ensure reliability and scalability of Tier-0 services.
Reddit is a community of communities, built on shared interests, passion, and trust. With over 100,000 active communities and approximately 130 million daily active users, it is one of the internet's largest sources of information.
You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.
Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.
Maintain observability platform and introduce observability on new projects.
Implement automated management features and configure solutions per security processes.
Manage CI systems and pipelines, and design and implement infrastructure.
Lingaro is a global technology company providing data, cloud, and DevOps solutions. With over 1,500 employees across 7 sites, they foster a diverse and inclusive culture.
Empower engineers on other teams by maintaining monitoring tooling and collaborating on observability best practices.
Enhance reliability of Kubernetes applications through resource optimization, streamlined upgrades, and scalability.
Participate in on-call and incident response processes, occasionally diving into application code to debug production issues.
Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. It serves over 2 million users worldwide across 190 countries, with tens of thousands of projects launched each month, and fosters a culture of grit, speed, and craft.
Design, build, and operate reconciliation systems for Grafana Cloud stacks at scale.
Collaborate across teams to improve reliability, deployment complexity, and incident response.
Contribute to roadmap planning, technical design, and long-term simplification of stack operations.
Grafana Labs is the company behind the open source observability platform Grafana, providing a fully managed observability cloud. With over 1,600 team members across 40+ countries, the company fosters a global, collaborative culture rooted in open source principles.
Serve as the primary technical point of contact for a portfolio of Grafana customers, designing and guiding their observability maturity journey.
Conduct regular technical reviews, health checks, and root cause analysis to drive adoption and ensure customer success.
Act as the voice of the customer internally, shaping product feedback and roadmap priorities while building long-term strategic relationships.
Grafana Labs builds the open source observability platform Grafana and its fully managed cloud service. The company has over 1,600 team members across 40+ countries, serving more than 7,000 customers including major enterprises, and fosters a remote-first, transparent, and innovation-driven culture.
Help design, build, and operate the Kubernetes platform used across PulsePoint.
Own reliability, observability, and incident response across platform services.
Build infrastructure automation and GitOps workflows to reduce operational toil.
PulsePoint sits at the intersection of healthcare and adtech, helping brands interpret health signals using real-world data. With over 300 employees, the company is a post-acquisition profitable leader in the US healthcare ad market, known for a flat hierarchy and high engineering bar.
Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.
Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.
Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Design and develop a highly available, scalable, and secure ClickHouse Cloud platform for regulated environments.
Build innovative deployment automation across cloud, hybrid, and on-prem systems, including disconnected environments.
Collaborate with Security, Dataplane, and Infrastructure teams to ensure compliance with NIST and FedRAMP frameworks.
ClickHouse is a fast-growing private cloud company offering real-time analytics, data warehousing, observability, and AI workloads. With over 3,000 customers and a $400M Series D, the company fosters a culture of innovation and collaboration.
Support the deployment, operation, and maintenance of the Karuna service running on Kubernetes.
Monitor production environments to ensure high availability, reliability, and performance.
Investigate, troubleshoot, and resolve production incidents, performing root cause analysis.
Software Mind develops solutions that make an impact for companies around the globe. They build cross-functional engineering teams with a culture of openness, respect, grit, and enjoyment.
Design, build, and operate data and analytics systems supporting advanced model development workflows.
Own the complete lifecycle of experiment data, including modeling, ingestion, retrieval, and analysis.
Develop high-performance, scalable data infrastructure supporting increasing scale and reliability requirements.
The company focuses on building foundational data and analytics infrastructure for advanced AI research and development. They have a globally distributed team with a people-first culture focused on collaboration and innovation.
Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Shape the future of Reddit by adapting platforms to evolving privacy, security, and regulatory landscapes.
Partner with product, design, and engineering teams to build trusted, compliant experiences for millions of users.
Own product areas end-to-end from technical design to launch, influencing technical and product strategy.
Reddit is a community of communities built on shared interests and authentic conversations. With over 100,000 active communities and approximately 130 million daily active users, Reddit is one of the largest sources of information on the internet.
Design and architect observability solutions leveraging OpenTelemetry, Kubernetes, and cloud-native technologies.
Develop and execute Proofs of Concept (POCs) that highlight Dash0's differentiated technical capabilities.
Deliver engaging technical demos and presentations tailored to engineering and executive audiences.
Dash0 is building an OpenTelemetry-native observability platform that eliminates vendor lock-in and provides transparent pricing. Backed by top-tier investors including Balderton Capital, Accel and Cherry Ventures, the company has a collaborative, fast-moving team culture with a builder mindset.
Serve as a trusted technical advisor guiding customers through their observability journey.
Design and guide customer observability maturity strategies to improve reliability and operational visibility.
Provide expert troubleshooting and technical recommendations to resolve complex challenges.
Jobgether is a platform that uses AI to match candidates with jobs. They focus on remote work and have a collaborative culture built around transparency, autonomy, and trust.