Source Job

US Unlimited PTO

  • Design, build, and operate distributed systems that ingest, process, and store telemetry at very high scale.
  • Own the reliability, performance, capacity, and cost-efficiency of telemetry pipelines and storage systems.
  • Participate in the on-call rotation, help resolve production incidents, and drive root-cause fixes through to completion.

Go Kubernetes Terraform AWS OpenTelemetry

20 jobs similar to Cloud Software Engineer - Observability Platform

Jobs ranked by similarity.

$87,480–$110,160/yr
Europe 6w PTO

  • Design, build, and operate reconciliation systems for Grafana Cloud stacks at scale.
  • Collaborate across teams to improve reliability, deployment complexity, and incident response.
  • Contribute to roadmap planning, technical design, and long-term simplification of stack operations.

Grafana Labs is the company behind the open source observability platform Grafana, providing a fully managed observability cloud. With over 1,600 team members across 40+ countries, the company fosters a global, collaborative culture rooted in open source principles.

US Unlimited PTO

  • Architect and build robust, scalable, and highly available distributed infrastructure.
  • Build a cutting-edge cloud-native platform on public cloud and automate resource management.
  • Optimize cost efficiency systematically through distributed systems and industry best practices.

ClickHouse is a private cloud company that provides real-time analytics, data warehousing, observability, and AI workloads. With over 4,000 customers and rapid growth, the company fosters a collaborative, globally distributed culture.

$217,000–$303,900/yr
US Unlimited PTO

  • Work collaboratively with a team to create and maintain the foundational platform for Reddit's infrastructure.
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
  • Contribute upstream changes to open source projects and share on-call responsibilities.

Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information, employing a flexible-first workforce that values open-source contributions.

Global

  • Design and develop a highly available, scalable, and secure ClickHouse Cloud platform for regulated environments.
  • Build innovative deployment automation across cloud, hybrid, and on-prem systems, including disconnected environments.
  • Collaborate with Security, Dataplane, and Infrastructure teams to ensure compliance with NIST and FedRAMP frameworks.

ClickHouse is a fast-growing private cloud company offering real-time analytics, data warehousing, observability, and AI workloads. With over 3,000 customers and a $400M Series D, the company fosters a culture of innovation and collaboration.

$200,000–$240,000/yr
US Unlimited PTO 16w maternity 16w paternity

  • You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.

Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.

US 6w PTO

  • Take an active role in influencing our roadmap and your own career objectives.
  • Design, build, operate, and maintain critical systems, owning reliability, performance, and availability.
  • Collaborate with your team to deliver new features and iterate based on results.

Grafana Labs is the company behind the open-source observability platform Grafana, providing a fully managed observability cloud. With over 1,600 team members across 40+ countries and 35 million users, the company thrives on a transparent, collaborative, and open-source culture.

Europe 6w PTO

  • Own features end-to-end across the stack, from backend Go services to frontend TypeScript/React.
  • Design and evolve systems for scheduling, probe lifecycle, and telemetry ingestion at scale.
  • Build cross-product and AI-assisted workflows that integrate Synthetic Monitoring into the Grafana Cloud platform.

Grafana Labs is the company behind the open source observability cloud, helping organizations see and act on their data. With over 1,600 employees across 40+ countries, the company fosters an open, collaborative culture.

$176,000–$231,000/yr
US Unlimited PTO

  • Lead the design and delivery of business-critical systems powering customer acquisition, onboarding, billing, and revenue operations.
  • Drive architectural excellence for high-throughput, fault-tolerant metering and billing pipelines.
  • Mentor engineers and foster a culture of ownership, learning, and pragmatic excellence.

Temporal is an open source programming model that simplifies code, making applications more reliable and developers more productive. It is a growing, values-driven company with a collaborative culture focused on quality and pragmatism.

US

  • Define and evolve enterprise observability vision, standards, and roadmap.
  • Lead implementation of Dynatrace SaaS and platform capabilities.
  • Design observability solutions for cloud-native infrastructure.

FreedomPay provides commerce solutions. They are a large company with a culture focused on innovation and collaboration.

Europe US

  • Design and architect observability solutions leveraging OpenTelemetry, Kubernetes, and cloud-native technologies.
  • Develop and execute Proofs of Concept (POCs) that highlight Dash0's differentiated technical capabilities.
  • Deliver engaging technical demos and presentations tailored to engineering and executive audiences.

Dash0 is building an OpenTelemetry-native observability platform that eliminates vendor lock-in and provides transparent pricing. Backed by top-tier investors including Balderton Capital, Accel and Cherry Ventures, the company has a collaborative, fast-moving team culture with a builder mindset.

APAC

  • Design and deliver significant components and core subsystems of our Kubernetes platform, such as secrets management, workload identity, storage, or cluster networking, from design through production operation.
  • Contribute to the architecture of distributed workloads, working with dependent teams to get runtime and isolation models right, while spending most time hands-on in code.
  • Own operability of built systems including SLOs, failure modes, upgrades, migrations, and on-call, and mentor earlier-career engineers.

ServiceNow is the AI control tower for business reinvention, bringing together any AI, any data, and any workflow to help 85% of the Fortune 500 work smarter, faster, and better. The company fosters an AI-native culture where technology and talent are unstoppable together, with a focus on freeing people from busywork.

$151,000–$206,000/yr
US Unlimited PTO

  • Help build large scale, real-time services and applications leveraging massive datasets to power internal APIs and external applications.
  • Build tooling, libraries, frameworks, and services that support security, research and data platform initiatives.
  • Participate in planning and technical discussions with engineering and product teams to build the right solutions.

Censys provides real-time Internet intelligence and actionable threat insights to global governments, over 50% of the Fortune 500, and leading threat intelligence providers worldwide. They are a distributed team with roots in Ann Arbor, Michigan, and are committed to creating an inclusive environment.

Argentina

  • Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
  • Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
  • Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.

Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.

$163,000–$263,670/yr
US

  • Design, build, and operate the distributed systems that deliver feature-flag configuration to LaunchDarkly SDKs.
  • Own and improve the reliability, latency, and scalability of our streaming and polling infrastructure.
  • Debug and resolve complex production issues while sharing the team's on-call rotation.

LaunchDarkly builds a feature management platform that enables safe and gradual software releases. The company is growing and fosters a humble, open, collaborative culture.

Canada

  • Design, build, and enhance the infrastructure powering a large-scale real-time data ingestion platform.
  • Scale backend services to support increasing data volumes while maintaining high availability and performance.
  • Automate infrastructure provisioning and deployment pipelines using Infrastructure as Code.

Our partner builds a high-performance real-time data platform operating at petabyte scale. They are a remote-first organization with a collaborative engineering culture.

US Unlimited PTO

  • Architect and operate distributed, fault-tolerant systems for a large-scale cloud data platform.
  • Lead optimization initiatives across compute, storage, networking, and infrastructure efficiency.
  • Collaborate with engineering teams and stakeholders to build tooling for cloud cost and resource visibility.

This company builds a highly scalable cloud data platform and focuses on infrastructure efficiency and multi-cloud environments. With operations across more than 20 countries, it fosters a globally distributed, remote-friendly culture that values ownership, collaboration, and innovation.

$235,000–$275,000/yr

  • Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.

Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

Global 7w PTO

  • Lead the design and operation of LivePerson's observability platforms across logs, metrics, traces, alerting, and synthetic monitoring.
  • Own large-scale observability pipelines using technologies like Elastic Cloud, Grafana, Prometheus, and Kafka.
  • Provide technical leadership and mentorship while driving best practices in DevOps, cloud engineering, and observability.

LivePerson is a leader in trusted enterprise conversational AI and digital transformation, powering nearly a billion conversational interactions every month. The company is recognized as the #1 Most Innovative AI Company by Fast Company and fosters a diverse, inclusive culture that empowers employees globally.

$160,000–$200,000/yr
US Unlimited PTO

  • Deliver production-ready Temporal implementations by co-building workflows with customer and partner engineering teams.
  • Identify and remove early activation blockers and establish operational standards.
  • Define observability, reliability, and deployment strategies for production systems.

Temporal is an open source programming model that simplifies code, makes applications more reliable, and helps developers focus on delivering features faster. The company is mission-driven, building a team that values curiosity, drive, collaboration, genuineness, and humility.