Source Job

US

  • Design, build, and operate high-scale observability pipelines for logs, metrics, traces, and exceptions.
  • Lead cross-functional initiatives to resolve scaling bottlenecks and evolve production infrastructure safely.
  • Partner with engineering teams to improve observability tools and provide technical leadership across teams.

Go Ruby Kubernetes Azure OpenTelemetry

20 jobs similar to Senior Software Engineer, Observability Delivery

Jobs ranked by similarity.

US 24w maternity 24w paternity

  • Architect and build the observability platform for metrics, logs, traces, and events across global infrastructure.
  • Drive instrumentation with OpenTelemetry, building shared libraries and collector deployments for correlated signals.
  • Run observability as an internal product with published interfaces, versioned clients, and SLOs to ensure adoption.

Smartsheet empowers teams to manage work and scale solutions, uniting human teams with AI agents to automate tasks and uncover insights. With over 20 years of experience, the company fosters a collaborative, innovative culture focused on employee well-being and professional growth.

US Unlimited PTO

  • Lead a two-month observability maturity assessment across metrics, logs, and traces at massive scale.
  • Drive consolidation to AWS-native observability on OpenTelemetry, including pipelines, dashboards, and alerts.
  • Act as Pod Leader: set technical direction, stay hands-on, and own the customer relationship.

EverOps is the premier Embedded Service Provider, partnering directly with customer engineering teams to assess and address mission-critical infrastructure, cloud, and delivery challenges. We've been remote since day one, and our culture centers on ownership, technical leadership, and continuous professional growth.

US

  • Design, build, and operate shared platform foundations including GCP, Kubernetes, networking, CI/CD, and observability.
  • Diagnose and troubleshoot complex distributed systems running at high request volume.
  • Raise the reliability bar through dashboards, alerting, on-call readiness, and automation.

Sanity.io builds an AI-powered content operating system that helps teams model, create, and automate content. The company has 200+ employees and a positive, flexible, trust-based culture that supports growth and work-life balance.

$139,200–$235,200/yr
Canada United States Unlimited PTO

  • Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment.
  • Improve deployment, monitoring, and operations using Kubernetes, Helm, Terraform, and cloud services from AWS or GCP.
  • Automate operational work, strengthen observability, and manage production issues to reduce single points of failure.

GitLab is the intelligent orchestration platform for DevSecOps, helping organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. Trusted by more than 50 million registered users and over 50% of the Fortune 100, GitLab fosters a high-performance culture driven by shared values and continuous knowledge exchange.

$180,000–$240,000/yr
US

  • Own large slices of the system end to end, from approach to operation.
  • Turn Beads into a platform and take Gas City to the cloud.
  • Define SLOs, observability, backups, and security baseline for enterprise readiness.

Gas City builds the open-source stack teams use to run coding agents at scale, including the Beads work graph and Gas City agent orchestration. It's a small, flat organization moving toward revenue with a focus on reliability and agent-driven development.

$168,675–$229,900/yr
US

  • Establish and employ continuous integration and delivery (CI/CD) patterns for successful software solutions.
  • Design secure, operationally sound solutions across AWS, Azure, OpenShift, and IBM Cloud.
  • Implement observability stacks, manage Kubernetes clusters, and automate infrastructure with Terraform.

Conga unifies commercial operations by aligning pricing, quoting, contracting, rebates, and communications so companies run as connected, smarter businesses. With more than 10,000 customers worldwide, including over 50% of the Fortune 100, Conga fosters a collaborative culture where every voice is heard.

United States Canada 6w PTO

  • Build and ship end-to-end features for cloud integrations and observability applications, spanning frontend, backend, dashboards, and alerts.
  • Work with TypeScript/React, Go, and Jsonnet, owning projects from concept through production.
  • Contribute to open-source observability technologies like Prometheus and OpenTelemetry, and participate in on-call rotations.

Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. It operates with a remote-first, globally distributed team, and focuses on transparency and efficient hiring.

Global Unlimited PTO

  • Lead a globally distributed observability team building and operating metrics, logging, alerting, and capacity planning platforms.
  • Set priorities with Site Reliability Engineering, Product Engineering, and GitLab Dedicated teams while owning reliability, scalability, and cost.
  • Participate in incident response and on-call rotations, using SLOs, error budgets, and AI tools to improve alerting and sustain operational load.

GitLab is an intelligent orchestration platform for DevSecOps that helps organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. With more than 50 million registered users and a high-performance culture driven by shared values and AI adoption, GitLab's globally distributed team collaborates to solve complex problems.

Germany Sweden Spain Ireland UK 6w PTO

  • Manage and develop a distributed team of backend and frontend engineers, providing regular feedback and supporting career growth.
  • Collaborate closely with go-to-market and engineering leadership to ensure seamless integration of Fleet Management, Kubernetes Helm Chart, and Instrumentation Hub.
  • Foster a psychologically safe environment that encourages learning, experimentation, and continuous improvement.

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations for reliability and telemetry optimization. As a 100% remote company with team members across 40+ countries, Grafana Labs fosters a global, collaborative culture that values transparency, autonomy, and meaningful work.

$170,000–$235,000/yr
US

  • Design and implement backend services for licensing, entitlements, feature access, and usage limits across NodeZero's product and APIs.
  • Build and evolve provisioning, admin experience, MSP/MSSP capabilities, and audit logging for a multi-tenant SaaS platform.
  • Operate production services with monitoring, incident response, and a high bar for design quality and test coverage.

Horizon3 is a fast-growing, remote cybersecurity company that helps organizations proactively find, fix, and verify exploitable attack vectors through its NodeZero autonomous pentesting platform. The team is a fusion of former special operations cyber operators and startup engineers, fostering a culture of respect, collaboration, ownership, and results.

UK

  • Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
  • Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.

$135,000–$170,000/yr
US Unlimited PTO

  • Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
  • Build and own the observability layer for the fleet, surfacing data in shared dashboards and alerting.
  • Design and build automated recovery and self-healing for production systems.

Climavision rebuilds climate technology with a high-resolution weather radar and satellite network to reduce coverage gaps and improve forecasting. Backed by The Rise Fund, they are headquartered in Louisville, KY with R&D in Raleigh, NC, operating with a startup culture.

Canada USA Unlimited PTO

  • Own the observability, logging and alerting for Kubernetes clusters and critical workloads.
  • Build and maintain automation for lifecycle management of Kubernetes clusters.
  • Identify and root-fix reliability bottlenecks before they become incidents.

Wrapbook is an AI platform for production finance, built for feature films and TV, trusted by Netflix and Paramount. Backed by top investors, our team of over 350 employees uses AI to transform how finance teams work.

US

  • Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, networking, storage, and workload schedulers.
  • Translate requirements from diverse customers into clear product direction and partner with engineering to define requirements.
  • Manage the observability backlog using feedback from production deployments and design partners to refine priorities.

Mirantis is a Kubernetes-native AI infrastructure company that enables organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI workloads. It is a distributed team committed to openness and technical excellence.

  • Own the technical direction and architecture of the releases platform, including the canonical release model.
  • Lead key technical decisions around security, reliability, scalability, and release governance.
  • Build and scale tools that support large-scale engineering teams and critical software delivery workflows.

This company builds a foundational release control platform used by leading engineering teams to ship software safely and efficiently. It fosters a collaborative, inclusive, and remote-first culture, with significant autonomy and opportunities for technical leadership.

$175,000–$250,000/yr
US 2w maternity 2w paternity

  • Design, develop, and maintain internal software, services, and automation using Go and Rust.
  • Build and operate Kubernetes-based infrastructure and improve developer workflows and CI/CD.
  • Collaborate across teams to solve ambiguous technical challenges and improve system reliability.

Our partner is a high-growth technology organization building internal platforms to enable efficient engineering. They operate with small, autonomous teams in a high-trust, collaborative remote environment with a focus on technical excellence.

$207,000–$256,000/yr
Australia 6w PTO

  • Serve as primary technical contact for strategic customers, guiding observability maturity and architecting solutions.
  • Combine hands-on troubleshooting with roadmap development, training, and stakeholder engagement.
  • Influence product direction by bringing customer feedback to internal teams.

The company specializes in observability solutions, helping organizations build modern monitoring practices. It is a remote-first organization with a global team, emphasizing autonomy, collaboration, and sustainable work-life balance.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

Global Unlimited PTO

  • Architect and build robust, scalable, and highly available distributed infrastructure.
  • Build a cutting-edge cloud-native platform on public cloud and automate resource management.
  • Improve reliability, security, and cost efficiency of cloud services.

ClickHouse builds a real-time analytics database platform and manages ClickHouse Cloud data plane end-to-end with compute, networking, and security. As a rapidly scaling global startup, the company operates across 25+ countries and fosters a flexible, remote-friendly culture with equity and healthcare benefits.

$150,000–$165,000/yr
US Unlimited PTO

  • Design and manage high-availability platforms using Kubernetes, Terraform, and Ansible with native-AI capabilities.
  • Develop and operate the observability stack: Grafana, Mimir, Loki, Tempo, and Prometheus on Kubernetes via GitLab CI/CD.
  • Build automation scripts in Python, maintain GitOps pipelines, and mentor mid-level engineers.

Flexential builds and operates critical IT platforms including observability, DevOps, and ITSM technologies. The company fosters a collaborative engineering culture and values diversity.