Source Job

US 24w maternity 24w paternity

  • Architect and build the observability platform for metrics, logs, traces, and events across global infrastructure.
  • Drive instrumentation with OpenTelemetry, building shared libraries and collector deployments for correlated signals.
  • Run observability as an internal product with published interfaces, versioned clients, and SLOs to ensure adoption.

Go Python OpenTelemetry AWS Kubernetes

20 jobs similar to Senior Software Engineer I (Observability Platform)

Jobs ranked by similarity.

EMEA

  • Design and operate scalable telemetry pipelines for metrics, logs, and traces across distributed GPU and edge infrastructure.
  • Architect and maintain telemetry storage systems optimized for large-scale time-series and event data.
  • Build comprehensive observability across compute, storage, networking, GPU clusters, and inference workloads.

Radian Arc builds and operates large-scale GPU cloud and edge infrastructure for AI workloads. The company is a fast-growing international scale-up with a focus on innovative infrastructure solutions, offering an inclusive and diverse working environment.

$104,000–$166,000/yr
US

  • Build and maintain telemetry pipelines across Dynatrace, Datadog, and Splunk for metrics, logs, and traces.
  • Design observability for distributed systems in AWS/GovCloud, including dashboards and golden-signal monitoring.
  • Build and tune alert definitions, support on-call rotations, and integrate with ServiceNow.

Peraton is a next-generation national security company that drives missions of consequence globally. They are a leading mission capability integrator and enterprise IT provider, serving essential government agencies and supporting the U.S. armed forces.

$139,200–$235,200/yr
Canada United States Unlimited PTO

  • Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment.
  • Improve deployment, monitoring, and operations using Kubernetes, Helm, Terraform, and cloud services from AWS or GCP.
  • Automate operational work, strengthen observability, and manage production issues to reduce single points of failure.

GitLab is the intelligent orchestration platform for DevSecOps, helping organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. Trusted by more than 50 million registered users and over 50% of the Fortune 100, GitLab fosters a high-performance culture driven by shared values and continuous knowledge exchange.

UK

  • Architect and build a robust, scalable, and highly available distributed infrastructure.
  • Build a cutting-edge cloud-native platform on top of the public cloud and automate cloud resource management.
  • Work closely with core database development and security teams to produce the SaaS offering.

ClickHouse is a real-time analytics and data warehousing company recognized on the Forbes Cloud 100 list. With over 4,000 customers and rapid growth, the company is a leader in its space.

North America 6w PTO

  • Develop tooling for cloud integrations and observability apps across the full stack.
  • Ship features end to end, from UI to backend, and contribute to open source.
  • Own projects from problem statement to production and join on-call.

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by 10,000+ organizations. We are a 100% remote company with 40+ countries, backed by leading investors and open-source values.

Germany Sweden Spain Ireland UK 6w PTO

  • Manage and develop a distributed team of backend and frontend engineers, providing regular feedback and supporting career growth.
  • Collaborate closely with go-to-market and engineering leadership to ensure seamless integration of Fleet Management, Kubernetes Helm Chart, and Instrumentation Hub.
  • Foster a psychologically safe environment that encourages learning, experimentation, and continuous improvement.

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations for reliability and telemetry optimization. As a 100% remote company with team members across 40+ countries, Grafana Labs fosters a global, collaborative culture that values transparency, autonomy, and meaningful work.

United States Canada 6w PTO

  • Build and ship end-to-end features for cloud integrations and observability applications, spanning frontend, backend, dashboards, and alerts.
  • Work with TypeScript/React, Go, and Jsonnet, owning projects from concept through production.
  • Contribute to open-source observability technologies like Prometheus and OpenTelemetry, and participate in on-call rotations.

Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. It operates with a remote-first, globally distributed team, and focuses on transparency and efficient hiring.

$180,000–$240,000/yr
US

  • Own large slices of the system end to end, from approach to operation.
  • Turn Beads into a platform and take Gas City to the cloud.
  • Define SLOs, observability, backups, and security baseline for enterprise readiness.

Gas City builds the open-source stack teams use to run coding agents at scale, including the Beads work graph and Gas City agent orchestration. It's a small, flat organization moving toward revenue with a focus on reliability and agent-driven development.

Poland

  • Design and develop highly performant backend services for real-time data processing and web APIs.
  • Define and own reliability objectives and error budgets for core API services.
  • Build observability through metrics, logging, tracing, and dashboards, and participate in on-call rotation.

The company provides technology that protects businesses and users from online fraud. It is a fully remote, globally distributed organization.

$157,000–$184,000/yr
US Canada UK Unlimited PTO 18w maternity 12w paternity

  • Build and maintain core components of the clearing house in Go on GCP, including customer onboarding flows and data ingestion pipelines.
  • Take ownership of ambiguous problems and specific features, driving them from design through production with appropriate testing and observability.
  • Design, implement, and operate reliable agentic workflows that automate complex, multi-step tasks across production systems.

Chainguard is the trusted source for open source, delivering hardened, secure, and production-ready builds of open source software. The company is venture-backed by leading investors and serves Fortune 500 enterprises, including OpenAI and Snowflake.

Canada USA Unlimited PTO

  • Own the observability, logging and alerting for Kubernetes clusters and critical workloads.
  • Build and maintain automation for lifecycle management of Kubernetes clusters.
  • Identify and root-fix reliability bottlenecks before they become incidents.

Wrapbook is an AI platform for production finance, built for feature films and TV, trusted by Netflix and Paramount. Backed by top investors, our team of over 350 employees uses AI to transform how finance teams work.

$145,000–$177,000/yr
US

  • Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
  • Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, and networking.
  • Design and implement solutions that improve availability, scalability, performance, and resilience of the platform.

Everbridge empowers enterprises and government organizations to anticipate, mitigate, respond to, and recover from critical events. The company focuses on building resilient systems and fostering a culture of ownership, continuous improvement, and operational excellence.

UK

  • Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
  • Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.

Global Unlimited PTO

  • Lead a globally distributed observability team building and operating metrics, logging, alerting, and capacity planning platforms.
  • Set priorities with Site Reliability Engineering, Product Engineering, and GitLab Dedicated teams while owning reliability, scalability, and cost.
  • Participate in incident response and on-call rotations, using SLOs, error budgets, and AI tools to improve alerting and sustain operational load.

GitLab is an intelligent orchestration platform for DevSecOps that helps organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. With more than 50 million registered users and a high-performance culture driven by shared values and AI adoption, GitLab's globally distributed team collaborates to solve complex problems.

Global

  • Design and improve the core orchestration engine for managing Supabase Branches lifecycle.
  • Provision ephemeral sandboxed execution environments on EKS/ECS for untrusted build workloads.
  • Implement job scheduling, pipeline optimization, and observability for thousands of concurrent builds.

Supabase is the Postgres development platform, built by developers for developers. They are a globally distributed team of ~400 members across 60+ countries, operating fully remote with an open-source-first culture.

$140,000–$160,000/yr
US

  • Manage assigned technical projects, guiding teams on Agile/Scrum practices to delight clients.
  • Remove impediments and build a trusting environment for problem-solving without blame or retribution.
  • Plan and coordinate project activities, schedules, and budgets to ensure delivery on time and within scope.

AHEAD builds platforms for digital business by weaving together cloud infrastructure, automation, analytics, and software delivery. They prioritize creating a culture of belonging, are an equal opportunity employer, and value diverse perspectives.

$150,000–$165,000/yr
US Unlimited PTO

  • Design and manage high-availability platforms using Kubernetes, Terraform, and Ansible with native-AI capabilities.
  • Develop and operate the observability stack: Grafana, Mimir, Loki, Tempo, and Prometheus on Kubernetes via GitLab CI/CD.
  • Build automation scripts in Python, maintain GitOps pipelines, and mentor mid-level engineers.

Flexential builds and operates critical IT platforms including observability, DevOps, and ITSM technologies. The company fosters a collaborative engineering culture and values diversity.

US

  • Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, networking, storage, and workload schedulers.
  • Translate requirements from diverse customers into clear product direction and partner with engineering to define requirements.
  • Manage the observability backlog using feedback from production deployments and design partners to refine priorities.

Mirantis is a Kubernetes-native AI infrastructure company that enables organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI workloads. It is a distributed team committed to openness and technical excellence.

$119,380–$165,100/yr
Spain UK

  • Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
  • Define SLOs and SLIs to drive architectural decisions and error budget policies.
  • Conduct blameless post-incident reviews and implement long-term preventive measures.

Airalo is the world's first eSIM store, helping travelers access affordable mobile data in 200+ countries. They are a fully remote team of 400+ people across 60+ countries, with a culture of trust, ownership, and freedom.

US Unlimited PTO

  • Build scalable back-end services for the next generation of applications using Kotlin and Java.
  • Lead code reviews and architectural discussions while mentoring other engineers on best practices.
  • Collaborate with product managers to solve distributed systems challenges and drive critical initiatives.

Smartsheet empowers teams to manage work seamlessly and scale solutions smarter, now uniting human teams with AI agents. With over 20 years of experience, they foster an inclusive culture that values diverse perspectives and supports professional growth.