Source Job

APAC

  • Design, build, and maintain software, APIs, and automation to enhance platform reliability and observability.
  • Support monitoring, reliability, and continuous improvement in Kubernetes-based environments with a focus on Datadog.
  • Integrate observability into CI/CD pipelines and automate operational tasks using scripting languages like Python.

Python Kubernetes Datadog AWS CI/CD

20 jobs similar to Senior Software Engineer / SRE (Observability Focus)

Jobs ranked by similarity.

Canada Europe

  • Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
  • Partnering with development teams to establish production readiness and operational readiness.
  • Building tooling to automate observability and operational workflows, eliminating manual toil.

Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.

US

  • Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
  • Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
  • Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.

The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.

Europe 6w PTO

  • Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
  • Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
  • Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.

Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.

$200,000–$215,000/yr
US

  • Design and deploy a modern observability stack including logging, metrics, and distributed tracing across hundreds of services.
  • Build alerting policies and incident response workflows to reduce manual escalations and improve mean time to detect and resolve.
  • Automate toil and set SRE standards while mentoring engineers on observability tooling.

WeightWatchers is a global digital health company and the world's #1 doctor-recommended behavioral weight health program. They have a cloud architecture supporting hundreds of services and are expanding their platform engineering team.

Global

  • Manage and optimize multi-cloud infrastructure (AWS required, GCP optional) with Kubernetes and CI/CD pipelines.
  • Improve observability through monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Coralogix).
  • Drive automation and Infrastructure as Code (IaC) using Terraform and Helm, and provide architectural guidance.

NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior. In 2023, NIQ combined with GfK, bringing together two industry leaders with operations in 100+ markets and covering more than 90% of the world's population.

US

  • Design and implement monitoring and alerting systems using tools like Prometheus, Grafana, and DataDog to ensure high availability and reliability.
  • Optimize performance and reliability of healthcare payment applications, lead incident response, and develop SLOs/SLIs.
  • Automate CI/CD pipelines, infrastructure provisioning with Terraform, and manage cloud infrastructure on AWS with Kubernetes.

LMI is a digital solutions provider accelerating government impact with innovation and speed, bringing commercial-grade platforms and mission-ready AI to federal agencies. Headquartered in Tysons, Virginia, LMI serves the defense, space, healthcare, and energy sectors, focusing on agility and collaboration to drive impactful results.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

India

  • Design, implement, and manage scalable, highly available systems on Azure Cloud, optimizing Kubernetes and containerized workloads.
  • Build and maintain robust CI/CD pipelines with GitHub Actions and implement infrastructure as code using Helm Charts.
  • Monitor system performance, troubleshoot issues, ensure uptime, and perform root cause analysis to improve reliability.

Resilinc is pioneering intelligent, autonomous systems that redefine supply chain risk management using agentic AI, trusted by top companies in life sciences, aerospace, high tech, and automotive. We are a fully remote, mission-led team with a collaborative culture focused on high-impact work.

Europe US

  • Design and architect observability solutions leveraging OpenTelemetry, Kubernetes, and cloud-native technologies.
  • Develop and execute Proofs of Concept (POCs) that highlight Dash0's differentiated technical capabilities.
  • Deliver engaging technical demos and presentations tailored to engineering and executive audiences.

Dash0 is building an OpenTelemetry-native observability platform that eliminates vendor lock-in and provides transparent pricing. Backed by top-tier investors including Balderton Capital, Accel and Cherry Ventures, the company has a collaborative, fast-moving team culture with a builder mindset.

$190,000–$225,000/yr
US

  • Design, build, and operate core cloud infrastructure on AWS, including compute, networking, and container orchestration.
  • Own the CI/CD platform used across engineering teams, including build pipelines, environment promotion, and progressive rollout.
  • Build and maintain the observability stack across the organization, including logging, metrics, distributed tracing, and alerting.

RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company is an Equal Opportunity and Affirmative Action employer committed to diversity and collaboration.

$62,640–$104,760/yr
Europe 4w PTO

  • Design, build, and run distributed cloud architectures and large-scale production systems.
  • Ensure reliability, observability, performance, and cost efficiency of the platform.
  • Collaborate with product and backend teams to design system architecture and optimize resource use.

Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.

$154,700–$208,000/yr
US Canada

  • Design and build automated reliability and self-healing systems to protect production at scale.
  • Own and improve incident management tooling and on-call health, reducing alert noise and empowering teams.
  • Develop and evolve observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection.

Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for improved safety, efficiency, and sustainability. As a public company with over 2.3 million IoT devices deployed, we foster a culture of growth mindset and collaboration.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.

Global

  • Ensure availability, performance, scalability, and resilience of production services in AWS.
  • Automate infrastructure provisioning and management using Infrastructure as Code (IaC).
  • Collaborate with development, architecture, security, and product teams to promote reliability best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide, operating in markets such as financial services, healthcare, automotive, and insurance. The company has over 25,200 employees across 32 countries and is recognized as a Top 25 global workplace by Fortune.

US 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
  • Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.

Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.

APAC Singapore Hong Kong

  • Design and maintain highly available cloud infrastructure across AWS, GCP, and Azure to support blockchain services and distributed systems.
  • Automate infrastructure and improve system reliability using Terraform, Golang, Python, and CI/CD pipelines.
  • Operate Kubernetes clusters and middleware platforms like Kafka, Redis, and NGINX while ensuring observability and disaster recovery.

BNB Chain is a community-first and open-source blockchain ecosystem focused on mass adoption through permissionless and decentralized principles. With a collaborative and dedicated team, it aims to onboard a billion new users to Web3.

US

  • Design, scale, and maintain enterprise monitoring and alerting ecosystems across multi-cloud and native systems.
  • Bridge development and operations to ensure high availability, performance tuning, and deep visibility.
  • Automate infrastructure and build robust observability pipelines using cloud-native tools like Prometheus, Grafana, and GCP.

Ontrac Solutions is a leading technology consulting firm specializing in cutting-edge solutions that drive business transformation. Their team is committed to innovation, collaboration, and excellence, empowering clients to succeed in an evolving digital landscape.

Europe 6w PTO

  • Run and evolve the Kubernetes landscape (Amazon EKS, on-prem via Rancher) for all deployments.
  • Automate deployment and scaling, and build observability to spot problems before users notice.
  • Design abstractions that let backend and data engineers ship without opening tickets.

Yazio is a nutrition app company that helps millions of users in over 150 countries lead healthier lives through diet tracking. The platform engineering team is small and senior, with a remote-first culture that values efficiency and work-life balance.

UK

  • Design, build, and operate reliable infrastructure supporting AI-powered products.
  • Own and improve Kubernetes environments and cloud infrastructure.
  • Enhance production reliability through observability, automation, and incident response.

The company builds advanced AI-driven products and services. It values engineering excellence, autonomy, and individual contribution, with a global team of skilled engineers.