Source Job

US

  • Design and implement monitoring and alerting systems using tools like Prometheus, Grafana, and DataDog to ensure high availability and reliability.
  • Optimize performance and reliability of healthcare payment applications, lead incident response, and develop SLOs/SLIs.
  • Automate CI/CD pipelines, infrastructure provisioning with Terraform, and manage cloud infrastructure on AWS with Kubernetes.

DevOps Observability Kubernetes CI/CD

20 jobs similar to Health DevOps Engineer - Observability/Reliability

Jobs ranked by similarity.

Global

  • Manage and optimize multi-cloud infrastructure (AWS required, GCP optional) with Kubernetes and CI/CD pipelines.
  • Improve observability through monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Coralogix).
  • Drive automation and Infrastructure as Code (IaC) using Terraform and Helm, and provide architectural guidance.

NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior. In 2023, NIQ combined with GfK, bringing together two industry leaders with operations in 100+ markets and covering more than 90% of the world's population.

Canada Europe

  • Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
  • Partnering with development teams to establish production readiness and operational readiness.
  • Building tooling to automate observability and operational workflows, eliminating manual toil.

Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.

Global

  • Ensure availability, performance, scalability, and resilience of production services in AWS.
  • Automate infrastructure provisioning and management using Infrastructure as Code (IaC).
  • Collaborate with development, architecture, security, and product teams to promote reliability best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide, operating in markets such as financial services, healthcare, automotive, and insurance. The company has over 25,200 employees across 32 countries and is recognized as a Top 25 global workplace by Fortune.

Global

  • Design and maintain AWS infrastructure using Terraform, with a focus on scalability cost and PCI-scoped network segmentation
  • Build and evolve the observability stack and CI/CD pipelines to ensure smooth production operations and rapid deployment
  • Lead incident response define SLOs and run performance tests to optimize payment-critical services

Xplor Technologies provides vertical software, embedded payments, and AI tools for membership-based and service-based industries. With over 130,000 businesses in 72+ countries and processing $47 billion in payments annually, the company values diversity, collaboration, and a people-first culture.

$109,000–$194,000/yr
US 4w PTO 16w maternity 8w paternity

  • Build, deploy, and maintain scalable, highly available systems on AWS and own CI/CD pipelines and infrastructure as code.
  • Improve system reliability with alerting, runbooks, and observability using Grafana, Loki, Mimir, and Tempo.
  • Mentor junior engineers, collaborate on infrastructure design, and participate in on-call rotation with rare off-hours incidents.

Waymark is a mission-driven team of healthcare providers, technologists, and builders working to transform care for people with Medicaid benefits. We partner with communities to deliver technology-enabled, human-centered support that helps patients stay healthy and thrive.

$120,000–$120,000/yr
Canada Unlimited PTO

  • Design and maintain scalable, reliable infrastructure using AWS services like ECS and RDS
  • Build and manage infrastructure as code with Terraform, and optimize CI/CD pipelines for secure software delivery
  • Improve system observability using Datadog, Sentry, and Coralogix, and enable a seamless developer experience

Fullscript provides a platform for healthcare practitioners to prescribe and manage supplements and lab tests for patients. They serve over 125,000 practitioners and 10 million patients across North America, with a mission to help people get better.

$154,700–$208,000/yr
US Canada

  • Design and build automated reliability and self-healing systems to protect production at scale.
  • Own and improve incident management tooling and on-call health, reducing alert noise and empowering teams.
  • Develop and evolve observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection.

Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for improved safety, efficiency, and sustainability. As a public company with over 2.3 million IoT devices deployed, we foster a culture of growth mindset and collaboration.

$180,000–$195,000/yr
US

  • Architect, build, and evolve secure, scalable cloud infrastructure on AWS and AWS GovCloud to power the Hypori SaaS platform.
  • Independently own ambiguous, high-impact infrastructure problems, guide technical direction, and act as a senior escalation point during production incidents.
  • Drive strategy and execution of Infrastructure as Code, observability, CI/CD, and operational frameworks while mentoring engineers and raising the technical bar across the organization.

Hypori Inc. is a high-growth cybersecurity SaaS company transforming secure mobility through a virtual workspace platform that enables users to access enterprise apps and data from any mobile device with zero data on the endpoint and total personal privacy. Backed by $55M in funding from investors including UBS, AE Industrial Partners, Hale Capital Partners, and GreatPoint Ventures, the company is expanding into new commercial and regulated markets.

$185,000–$280,000/yr
US 4w PTO

  • Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
  • Scale single-tenant deployments and build observability, incident response, and compliance practices.
  • Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.

Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.

Europe 6w PTO

  • Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
  • Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
  • Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.

Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.

Global

  • Own the reliability of Supabase's deployment and release systems against clear SLOs and error budgets.
  • Standardize and instrument pre-production deployment workflows for trustworthy signal.
  • Drive disaster-recovery readiness and reduce mean-time-to-detect and recover for deploy-related incidents.

Supabase is the Postgres development platform providing a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. They are a globally distributed team of ~400 members across 60+ countries, with over $1B raised and 540,000+ community members.

$200,000–$215,000/yr
US

  • Design and deploy a modern observability stack including logging, metrics, and distributed tracing across hundreds of services.
  • Build alerting policies and incident response workflows to reduce manual escalations and improve mean time to detect and resolve.
  • Automate toil and set SRE standards while mentoring engineers on observability tooling.

WeightWatchers is a global digital health company and the world's #1 doctor-recommended behavioral weight health program. They have a cloud architecture supporting hundreds of services and are expanding their platform engineering team.

SRE Engineer

IPSY
Mexico Colombia

  • Build and maintain observability across the platform in Datadog, including dashboards, monitors, APM, and log pipelines.
  • Participate in on-call rotation and incident response, driving blameless post-incident reviews and automating toil.
  • Leverage AI tools to accelerate debugging, generate runbooks, and build automation for operational efficiency.

IPSY is a beauty subscription platform that connects brands and consumers through curated beauty products. It is a remote-first company with a focus on community and engagement.

$120,000–$165,000/yr
US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

Brazil

  • Ensure system architecture meets technical requirements by collaborating with IT teams (Architecture, Security, Infrastructure).
  • Maintain and evolve the microservices environment on AWS with a focus on information security.
  • Implement DevOps practices, automation, and monitoring tools to ensure system reliability and scalability.

Experian is a global data and technology company that powers opportunities for people and businesses worldwide. With 25,200 employees across 32 countries, it fosters a people-centric, inclusive culture recognized as a World's Best Workplace.

$164,000–$218,000/yr
US Unlimited PTO

  • Lead design and evolution of secure cloud infrastructure and deployment systems for critical decentralized applications.
  • Drive improvements across CI/CD pipelines, deployment workflows, and engineering productivity practices.
  • Collaborate with developers, security specialists, product leaders, and infrastructure teams in a remote-first environment.

Our partner is building and scaling secure, high-performance infrastructure powering one of the most widely used decentralized technology platforms in the world. They operate as a fully remote, globally distributed team with a focus on DevOps, security, and blockchain technology.

US Unlimited PTO

  • Serve as the first responder for production incidents, triaging and resolving issues.
  • Monitor application health and system availability using Datadog.
  • Develop automation scripts using Python or PowerShell to improve operational efficiency.

NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. Recognized as one of the fastest-growing companies in America, it offers a fulfilling work environment with career advancement opportunities across multiple locations in the US, South America, and India.

US

  • Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
  • Use AI agents as force multipliers to automate manual processes and improve developer experience.
  • Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.

Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.