Source Job

US

  • Define and evolve enterprise observability vision, standards, and roadmap.
  • Lead implementation of Dynatrace SaaS and platform capabilities.
  • Design observability solutions for cloud-native infrastructure.

Dynatrace OpenTelemetry Kubernetes AIOps AWS

20 jobs similar to Sr. Observability Engineer

Jobs ranked by similarity.

Europe US

  • Design and architect observability solutions leveraging OpenTelemetry, Kubernetes, and cloud-native technologies.
  • Develop and execute Proofs of Concept (POCs) that highlight Dash0's differentiated technical capabilities.
  • Deliver engaging technical demos and presentations tailored to engineering and executive audiences.

Dash0 is building an OpenTelemetry-native observability platform that eliminates vendor lock-in and provides transparent pricing. Backed by top-tier investors including Balderton Capital, Accel and Cherry Ventures, the company has a collaborative, fast-moving team culture with a builder mindset.

  • Serve as a trusted technical advisor guiding customers through their observability journey.
  • Design and guide customer observability maturity strategies to improve reliability and operational visibility.
  • Provide expert troubleshooting and technical recommendations to resolve complex challenges.

Jobgether is a platform that uses AI to match candidates with jobs. They focus on remote work and have a collaborative culture built around transparency, autonomy, and trust.

Europe

  • Design and architect enterprise observability solutions using cloud-native technologies and OpenTelemetry.
  • Deliver technical demonstrations, proof-of-concepts, and architecture sessions to showcase product value.
  • Support enterprise sales cycles by partnering with account teams and advising on best practices.

The company specializes in modern observability solutions for cloud-native environments. It is a high-growth technology company with a collaborative, fast-paced culture emphasizing ownership and impact.

$182,000–$217,000/yr
United States 6w PTO

  • Serve as the primary technical point of contact for a portfolio of Grafana customers, designing and guiding their observability maturity journey.
  • Conduct regular technical reviews, health checks, and root cause analysis to drive adoption and ensure customer success.
  • Act as the voice of the customer internally, shaping product feedback and roadmap priorities while building long-term strategic relationships.

Grafana Labs builds the open source observability platform Grafana and its fully managed cloud service. The company has over 1,600 team members across 40+ countries, serving more than 7,000 customers including major enterprises, and fosters a remote-first, transparent, and innovation-driven culture.

Global 7w PTO

  • Lead the design and operation of LivePerson's observability platforms across logs, metrics, traces, alerting, and synthetic monitoring.
  • Own large-scale observability pipelines using technologies like Elastic Cloud, Grafana, Prometheus, and Kafka.
  • Provide technical leadership and mentorship while driving best practices in DevOps, cloud engineering, and observability.

LivePerson is a leader in trusted enterprise conversational AI and digital transformation, powering nearly a billion conversational interactions every month. The company is recognized as the #1 Most Innovative AI Company by Fast Company and fosters a diverse, inclusive culture that empowers employees globally.

APAC

  • Design, build, and maintain software, APIs, and automation to enhance platform reliability and observability.
  • Support monitoring, reliability, and continuous improvement in Kubernetes-based environments with a focus on Datadog.
  • Integrate observability into CI/CD pipelines and automate operational tasks using scripting languages like Python.

US

  • Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
  • Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
  • Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.

The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.

Germany

  • Serve as a trusted technical advisor for enterprise customers in observability.
  • Design architectures and deliver demonstrations using cloud-native technologies.
  • Guide complex technical decisions to maximize observability value.

Our partner specializes in observability solutions for cloud-native environments. They are a remote-first, high-growth company with a collaborative culture.

US

  • Lead observability and monitoring operations integration and workflow support, including event-to-incident patterns and dashboard visualization.
  • Manage events and incidents tied to monitoring platforms, and support OpenTelemetry implementation and automation using Splunk SOAR/Ansible.
  • Provide operational reporting views and ensure hands-on experience with enterprise monitoring, observability, or APM engineering for technical and non-technical stakeholders.

Makpar is a comprehensive professional and technical solutions provider for the Federal government, combining cloud engineering, data management, cybersecurity, and emerging technologies. They are an Equal Opportunity Employer with a connected and engaged workforce dedicated to delivering mission success for government clients.

$235,000–$275,000/yr

  • Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.

Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.

$200,000–$240,000/yr
US Unlimited PTO 16w maternity 16w paternity

  • You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.

Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.

US 16w maternity 16w paternity

  • You will define the product vision, strategy, and roadmap for Elastic Agent, Fleet Server, and telemetry collectors.
  • You will drive data-driven decisions by defining and tracking KPIs for agent deployment success and pipeline efficiency.
  • You will lead cross-functional initiatives with engineering, UX, and marketing to deliver compelling collector experiences.

Elastic, the Search AI Company, enables everyone to find answers in real time using all their data at scale. Used by more than 50% of the Fortune 500, Elastic's cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI.

Argentina

  • Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
  • Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
  • Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.

Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.

India

  • Maintain observability platform and introduce observability on new projects.
  • Implement automated management features and configure solutions per security processes.
  • Manage CI systems and pipelines, and design and implement infrastructure.

Lingaro is a global technology company providing data, cloud, and DevOps solutions. With over 1,500 employees across 7 sites, they foster a diverse and inclusive culture.

$190,000–$225,000/yr
US

  • Design, build, and operate core cloud infrastructure on AWS, including compute, networking, and container orchestration.
  • Own the CI/CD platform used across engineering teams, including build pipelines, environment promotion, and progressive rollout.
  • Build and maintain the observability stack across the organization, including logging, metrics, distributed tracing, and alerting.

RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company is an Equal Opportunity and Affirmative Action employer committed to diversity and collaboration.

$217,000–$303,900/yr
US Unlimited PTO

  • Work collaboratively with a team to create and maintain the foundational platform for Reddit's infrastructure.
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
  • Contribute upstream changes to open source projects and share on-call responsibilities.

Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information, employing a flexible-first workforce that values open-source contributions.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

US

  • Design and implement monitoring and alerting systems using tools like Prometheus, Grafana, and DataDog to ensure high availability and reliability.
  • Optimize performance and reliability of healthcare payment applications, lead incident response, and develop SLOs/SLIs.
  • Automate CI/CD pipelines, infrastructure provisioning with Terraform, and manage cloud infrastructure on AWS with Kubernetes.

LMI is a digital solutions provider accelerating government impact with innovation and speed, bringing commercial-grade platforms and mission-ready AI to federal agencies. Headquartered in Tysons, Virginia, LMI serves the defense, space, healthcare, and energy sectors, focusing on agility and collaboration to drive impactful results.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

Poland

  • Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
  • Define and drive SRE platform strategy, incident management, and observability engineering.
  • Mentor team members, foster collaboration, and ensure operational excellence.

XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.