Source Job

US UK Ireland Poland Germany Australia

  • Design and implement a scalable observability platform for Whatnot's growing infrastructure.
  • Work with core infrastructure, platform, and developer tools teams to redesign data collection to visualization.
  • Utilize AI agents and open standards to ensure visibility into software stack performance and reliability.

Python Elixir Go OpenTelemetry Kubernetes

20 jobs similar to Software Engineer - Infrastructure Reliability Engineering

Jobs ranked by similarity.

$4,538–$5,772/mo
Poland

  • Define and drive reliability of systems at the scale of millions of clients, strengthening SRE practices. - Develop observability platforms and serve as a strategic partner to product engineering teams. - Enhance proactive resilience through early-warning systems, AI/ML, and incident management.

XTB is a global FinTech company specializing in online trading of financial instruments. As the largest FinTech in Poland and a leader in Central and Eastern Europe, we operate across multiple continents and are a certified Great Place to Work, focusing on employee development and training.

$217,000–$303,900/yr
US Unlimited PTO

  • Work collaboratively with a team to create and maintain the foundational platform for Reddit's infrastructure.
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
  • Contribute upstream changes to open source projects and share on-call responsibilities.

Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information, employing a flexible-first workforce that values open-source contributions.

APAC

  • Design, build, and maintain software, APIs, and automation to enhance platform reliability and observability.
  • Support monitoring, reliability, and continuous improvement in Kubernetes-based environments with a focus on Datadog.
  • Integrate observability into CI/CD pipelines and automate operational tasks using scripting languages like Python.

US

  • Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
  • Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
  • Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.

They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.

US Unlimited PTO

  • Design, build, and operate distributed systems that ingest, process, and store telemetry at very high scale.
  • Own the reliability, performance, capacity, and cost-efficiency of telemetry pipelines and storage systems.
  • Participate in the on-call rotation, help resolve production incidents, and drive root-cause fixes through to completion.

ClickHouse provides a real-time analytics database for data warehousing, observability, and AI workloads. Recognized on the 2025 Forbes Cloud 100 list, it has over 4,000 customers and significant year-over-year growth.

US

  • Define and evolve enterprise observability vision, standards, and roadmap.
  • Lead implementation of Dynatrace SaaS and platform capabilities.
  • Design observability solutions for cloud-native infrastructure.

FreedomPay provides commerce solutions. They are a large company with a culture focused on innovation and collaboration.

US

  • Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
  • Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
  • Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.

The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.

$235,000–$275,000/yr

  • Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.

Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.

$118,800–$237,600/yr
Europe

  • Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
  • Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
  • Manage distributed systems, observability, incident response, and automation with a security-first mindset.

Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.

India

  • Ensure production system reliability, scalability, and performance through automation and AI-driven operations.
  • Design and implement AIOps workflows for event ingestion, anomaly detection, and incident response.
  • Collaborate with engineering teams to embed reliability into the software lifecycle and improve system resilience.

Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. The company is a partner-focused recruitment service with a remote-first, globally distributed team.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

$74,000–$111,000/yr
Canada Unlimited PTO

  • Define and implement observability strategies, standards, and governance across applications and platforms.
  • Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms.
  • Establish and drive SRE best practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.

Valtech is the experience innovation company that helps brands unlock new value in an increasingly digital world by blending crafts, categories, and cultures. They have a workplace culture that fosters creativity, diversity, and autonomy, with a borderless global framework enabling seamless collaboration.

US

  • Improve system availability, scalability, and resilience across Flowcode's platforms.
  • Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
  • Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.

Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.

$87,480–$110,160/yr
Europe 6w PTO

  • Design, build, and operate reconciliation systems for Grafana Cloud stacks at scale.
  • Collaborate across teams to improve reliability, deployment complexity, and incident response.
  • Contribute to roadmap planning, technical design, and long-term simplification of stack operations.

Grafana Labs is the company behind the open source observability platform Grafana, providing a fully managed observability cloud. With over 1,600 team members across 40+ countries, the company fosters a global, collaborative culture rooted in open source principles.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

$200,000–$240,000/yr
US Unlimited PTO 16w maternity 16w paternity

  • You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.

Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.

$53,300–$119,850/yr
Global Unlimited PTO 16w maternity 16w paternity

  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.

Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.

Europe US

  • Design and architect observability solutions leveraging OpenTelemetry, Kubernetes, and cloud-native technologies.
  • Develop and execute Proofs of Concept (POCs) that highlight Dash0's differentiated technical capabilities.
  • Deliver engaging technical demos and presentations tailored to engineering and executive audiences.

Dash0 is building an OpenTelemetry-native observability platform that eliminates vendor lock-in and provides transparent pricing. Backed by top-tier investors including Balderton Capital, Accel and Cherry Ventures, the company has a collaborative, fast-moving team culture with a builder mindset.

Poland

  • Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
  • Define and drive SRE platform strategy, incident management, and observability engineering.
  • Mentor team members, foster collaboration, and ensure operational excellence.

XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.

$175,000–$195,000/yr

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
  • Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
  • Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.

Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.