Source Job

US

  • Ensure Kafka reliability through topic design, governance, and latency optimization.
  • Own Ceph reliability including operations, pool design, and capacity planning.
  • Build operational automation to reduce manual toil and speed up incident response.

Kafka Hadoop Ceph Terraform Ansible

20 jobs similar to Site Reliability Engineer (Hadoop, Kafka)

Jobs ranked by similarity.

$0–$150,000/yr
US EU UK

  • Help design, build, and operate the Kubernetes platform used across PulsePoint.
  • Own reliability, observability, and incident response across platform services.
  • Build infrastructure automation and GitOps workflows to reduce operational toil.

PulsePoint sits at the intersection of healthcare and adtech, helping brands interpret health signals using real-world data. With over 300 employees, the company is a post-acquisition profitable leader in the US healthcare ad market, known for a flat hierarchy and high engineering bar.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

$0–$150,000/yr
US EU UK

  • Maintain and monitor Kafka, Hadoop, Presto, and RDBMS systems.
  • Ingest, validate, and process internal and third-party data flows.
  • Build Kafka consumers using Spark Streaming for near-real-time aggregation.

PulsePoint sits at the intersection of healthcare and adtech, helping brands interpret health journey signals. We are 300+ employees, growing, and a leading player in the US healthcare ad market.

Germany

  • Lead a team of data engineers to design and optimize scalable data solutions.
  • Architect and implement large-scale streaming and batch data processing using Kafka, Spark, and Hadoop.
  • Collaborate with cross-functional teams to define data strategies and ensure data governance and quality.

The company is a technology-focused organization that builds scalable data platforms. It offers a fully remote work environment with a collaborative culture and international teams.

  • Drive company-wide reliability strategy, standards, and best practices across LinkedIn Engineering.
  • Lead adoption of service criticality models to set reliability expectations based on business impact.
  • Partner with engineering teams to improve system design, reduce incident risk, and strengthen operational readiness.

LinkedIn is the world's largest professional network, built to create economic opportunity for every member of the global workforce. We foster a culture of trust, care, inclusion, and fun, investing in employee growth to transform the way the world works.

$100,000–$150,000/yr
US

  • Design and build scalable backend platforms and shared infrastructure.
  • Architect high-throughput, low-latency systems with event-driven architectures.
  • Establish observability standards, mentor engineers, and drive technical excellence.

The company is a technology organization focused on building scalable backend platforms. They operate in a collaborative environment with experienced engineers and cross-functional teams.

$200,000–$240,000/yr
US Unlimited PTO 16w maternity 16w paternity

  • You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.

Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.

$151,000–$242,000/yr
US

  • Ensure reliability, scalability, and operational excellence of analytics and data systems.
  • Provide technical leadership and direction to an offshore contract team.
  • Drive incident response, automation, and data governance initiatives.

Workiva provides an AI-powered platform that unifies finance, risk, and sustainability for complex organizations. It is a large enterprise with a collaborative and innovative culture centered on data integrity and trust.

$217,000–$303,900/yr
US Unlimited PTO

  • Work collaboratively with a team to create and maintain the foundational platform for Reddit's infrastructure.
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
  • Contribute upstream changes to open source projects and share on-call responsibilities.

Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information, employing a flexible-first workforce that values open-source contributions.

Germany

  • Architect and improve real-time data streaming systems processing millions of events per second with ultra-low latency.
  • Lead development of core data infrastructure enabling real-time decision-making and personalization.
  • Solve complex distributed systems challenges including fault tolerance, scalability, and performance optimization.

The hiring company is a technology firm focused on building real-time data platforms and AI-driven systems. The company culture emphasizes innovation, autonomy, and high engineering standards, though specific employee count is not provided.

$235,000–$275,000/yr

  • Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.

Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.

US

  • Lead the long-term technical strategy and architectural roadmap for Catena's high-throughput infrastructure.
  • Architect and scale distributed event-streaming systems capable of processing millions of concurrent telemetry pings.
  • Design standardized abstractions, API gateways, and microservice patterns for rapid integration development.

Catena is the universal API for fleet telematics data, providing a single integration for insurers, factors, and freight-tech platforms to access carrier vehicle data. They are an early-stage startup with $8.25M in funding, offering high ownership and real impact.

$200,500–$236,000/yr
US Canada Mexico

  • Partner with PMs to translate business opportunities into scalable data and AI solutions.
  • Deliver reliable data products that support experimentation and decision-making across the company.
  • Build and scale batch and real-time pipelines for high-quality data processing.

Change.org is the world’s largest platform for democracy, enabling millions of people to create change. With over 100 million users annually, the company is investing heavily in AI to build powerful tools and is expanding its talented team at the intersection of technology and social impact.

$200,000–$350,000/yr
US

  • Define architecture for complex, high-impact systems.
  • Lead company-critical engineering initiatives.
  • Mentor senior and staff-level engineers.

They build sophisticated infrastructure and software systems for critical business operations. They are a rapidly scaling company with a collaborative, fast-moving culture where ownership is encouraged and decisions are made quickly.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

US

  • Establish performance, throughput, latency, and capacity baselines for critical platform workflows.
  • Define and maintain SLOs, error budgets, dashboards, alerts, and reliability thresholds.
  • Lead load, stress, soak, spike, failure, and recovery testing in representative environments.

Tech Holding is a full-service consulting firm that delivers predictable outcomes and high-quality solutions to clients. The company was founded by experienced industry professionals who have held senior positions at startups to Fortune 50 firms, fostering a culture of deep expertise, integrity, transparency, and dependability.

$140,000–$170,000/yr
US

  • Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
  • Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
  • Serve as the on-account technical presence, partnering with customer infrastructure and security teams.

We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.

US Canada Unlimited PTO

  • Lead a unified platform strategy defining and executing a cohesive roadmap aligning Production Engineering with Tenant Experience.
  • Drive operational excellence and reliability, maintaining deep accountability for GitLab’s production outcomes and incident management.
  • Scale engineering leadership by managing a senior management team, coaching on performance, hiring, and team structure.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With over 50 million registered users and a high-performance culture driven by values and continuous knowledge exchange, GitLab is where careers accelerate and innovation flourishes.

Portugal

  • Work with teams to define SLIs and SLOs, and create systems for observability.
  • Analyze failure scenarios, create runbooks, and reduce work that does not add value.
  • Participate in incident management and facilitate on-call duty to ensure reliable production environments.

Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.

US

  • Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
  • Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
  • Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.

They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.