Similar Jobs
See allSite Reliability Engineer
PulsePoint
US
Kubernetes
Terraform
Prometheus
Site Reliability Engineer
Yuno
Global
AWS
Kubernetes
Terraform
Data Engineer
PulsePoint
US
Python
Spark
Kafka
Senior Data Engineer (Team Lead)
Partner Company
Germany
Apache Kafka
Apache Spark
Hadoop
Principal Staff Software Engineer, Systems Infrastructure
Distributed Systems
Site Reliability Engineering
High Availability
Your role:
- Ensure Kafka reliability – architecture, topic design, governance, partition strategy, throughput and latency optimization.
- Own Ceph reliability – operations, pool design, placement optimization, capacity planning.
- Build operational automation to reduce manual toil, speed up incident response, and prevent failures before they happen.
Requirements:
- 5+ years running distributed systems reliability at scale in production.
- Deep expertise in Kafka, Ceph, or similar distributed infrastructure.
- Proven ability to design for scale, reliability, and failure recovery.
- Experience mentoring engineers and making technical decisions.
- Willingness to work 9am–6pm ET US hours.
We offer:
- Remote work, high engineering bar and comfortable culture.
- Flat hierarchy with easy access to business, product, and operations.
- Enormous scale (20PB+ cluster) with real growth potential.
- Ownership and direct impact, you have room to shift focus as your interests evolve.
PulsePoint
PulsePoint sits at the intersection of healthcare and adtech, helping brands and agencies interpret health journey signals and unify digital determinants of health with real-world data. We are a 300+ employee post-acquisition business, one of the leading players in the US healthcare ad market.