Similar Jobs
See allSite Reliability Engineer
XTB
Poland
Python
Kubernetes
Ansible
Staff Software Engineer, Observability
US
Distributed Systems
Kubernetes
Prometheus
Senior Software Engineer / SRE (Observability Focus)
Unknown
APAC
Python
Kubernetes
Datadog
Senior Site Reliability Engineer
Partner Company
US
Google Cloud Platform
Kubernetes
Terraform
Cloud Software Engineer - Observability Platform
ClickHouse
US
Go
Kubernetes
Terraform
Responsibilities:
- Rethink and redesign observability, from data collection through normalization, routing, processing, to querying and visualization.
- Work closely with Core Infrastructure, Platform, and Developer Tools teams.
- Ensure visibility into the state, performance, reliability, and user experiences of our software stack.
Requirements:
- 7+ years experience designing and building large-scale distributed systems.
- Strong understanding of observability principles (metrics, logs, events, traces).
- Experience with OpenTelemetry, OLAP databases, or developer-facing tools.
Nice to Have:
- Experience in high-traffic, real-time, or event-driven systems.
- Comfortable with cloud-native environments (AWS, GCP, Kubernetes).
- Strong written and verbal communication skills.
Whatnot
Whatnot is the largest live shopping platform in North America and Europe, enabling sellers to build businesses across hundreds of categories. They are a remote co-located team anchored in hubs across the US, UK, Ireland, Poland, Germany, and Australia, and were recently named the #1 Best Startup Employer in America by Forbes.