Similar Jobs
See allPrincipal Staff Software Engineer, Systems Infrastructure
Distributed Systems
Site Reliability Engineering
High Availability
Staff Site Reliability Engineer
Filevine
Kubernetes
Python
Go
Senior Site Reliability Engineer
Valtech
Portugal
Site Reliability Engineering
DevOps
Cloud Engineering
Engineering Team Leader (Site Reliability Engineering)
XTB
Poland
Python
Kubernetes
Ansible
Site Reliability Engineer
Yuno
Global
AWS
Kubernetes
Terraform
Key Responsibilities:
- Establish performance, throughput, latency, and capacity baselines for critical customer and platform workflows.
- Define and maintain SLOs, error budgets, dashboards, alerts, and reliability thresholds.
- Instrument and analyze the full request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services.
Required Skills:
- Significant experience in Site Reliability Engineering, performance engineering, or distributed systems.
- Deep understanding of observability, performance analysis, capacity planning, and reliability engineering.
- Strong hands-on experience with cloud infrastructure and production distributed systems.
What Success Looks Like:
- Established measurable throughput, latency, and capacity baselines for critical platform journeys.
- Identified the platform's most significant scalability and reliability constraints and created an actionable remediation roadmap.
- Validated representative high-scale scenarios through load, soak, stress, failure, and recovery testing.
Tech Holding
Tech Holding is a full-service consulting firm that delivers predictable outcomes and high-quality solutions to clients. The company was founded by experienced industry professionals who have held senior positions at startups to Fortune 50 firms, fostering a culture of deep expertise, integrity, transparency, and dependability.