Lead Site Reliability Engineer (Performance & Scalability)

Tech Holding

Remote regions

US

Benefits

Similar Jobs

See all

Key Responsibilities:

  • Establish performance, throughput, latency, and capacity baselines for critical customer and platform workflows.
  • Define and maintain SLOs, error budgets, dashboards, alerts, and reliability thresholds.
  • Instrument and analyze the full request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services.

Required Skills:

  • Significant experience in Site Reliability Engineering, performance engineering, or distributed systems.
  • Deep understanding of observability, performance analysis, capacity planning, and reliability engineering.
  • Strong hands-on experience with cloud infrastructure and production distributed systems.

What Success Looks Like:

  • Established measurable throughput, latency, and capacity baselines for critical platform journeys.
  • Identified the platform's most significant scalability and reliability constraints and created an actionable remediation roadmap.
  • Validated representative high-scale scenarios through load, soak, stress, failure, and recovery testing.

Tech Holding

Tech Holding is a full-service consulting firm that delivers predictable outcomes and high-quality solutions to clients. The company was founded by experienced industry professionals who have held senior positions at startups to Fortune 50 firms, fostering a culture of deep expertise, integrity, transparency, and dependability.

Apply for This Position