Engineering Manager, Production Engineering - Observability

GitLab

Remote regions

Global

Benefits

Unlimited PTO

Similar Jobs

See all

Overview:

  • Lead the globally distributed Observability team that builds and operates the metrics, logging, alerting, and capacity planning platforms GitLab engineers rely on.
  • Work with Site Reliability Engineering, Product Engineering, and other Infrastructure Platforms teams to make it easier to observe services.

What You'll Do:

  • Own the reliability, scalability, and cost of observability platforms and guide improvements to Prometheus pipelines, log ingestion, alerting, and capacity forecasting.
  • Participate in the Incident Manager On Call rotation, coordinating response to high-severity incidents and keeping on-call sustainable.
  • Use AI tools and agents to support engineering workflows and incident triage while engineers retain responsibility for decisions.

What You'll Bring:

  • Experience leading an observability, platform engineering, or site reliability engineering team operating at scale in a distributed, asynchronous environment.
  • Technical knowledge of metrics systems like Prometheus, logging platforms, alerting design, and long-term storage.
  • Experience with SLOs, error budgets, capacity forecasts, and improving production on-call rotations.

GitLab

GitLab is an intelligent orchestration platform for DevSecOps that helps organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. With more than 50 million registered users and a high-performance culture driven by shared values and AI adoption, GitLab's globally distributed team collaborates to solve complex problems.

Apply for This Position