Source Job

Spain

  • Report to the Manager of Customer Reliability and provide senior level technical support on customer issues.
  • Work with development engineering to track escalations, bugs, and feature requests.
  • Develop technical troubleshooting sessions, trainings, and maintain internal knowledge base articles.

Kubernetes Linux Python AWS Elasticsearch

20 jobs similar to Senior Customer Reliability Engineer

Jobs ranked by similarity.

$130,000–$163,000/yr
US

  • Own executive-level relationships and strategic engagement for a portfolio of enterprise and high-value customers.
  • Collaborate with customer teams on Sysdig deployments, architecture, and operational best practices.
  • Lead Customer Business Reviews and strategic touchpoints with both technical and non-technical stakeholders.

Sysdig is a cloud security company that created Falco, the open standard for cloud threat detection, and leads the cloud security market with runtime insights and open innovation. Trusted by over 60% of the Fortune 500, Sysdig is recognized as a Best Place to Work and one of Deloitte's fastest-growing companies.

Ireland

  • Diagnose and resolve complex production issues across Linux, Kubernetes, networking, storage, and GPU systems.
  • Act as a senior escalation point for critical incidents, collaborating with engineering teams on root cause analysis.
  • Develop tools and automation in Python, Bash, or Go to improve troubleshooting efficiency and observability.

The partner company provides advanced AI and cloud infrastructure solutions, supporting large-scale distributed computing and AI workloads. They operate in a fast-moving, collaborative environment with highly skilled engineering teams focused on cutting-edge technology and operational excellence.

UK

  • Lead cross-team incident triage for high-impact customer outages, coordinating Engineering, Product, and Customer Experience response and contributing to root cause analysis.
  • Develop and maintain observability for cloud-hosted customer deployments by building and refining system monitors, dashboards, and alerting.
  • Serve as the senior escalation point for complex support cases in EMEA, working cases that involve deep platform internals and unusual failure modes.

Dragos defends industrial organizations that provide modern civilization necessities like water, electricity, and safe working environments. As a market leader in ICS/OT Cybersecurity, we operate globally with a remote-first culture and are looking for mission-oriented teammates who value authenticity, transparency, and trust.

Ireland

  • Investigate and resolve customer technical issues across cloud security posture management, vulnerability scanning, and threat detection in AWS, Azure, and GCP environments.
  • Troubleshoot cloud connector and integration failures, including IAM, network connectivity, and API authentication issues.
  • Design and implement automation leveraging AI agents to improve triage accuracy and resolution efficiency.

Wiz provides an AI-powered cloud security platform that connects code, cloud, and runtime to secure cloud and AI applications, trusted by over 65% of the Fortune 100. As one of the fastest-growing startups, powered by Google, Wiz has a culture that values world-class talent and encourages creative thinking.

US

  • Troubleshoot and resolve complex technical issues for customers, using debugging, networking, and system administration skills.
  • Own and drive customer technical support experience, collaborating across teams and escalating when necessary.
  • Design and implement automation solutions to scale support offerings, while participating in on-call rotation for after-hours coverage.

Wiz is redefining security for the AI era by connecting code, cloud, and runtime into a single shared context, trusted by over 65% of the Fortune 100. As one of the fastest-growing startups, now powered by Google, we offer a culture that values world-class talent and creative freedom, with a global team scanning over 230 billion files daily.

United States Canada Unlimited PTO

  • Support self-managed and SaaS customers by resolving issues via Zendesk, merge requests, email, and video calls.
  • Collaborate with cross-functional teams to define product goals, fix bugs, and improve documentation.
  • Participate in on-call rotations and contribute to code, documentation, and support processes.

GitLab is an intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security and compliance risk. With over 50 million registered users and trust from more than 50% of the Fortune 100, GitLab fosters a high-performance, remote-first culture where every voice is valued.

US

  • Deliver outstanding support to open-source users and enterprise customers via multiple channels.
  • Develop deep expertise in GitOps, Argo, and Kubernetes to troubleshoot and resolve issues.
  • Contribute to support playbooks, processes, and documentation to improve customer experience.

Akuity provides enterprise support for Argo, an open-source project for automating application delivery on Kubernetes. Backed by $25 million funding, the company is experiencing rapid growth and maintains a culture of humility, authenticity, and diversity.

US 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
  • Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.

Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.

US Unlimited PTO

  • Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
  • Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
  • Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.

TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.

EMEA Unlimited PTO

  • Provide technical support to GitLab's largest Self-managed, GitLab Dedicated, and GitLab.com customers, ensuring stable and performant environments.
  • Troubleshoot complex issues using Linux, GitLab, CI/CD knowledge, and tools like Zendesk and strace, through email and video conferencing.
  • Collaborate with Product, Development, Infrastructure, Customer Success, and Sales teams to drive defect resolution and influence the roadmap.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100 trust GitLab to ship better software, and the company fosters a high-performance culture driven by values and continuous knowledge exchange.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.

Europe 6w PTO

  • Run and evolve the Kubernetes landscape (Amazon EKS, on-prem via Rancher) for all deployments.
  • Automate deployment and scaling, and build observability to spot problems before users notice.
  • Design abstractions that let backend and data engineers ship without opening tickets.

Yazio is a nutrition app company that helps millions of users in over 150 countries lead healthier lives through diet tracking. The platform engineering team is small and senior, with a remote-first culture that values efficiency and work-life balance.

Canada Unlimited PTO 20w maternity 16w paternity

  • Own high-severity technical escalations from intake through resolution or engineering handoff.
  • Partner with Support and Product/Engineering to close the gap between customer problems and engineering fixes.
  • Monitor patterns across escalations to catch systemic issues and translate into product improvements.

Tailscale builds software that makes it easy to securely interconnect people and devices. Founded in 2019, the company is fully distributed and backed by Accel, CRV, and others.

$127,008–$152,410/yr
UK Sweden Spain Germany Ireland 6w PTO

  • Partner with product engineering squads to own production reliability for high-SLA customer environments, designing automation and defining per-tenant SLOs.
  • Serve as a primary escalation point for incidents, leading response, post-incident reviews, and reducing SLO burn to prevent repeats.
  • Influence feature design for scalability and operability, improve alert quality, and eliminate toil through automation.

Grafana Labs is the company behind the open observability cloud, providing a fully managed observability platform for organizations to see, understand, and act on their data. With over 35 million users, 7,000+ customers, and 1,600+ team members across 40+ countries, we foster a remote, collaborative culture rooted in open-source values.

$62,640–$104,760/yr
Europe 4w PTO

  • Design, build, and run distributed cloud architectures and large-scale production systems.
  • Ensure reliability, observability, performance, and cost efficiency of the platform.
  • Collaborate with product and backend teams to design system architecture and optimize resource use.

Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.

India

  • Design, implement, and manage scalable, highly available systems on Azure Cloud, optimizing Kubernetes and containerized workloads.
  • Build and maintain robust CI/CD pipelines with GitHub Actions and implement infrastructure as code using Helm Charts.
  • Monitor system performance, troubleshoot issues, ensure uptime, and perform root cause analysis to improve reliability.

Resilinc is pioneering intelligent, autonomous systems that redefine supply chain risk management using agentic AI, trusted by top companies in life sciences, aerospace, high tech, and automotive. We are a fully remote, mission-led team with a collaborative culture focused on high-impact work.

$195,000–$195,000/yr
Unlimited PTO 18w maternity 12w paternity

  • Lead technical enablement programs for partners to implement Chainguard Containers and Libraries.
  • Develop partner-facing technical content like integration guides and reference architectures.
  • Act as technical liaison between partner ecosystem and Product and Engineering teams.

Chainguard is the trusted source for open source, delivering hardened, secure builds of open source software. The company is venture-backed by leading investors and serves Fortune 500 enterprises and global industry leaders.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.

Global Unlimited PTO

  • Own the architecture and evolution of the SMBP Operator's CRD API surface, designing extensions that meet enterprise expectations.
  • Design and implement flexible network configuration patterns that give customers choice of ingress controller, load balancer, and traffic management approach.
  • Build the observability integration enterprise customers expect: Prometheus-native metrics, ServiceMonitor CRDs, Kubernetes Events, and pre-built dashboards.

Ditto is redefining how data moves at the edge, making it seamless for developers to build resilient, real-time applications regardless of network conditions. With more than $145 million in funding and trusted by organizations like Chick-fil-A, Delta Airlines, and the U.S. military, Ditto is a globally distributed, fast-growing startup committed to building a diverse and inclusive team.