Source Job

Europe

  • Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
  • Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
  • Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.

Kubernetes Linux Networking Ansible Python

20 jobs similar to Senior Site Reliability Engineer / Kubernetes

Jobs ranked by similarity.

EU

  • Own and operate production infrastructure across Kubernetes, Linux, networking, and virtualization.
  • Lead incident response and implement observability to improve availability and performance.
  • Define SLOs and automate infrastructure with Ansible, Bash, Python, and GitOps.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective, data-driven processes. They foster a collaborative, international, and fully remote work environment, emphasizing autonomy and ownership for their small to mid-sized team.

US

  • Drive complex infrastructure migrations and build platform tooling and automation across multiple production environments.
  • Support development teams by consulting on infrastructure needs and improving observability and incident response.
  • Provide operational support and maintain platform reliability through structured debugging and on-call rotations.

PENN Entertainment is North America's leading provider of integrated entertainment, sports content, and casino gaming experiences. We operate across numerous locations in North America and foster a culture that cares about career growth and skill expansion.

Spain 5w PTO

  • Define SLIs, SLOs, and reliability targets for the platform.
  • Improve observability, alerting, and production readiness across services.
  • Automate operational work and support cloud/Kubernetes infrastructure.

Lodgify is a fast-growing scale-up in vacation rental technology, backed by $30M in funding. Headquartered in Barcelona, the 380+ person team of 60+ nationalities is passionate about transforming short-term rentals.

Global

  • Support deployment, operation, and reliability of production services on Kubernetes.
  • Monitor service health, investigate production incidents, and participate in on-call and postmortems.
  • Troubleshoot application runtime, networking, and service-to-service issues across Node.js and JVM.

Software Mind develops innovative solutions for global companies, partnering with tech giants and unicorns on transformative projects. They foster cross-functional engineering teams with a culture of openness, respect, and passion, combining employment with enjoyment.

US

  • Provide solutions to customers to make them successful using our products.
  • Troubleshoot customer environments and engage in active triaging with customers.
  • Participate in on-call rotation for weekend coverage.

Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life with its unified DataOps platform, Astro, powered by Apache Airflow. Trusted by more than 800 enterprises, the company fosters a diverse and inclusive culture as an equal opportunity employer.

Global 6w PTO

  • Own and improve production infrastructure reliability and stability.
  • Prepare, execute, and support deployments and infrastructure changes.
  • Build and maintain Infrastructure-as-Code solutions using Ansible and Terraform.

Social Discovery Group (SDG) is a group of social discovery companies that solve problems of loneliness, isolation, and disconnection by transforming virtual intimacy into the new normal. Our international team of digital nomads works remotely from all over the world and we are proud to be a two-time 'Great Place to Work' winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025).

$4,538–$5,772/mo
Poland

  • Define and drive reliability of systems at the scale of millions of clients, strengthening SRE practices. - Develop observability platforms and serve as a strategic partner to product engineering teams. - Enhance proactive resilience through early-warning systems, AI/ML, and incident management.

XTB is a global FinTech company specializing in online trading of financial instruments. As the largest FinTech in Poland and a leader in Central and Eastern Europe, we operate across multiple continents and are a certified Great Place to Work, focusing on employee development and training.

US

  • Improve system availability, scalability, and resilience across Flowcode's platforms.
  • Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
  • Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.

Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.

US

  • Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
  • Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
  • Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.

Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.

$119,380–$165,100/yr
Spain UK

  • Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
  • Define SLOs and SLIs to drive architectural decisions and error budget policies.
  • Conduct blameless post-incident reviews and implement long-term preventive measures.

Airalo is the world's first eSIM store, helping travelers access affordable mobile data in 200+ countries. They are a fully remote team of 400+ people across 60+ countries, with a culture of trust, ownership, and freedom.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

UK

  • Lead Cloud Platform and SRE teams to scale securely and efficiently.
  • Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
  • Champion SRE culture with SLOs, error budgets, and observability.

Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.

$145,000–$177,000/yr
US

  • Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
  • Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, and networking.
  • Design and implement solutions that improve availability, scalability, performance, and resilience of the platform.

Everbridge empowers enterprises and government organizations to anticipate, mitigate, respond to, and recover from critical events. The company focuses on building resilient systems and fostering a culture of ownership, continuous improvement, and operational excellence.

India

  • Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
  • Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
  • Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.

Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.

UK

  • Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
  • Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.

$114,700–$195,000/yr
North America

  • Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
  • Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
  • Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.

Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.

Europe 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and define per-tenant SLOs and reliability models.
  • Serve as a primary escalation point for incidents, lead response and post-incident reviews, and improve alert quality.

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations to ensure reliability and resolve incidents faster. We are a 100% remote company with team members across 40+ countries, backed by leading investors, and we foster a global collaborative culture and a passion for meaningful work.

$53,300–$119,850/yr
Global Unlimited PTO 16w maternity 16w paternity

  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.

Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.

US

  • Administer Linux and Windows VMs and support network observability platforms like Syslog-NG and Trapd.
  • Troubleshoot incidents, perform upgrades and patching, and automate tasks using Bash, PowerShell, or Python.
  • Collaborate with network and security engineers to analyze traffic patterns and ensure reliability, scalability, and security.

Blueprint is a technology solutions firm that helps organizations achieve meaningful outcomes across AI, cloud, data, and emerging technology. Its culture is built on ownership, high standards, and exceptional work, with teams across the United States.

Latin America

  • Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.

We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.