Source Job

Europe

  • Build and operate the infrastructure behind AI-powered products, improving reliability, security, scalability, and cost efficiency.
  • Write code, automate infrastructure, investigate production issues, and design systems that reduce operational complexity.
  • Take ownership of unfamiliar systems, identify highest-leverage improvements, and balance immediate production needs with long-term platform investments.

Kubernetes GCP Cloud Networking Infrastructure As Code Distributed Systems

20 jobs similar to Senior Engineer, SRE

Jobs ranked by similarity.

UK

  • Design, build, and operate reliable infrastructure supporting AI-powered products.
  • Own and improve Kubernetes environments and cloud infrastructure.
  • Enhance production reliability through observability, automation, and incident response.

The company builds advanced AI-driven products and services. It values engineering excellence, autonomy, and individual contribution, with a global team of skilled engineers.

$125,000–$250,000/yr
Global

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads.
  • Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
  • Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.

Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.

$62,640–$104,760/yr
Europe 4w PTO

  • Design, build, and run distributed cloud architectures and large-scale production systems.
  • Ensure reliability, observability, performance, and cost efficiency of the platform.
  • Collaborate with product and backend teams to design system architecture and optimize resource use.

Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.

$152,000–$195,000/yr
US Unlimited PTO

  • Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
  • Build and operate AI tooling infrastructure, including MCP servers and secure AI access.
  • Optimize CI/CD pipelines, implement progressive delivery, and advance Infrastructure as Code.

SecurityScorecard is the global leader in cybersecurity ratings, rating over 12 million companies across 64 countries. Headquartered in New York, it is recognized as a best workplace and funded by top investors.

Europe

  • Lead investigation and resolution of complex infrastructure, networking, and platform incidents.
  • Provide technical leadership for Kubernetes platform operations and drive automation initiatives.
  • Mentor engineers and develop operational standards, runbooks, and best practices.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Serving enterprises like Adobe, PayPal, and Volkswagen, Mirantis is committed to open standards and freedom from lock-in.

US

  • Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
  • Use AI agents as force multipliers to automate manual processes and improve developer experience.
  • Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.

Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.

Poland

  • Design, write and deliver software to implement and support large web-scale, highly-performant, highly-available infrastructure on GCP/AWS.
  • Monitor infrastructure, respond to incidents, correct and improve systems to prevent incidents, and plan capacity.
  • Tune large-scale clusters for optimal performance and efficiency and support system deployments and product releases.

OpenX develops digital advertising marketplaces and technologies to optimize ad delivery for publishers and advertisers. The company operates a large-scale cloud infrastructure in Poland and values teamwork, customer centricity, and continuous learning.

$150,000–$175,000/yr
US

  • Design and manage GCP project structure, networking, and core infrastructure.
  • Own Kubernetes cluster infrastructure and define infrastructure-as-code standards.
  • Improve system architecture for scalability, resilience, and performance.

UJET provides an AI-powered contact center platform that redefines customer experience with cloud-native architecture and mobile-first approach. The company is a growing tech firm with a collaborative culture focused on innovation and security.

Europe Middle East Asia North America

  • Build and operate frameworks to ensure reliable, sustainable solution delivery across Mistral-hosted and customer-hosted environments.
  • Operate Tier-1 customer environments, ensure SLO compliance, manage on-call and incident response.
  • Productize deployment, security, and scaling of Applied AI solutions with automation and security guardrails.

Mistral provides full-stack AI solutions from frontier models to developer tools, applications, and compute, partnering with enterprises across high-stakes industries. It is a dynamic, collaborative team with a diverse workforce distributed globally, known for being creative, low-ego, and team-spirited.

UK

  • Design and deliver solutions for cloud hosted production infrastructure.
  • Shape how mission-critical enterprise software solutions are developed and deployed using optimized CI/CD pipelines.
  • Design, build and support infrastructure and security technologies within the cloud.

Ping Identity provides an intelligent cloud identity platform that enables secure and seamless digital experiences. They serve over half of the Fortune 100 companies and have a global team that values diversity and individuality.

Global

  • Design, build, and operate scalable cloud infrastructure and infrastructure-as-code for globally distributed services.
  • Develop and maintain CI/CD pipelines to support rapid and reliable delivery of backend and client components.
  • Own service reliability by implementing observability (metrics, logs, tracing) and leading incident response with actionable improvements.

NetBird develops an open-source zero-trust network security platform that is easy to use and affordable for teams of all sizes. Since its launch in 2021, it has gained trust among thousands of companies and connects hundreds of thousands of users worldwide, driven by a community-focused culture.

$150,000–$175,000/yr
US Canada

  • Build and maintain cloud infrastructure across GCP, Kubernetes, and Terraform.
  • Own CI/CD pipelines and deploy fully automated, locked-down systems.
  • Strengthen security, access control, and observability for a growing platform.

Gauntlet builds the financial systems of the future, operating across the entire stack to offer best-in-class vault products. The team serves over $1.5B in client TVL and brings together traditional finance and crypto-native expertise.

US 4w PTO 14w maternity 14w paternity

  • Own core compute infrastructure across multiple cloud providers and regions.
  • Design capabilities for greater performance and flexibility in service deployment.
  • Investigate and resolve challenging cloud and compute issues across the stack.

Render is a cloud platform for developers building AI-native, full-stack, multi-service applications. Trusted by over 6 million developers, the company has raised $257M in funding and values craft, velocity, and user experience.

$164,000–$218,000/yr
US Unlimited PTO

  • Lead design and evolution of secure cloud infrastructure and deployment systems for critical decentralized applications.
  • Drive improvements across CI/CD pipelines, deployment workflows, and engineering productivity practices.
  • Collaborate with developers, security specialists, product leaders, and infrastructure teams in a remote-first environment.

Our partner is building and scaling secure, high-performance infrastructure powering one of the most widely used decentralized technology platforms in the world. They operate as a fully remote, globally distributed team with a focus on DevOps, security, and blockchain technology.

Global

  • Lead a high-leverage remote team of four infrastructure engineers, driving the evolution toward a scalable zero-toil platform.
  • Guide the team through an AI-driven engineering approach to reduce manual work and achieve zero-touch, scalable infrastructure.
  • Prepare and execute the strategy for CI/CD and artifact distribution systems to scale during a quality surge without increasing engineering toil.

Camunda is the enterprise platform for agentic orchestration, enabling organizations to coordinate AI agents, people, and systems across complex business processes. Trusted by over 700 organizations worldwide, including 9 of top 10 US banks, Camunda is a fully remote and global company with 150+ engineers across 20+ teams, and is transforming into an AI-first organization.

India

  • Integrate and automate security controls into CI/CD pipelines for continuous vulnerability scanning and secure software delivery.
  • Design, implement, and maintain secure Infrastructure as Code across GCP environments, enforcing IAM and network security policies.
  • Drive compliance and continuous auditing, acting as technical point for PCI-DSS and SOC 2 evidence collection.

Zact is a leading fintech innovator specializing in cutting-edge payments apps. They are a remote-first company with a focus on hiring in India, fostering a security-first engineering culture.

$100,000–$122,222/yr
Canada 8w PTO

  • End-to-end ownership of internal orchestration platform built on event-driven architecture with Redpanda, including code, architecture, and roadmap.
  • Own infrastructure-as-code using Terraform Cloud, manage Kubernetes workloads with Helm, and provide self-service tooling for engineering teams.
  • Set SLOs, handle production on-call, lead incident response, author design docs, and operate AI-natively using tools like Cursor and Notion AI.

Velora unifies Aplos, Raisely, and Keela into one company with a shared mission to help nonprofit organizations thrive by offering fundraising, donor management, financial tracking, and communications tools. We are a financially solid company with a combined team dedicated to making nonprofit work easier, more impactful, and more sustainable.

US

  • Drive the definition and adoption of SLIs and SLOs across services, reducing toil through automation and incident response.
  • Design and architect Infrastructure as Code solutions for large-scale environments using Docker, Kubernetes, and cloud-native services.
  • Serve as primary SRE liaison for development teams, influencing architecture and conducting training for clients.

Noctua Technology, LLC is a company that drives digital transformation by treating operations as a software engineering challenge, focusing on cloud native systems. They are a dynamic team seeking a Senior SRE to define strategy and bridge development and operations for clients.

Europe 6w PTO

  • Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
  • Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
  • Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.

Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.