Source Job

Global Unlimited PTO

  • Operate the Monad node fleet, including health, sync, upgrades, and incident response for validators, full nodes, and archive nodes.
  • Own infrastructure-as-code with Ansible, Terraform, and Kubernetes, and build observability with Prometheus, Grafana, and Loki.
  • Design and build AI agent tooling for automated operations, including runbooks-as-code and deterministic guardrails.

Linux Ansible Terraform Kubernetes Python

20 jobs similar to Senior DevOps / Infrastructure Engineer

Jobs ranked by similarity.

APAC Singapore Hong Kong

  • Design and maintain highly available cloud infrastructure across AWS, GCP, and Azure to support blockchain services and distributed systems.
  • Automate infrastructure and improve system reliability using Terraform, Golang, Python, and CI/CD pipelines.
  • Operate Kubernetes clusters and middleware platforms like Kafka, Redis, and NGINX while ensuring observability and disaster recovery.

BNB Chain is a community-first and open-source blockchain ecosystem focused on mass adoption through permissionless and decentralized principles. With a collaborative and dedicated team, it aims to onboard a billion new users to Web3.

$190,000–$230,000/yr
Global Unlimited PTO

  • Design and operate the infrastructure for a high-throughput messaging platform operating at 500K+ events/sec.
  • Build guardrails, runbooks, and validation gates that enable AI agents to safely execute deployments and operations.
  • Lead incident response and encode every fix as a new runbook and regression test.

Postscript is an AI messaging platform trusted by 20,000+ Shopify brands to drive revenue through SMS. The company is fully remote, backed by Greylock and Y Combinator, and has a culture of ownership and innovation.

$160,000–$180,000/yr
US

  • Own the infrastructure and platform powering the marketplace, focusing on reliability, observability, security, and automation.
  • Manage production AWS and EKS clusters, infrastructure as code with Terraform and GitOps, and CI/CD pipelines via GitHub Actions.
  • Build automation and internal tooling in Python, Bash, Go, and Node.js/TypeScript, and operate PostgreSQL, MongoDB, and Temporal.

Office Hours is an on-demand expert network that connects leading organizations with trusted experts across various knowledge domains. The company is hyper-growth, profitable, and expanding quickly, backed by top marketplace investors.

US

  • Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
  • Use AI agents as force multipliers to automate manual processes and improve developer experience.
  • Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.

Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.

Ireland

  • Design, build, and deploy production systems with focus on scalability, reliability, and security.
  • Develop and maintain automation to streamline operations and eliminate toil.
  • Proactively monitor systems and implement automated incident response to minimize downtime.

Arista Networks is an industry leader in data-driven networking for large data centers, campus, and routing. With over $8 billion in revenue and a culture valuing diversity, Arista is a Great Place to Work for Best Engineering Team and Best Company for Diversity.

$125,000–$250,000/yr
Global

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads.
  • Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
  • Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.

Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.

$185,000–$280,000/yr
US 4w PTO

  • Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
  • Scale single-tenant deployments and build observability, incident response, and compliance practices.
  • Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.

Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.

United States

  • Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
  • Drive reliability, monitoring, automation, and incident response for AI infrastructure.
  • Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.

Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.

US

  • Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
  • Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
  • Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.

ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.

US

  • Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.

$140,000–$165,000/yr
Global Unlimited PTO

  • Design, build, and optimize cloud infrastructure (AWS/Kubernetes/EKS) and CI/CD pipelines across multiple teams.
  • Troubleshoot and resolve production incidents of varying scope, ensuring reliability and performance.
  • Drive infrastructure projects end-to-end, mentor engineers, and establish standards that improve developer productivity.

Pacvue is a leading Commerce Media OS powering over $12B in advertising spend across 100+ global retail media networks. It enables over 70,000 brands and agencies with an inclusive global community that fosters innovation and career growth.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

$164,000–$218,000/yr
US Unlimited PTO

  • Lead design and evolution of secure cloud infrastructure and deployment systems for critical decentralized applications.
  • Drive improvements across CI/CD pipelines, deployment workflows, and engineering productivity practices.
  • Collaborate with developers, security specialists, product leaders, and infrastructure teams in a remote-first environment.

Our partner is building and scaling secure, high-performance infrastructure powering one of the most widely used decentralized technology platforms in the world. They operate as a fully remote, globally distributed team with a focus on DevOps, security, and blockchain technology.

Global

  • Design and maintain AWS infrastructure using Terraform, with a focus on scalability cost and PCI-scoped network segmentation
  • Build and evolve the observability stack and CI/CD pipelines to ensure smooth production operations and rapid deployment
  • Lead incident response define SLOs and run performance tests to optimize payment-critical services

Xplor Technologies provides vertical software, embedded payments, and AI tools for membership-based and service-based industries. With over 130,000 businesses in 72+ countries and processing $47 billion in payments annually, the company values diversity, collaboration, and a people-first culture.

$140,000–$170,000/yr
US

  • Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
  • Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
  • Serve as the on-account technical presence, partnering with customer infrastructure and security teams.

We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.

$180,000–$210,000/yr
US Unlimited PTO

  • Lead the design and development of automated, resilient platform technologies including Observability, DevOps, and ITSM. - Manage a team of platform engineers, driving technical roadmaps and ensuring platform reliability and security. - Build and operate OpenTelemetry observability platforms using LGTM stack on Kubernetes.

Flexential is a data center and IT services company building next-gen observability platforms for 40+ data center facilities. They value diversity and offer a collaborative culture focused on innovation.

Latin America

  • Build and operate the self-service infrastructure platform where developers and agents can validate changes in minutes.
  • Build golden paths for CI/CD, GitOps, and IaC to enable self-service provisioning and shipping.
  • Own reliability and observability, carrying on-call and turning recurring toil into automation.

Luxury Presence is building the AI growth platform for real estate. Backed by Bessemer Venture Partners, the company is a Series C firm with over 90,000 real estate professionals and has been ranked on the Inc. 5000 fastest-growing companies list three years in a row.

  • Manage and support complex on-premise Linux infrastructures, including lifecycle and configuration management.
  • Automate systems and processes using Ansible, shell scripting, and Python.
  • Work with virtualization, observability stacks, and security domains like IAM and IPAM.

Software Mind develops solutions that make an impact for companies around the globe, working with tech giants, unicorns, and transformative projects. The culture embraces openness, acts with respect, shows grit & guts, and combines employment with enjoyment.

Europe

  • Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.