Remote Devops Jobs · Observability

Job listings

  • Operate and evolve AWS infrastructure for Data/AI platforms, ensuring security, scalability, and high availability.
  • Build CI/CD pipelines, automate provisioning with Terraform, and implement observability.
  • Collaborate with Data, AI, and Infrastructure teams, document standards, and drive platform improvements.

The partner company is building a modern Data Platform team focused on secure, scalable, and highly available infrastructure for Data and AI workloads. The culture emphasizes DevOps, automation, and continuous improvement, with close collaboration across Data, AI, and Infrastructure teams.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

  • Design and operate CI/CD pipelines for repository intake, build execution, image creation, artefact promotion, environment deployment, and evidence generation.
  • Implement release controls for private registries, package sources, signing, SBOM generation, vulnerability scanning, dependency provenance, and approval workflows.
  • Create observability across build systems, workflow services, repositories, test environments, model endpoints, and platform components using logs, metrics, traces, and dashboards.

Deutsche Telekom IT Solutions is a subsidiary of the Deutsche Telekom Group providing a wide portfolio of IT and telecommunications services. With more than 5300 employees, it was recognized as Hungary's most attractive employer in 2025 and acknowledged as the Most Ethical Multinational Company in 2019.

$140,000–$170,000/yr

  • Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
  • Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
  • Serve as the on-account technical presence, partnering with customer infrastructure and security teams.

We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.

$109,800–$252,500/yr
US Unlimited PTO 16w maternity 8w paternity

  • Own the full platform stack for Veeam Data Cloud in Government and Sovereign Cloud environments, including incident response, reliability, and observability.
  • Design and implement high-availability, fault-tolerant infrastructure on Azure (including Azure Government) with SLIs, SLOs, and error budgets.
  • Drive reliability improvements through automation, chaos engineering, and cross-team collaboration, with a focus on compliance and security.

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide.

  • Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
  • Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
  • Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.

$235,000–$275,000/yr

  • Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.

Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.

$118,800–$237,600/yr

  • Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
  • Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
  • Manage distributed systems, observability, incident response, and automation with a security-first mindset.

Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.