Source Job

$6,500–$9,000/mo
Brazil

  • Architect, deliver, and maintain critical cloud platform components on AWS EKS, focusing on production reliability and observability.
  • Establish SRE standards including SLOs, error budgets, and automated tooling to reduce operational friction.
  • Provide technical advisory through code reviews and architecture recommendations to maintain high platform standards.

Terraform OpenTelemetry Kubernetes CI/CD

20 jobs similar to Sr. DevOps & Infrastructure Consultant

Jobs ranked by similarity.

Argentina

  • Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
  • Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
  • Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.

Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.

North America Canada Latin America

  • Design and operate scalable cloud infrastructure across AWS and GCP.
  • Build and improve Kubernetes, Linux, and cloud networking environments.
  • Strengthen security, disaster recovery, and platform resilience.

Hubstaff provides workforce analytics and time tracking for remote teams, serving over 200,000 global users. The company is a product-led organization with a winning culture and a fully remote team of experienced engineers.

Latin America

  • Serve as the technical backbone of cloud infrastructure operations, bridging incident detection and advanced architecture.
  • Build and maintain CI/CD pipelines, design IaC modules, and optimize cloud resources for performance and cost efficiency.
  • Lead observability initiatives, integrate DevSecOps practices, and collaborate with cross-functional teams to ensure robust cloud reliability.

CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. They operate with a nearshore model and focus on empowering businesses through staff augmentation, dedicated teams, and software engineering.

UK

  • Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
  • Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
  • Drive AI-specific observability, FinOps, and security practices across the platform.

We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.

$185,000–$280,000/yr
US 4w PTO

  • Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
  • Scale single-tenant deployments and build observability, incident response, and compliance practices.
  • Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.

Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.

Global

  • Design and maintain AWS infrastructure using Terraform, with a focus on scalability cost and PCI-scoped network segmentation
  • Build and evolve the observability stack and CI/CD pipelines to ensure smooth production operations and rapid deployment
  • Lead incident response define SLOs and run performance tests to optimize payment-critical services

Xplor Technologies provides vertical software, embedded payments, and AI tools for membership-based and service-based industries. With over 130,000 businesses in 72+ countries and processing $47 billion in payments annually, the company values diversity, collaboration, and a people-first culture.

Brazil

  • Support and evolve cloud infrastructure within a technology-driven financial services environment.
  • Automate operations, improve deployment processes, and strengthen platform reliability.
  • Collaborate with engineering teams to implement automation, monitoring, and security best practices.

This company operates in the financial services sector, leveraging technology to deliver high-performance digital solutions. It is a remote-first organization with a collaborative culture focused on autonomy, continuous learning, and technical excellence.

$150,000–$250,000/yr
US Europe Singapore

  • Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
  • Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
  • Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.

A fast-growing AI/ML platform startup building infrastructure for training, evaluating, and aligning AI models within reinforcement learning environments. The engineering team of ~15 includes competitive programming medalists, serial AI startup founders, and researchers published at top venues.

US

  • Design and maintain AWS cloud infrastructure using OpenTofu and Terraform.
  • Operate Kubernetes workloads on Amazon EKS, managing GitOps deployments with Argo CD and Helm.
  • Implement observability with Datadog, troubleshoot production incidents, and support on-call rotation.

PAR Technology Corporation provides innovative restaurant technology solutions, including point-of-sale, digital ordering, loyalty, and back-office software, as well as hardware and drive-thru offerings. With over 40 years of experience, the company serves more than 100,000 restaurants globally and fosters a collaborative culture centered on its 'Better Together' ethos.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

$140,000–$175,000/yr
US

  • Lead and grow a team of platform engineers, coaching them on infrastructure and cloud challenges.
  • Drive the platform roadmap, balancing reliability, cost, security, and developer experience with AWS and Kubernetes.
  • Partner cross-functionally to align platform priorities with business goals and ensure system reliability.

PerfectServe is a leading provider of clinical communication and physician scheduling solutions in the health IT space. The company has 400+ employees and 30,000+ customers, with over $100 million in annual revenue, and has received multiple Best in KLAS awards.

LATAM

  • Own and evolve Kubernetes and cloud infrastructure on AWS for scalability, reliability, and usability.
  • Design and improve CI/CD pipelines and developer workflows to enable fast, safe, repeatable deployments.
  • Work cross-functionally with product engineers to understand needs and enable them through tooling and best practices.

Artsy is an online platform that connects collectors, artists, and gallerists to make the art world more accessible. The company values an inclusive culture and a diverse workforce, with a team that operates with open-source principles and a focus on impact.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.

Global Unlimited PTO

  • Own Primer's internal developer platform end to end, including CI/CD pipelines, deployment workflows, and self-service tooling.
  • Build the human-AI development loop, creating tooling and automation for coding agent workflows.
  • Treat developer productivity as a measurable system, using frameworks like DORA to identify and fix delivery bottlenecks.

Primer provides a unified infrastructure for global payments, enabling finance and payments teams to reduce complexity and capture revenue. Backed by top investors like Sofina and Accel, they operate as a remote-first, async culture with high autonomy and low bureaucracy.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

Global

  • Manage and optimize multi-cloud infrastructure (AWS required, GCP optional) with Kubernetes and CI/CD pipelines.
  • Improve observability through monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Coralogix).
  • Drive automation and Infrastructure as Code (IaC) using Terraform and Helm, and provide architectural guidance.

NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior. In 2023, NIQ combined with GfK, bringing together two industry leaders with operations in 100+ markets and covering more than 90% of the world's population.

$200,000–$240,000/yr
US Unlimited PTO 16w maternity 16w paternity

  • You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.

Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

$140,000–$165,000/yr
Global Unlimited PTO

  • Design, build, and optimize cloud infrastructure (AWS/Kubernetes/EKS) and CI/CD pipelines across multiple teams.
  • Troubleshoot and resolve production incidents of varying scope, ensuring reliability and performance.
  • Drive infrastructure projects end-to-end, mentor engineers, and establish standards that improve developer productivity.

Pacvue is a leading Commerce Media OS powering over $12B in advertising spend across 100+ global retail media networks. It enables over 70,000 brands and agencies with an inclusive global community that fosters innovation and career growth.

$120,000–$155,000/yr
Global

  • Own infrastructure as code across development, staging, and production environments
  • Build, maintain, and improve CI/CD pipelines for reliable and efficient deployments
  • Manage cloud infrastructure, establish scalable engineering practices, and lead incident response

CelebriOS is a software company building B2B SaaS products that help businesses make better decisions and streamline operations. The company has a remote-first working environment and a benefits package designed to support their team.