Source Job

Argentina

  • Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
  • Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
  • Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.

Terraform Kubernetes SRE OpenTelemetry

20 jobs similar to Sr. DevOps & Infrastructure Consultant (remote)

Jobs ranked by similarity.

$6,500–$9,000/mo
Brazil

  • Architect, deliver, and maintain critical cloud platform components on AWS EKS, focusing on production reliability and observability.
  • Establish SRE standards including SLOs, error budgets, and automated tooling to reduce operational friction.
  • Provide technical advisory through code reviews and architecture recommendations to maintain high platform standards.

Inflect is a US-based advisory and marketplace revolutionizing digital infrastructure procurement. They are a small team focused on reducing friction in buying datacenter, cloud, and network services through automation and better deal terms.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

Latin America

  • Serve as the technical backbone of cloud infrastructure operations, bridging incident detection and advanced architecture.
  • Build and maintain CI/CD pipelines, design IaC modules, and optimize cloud resources for performance and cost efficiency.
  • Lead observability initiatives, integrate DevSecOps practices, and collaborate with cross-functional teams to ensure robust cloud reliability.

CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. They operate with a nearshore model and focus on empowering businesses through staff augmentation, dedicated teams, and software engineering.

North America Canada Latin America

  • Design and operate scalable cloud infrastructure across AWS and GCP.
  • Build and improve Kubernetes, Linux, and cloud networking environments.
  • Strengthen security, disaster recovery, and platform resilience.

Hubstaff provides workforce analytics and time tracking for remote teams, serving over 200,000 global users. The company is a product-led organization with a winning culture and a fully remote team of experienced engineers.

Global

  • Design and maintain AWS infrastructure using Terraform, with a focus on scalability cost and PCI-scoped network segmentation
  • Build and evolve the observability stack and CI/CD pipelines to ensure smooth production operations and rapid deployment
  • Lead incident response define SLOs and run performance tests to optimize payment-critical services

Xplor Technologies provides vertical software, embedded payments, and AI tools for membership-based and service-based industries. With over 130,000 businesses in 72+ countries and processing $47 billion in payments annually, the company values diversity, collaboration, and a people-first culture.

$200,000–$240,000/yr
US Unlimited PTO 16w maternity 16w paternity

  • You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.

Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.

UK

  • Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
  • Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
  • Drive AI-specific observability, FinOps, and security practices across the platform.

We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

$185,000–$280,000/yr
US 4w PTO

  • Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
  • Scale single-tenant deployments and build observability, incident response, and compliance practices.
  • Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.

Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.

US

  • Design and maintain AWS cloud infrastructure using OpenTofu and Terraform.
  • Operate Kubernetes workloads on Amazon EKS, managing GitOps deployments with Argo CD and Helm.
  • Implement observability with Datadog, troubleshoot production incidents, and support on-call rotation.

PAR Technology Corporation provides innovative restaurant technology solutions, including point-of-sale, digital ordering, loyalty, and back-office software, as well as hardware and drive-thru offerings. With over 40 years of experience, the company serves more than 100,000 restaurants globally and fosters a collaborative culture centered on its 'Better Together' ethos.

$150,000–$250,000/yr
US Europe Singapore

  • Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
  • Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
  • Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.

A fast-growing AI/ML platform startup building infrastructure for training, evaluating, and aligning AI models within reinforcement learning environments. The engineering team of ~15 includes competitive programming medalists, serial AI startup founders, and researchers published at top venues.

Global

  • Manage and optimize multi-cloud infrastructure (AWS required, GCP optional) with Kubernetes and CI/CD pipelines.
  • Improve observability through monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Coralogix).
  • Drive automation and Infrastructure as Code (IaC) using Terraform and Helm, and provide architectural guidance.

NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior. In 2023, NIQ combined with GfK, bringing together two industry leaders with operations in 100+ markets and covering more than 90% of the world's population.

Hungary

  • Build and maintain scalable, reliable, and secure environments on AWS using Infrastructure as Code tools.
  • Design and manage CI/CD pipelines, oversee Kubernetes clusters, and ensure GitOps practices.
  • Monitor system health with OpenTelemetry and Grafana, enforce security best practices, and mentor junior engineers.

Deutsche Telekom IT Solutions is a subsidiary of the Deutsche Telekom Group, providing IT and telecommunications services with over 5,300 employees. Recognized as Hungary's most attractive employer, it serves large corporate clients across Europe.

US

  • Design and implement monitoring and alerting systems using tools like Prometheus, Grafana, and DataDog to ensure high availability and reliability.
  • Optimize performance and reliability of healthcare payment applications, lead incident response, and develop SLOs/SLIs.
  • Automate CI/CD pipelines, infrastructure provisioning with Terraform, and manage cloud infrastructure on AWS with Kubernetes.

LMI is a digital solutions provider accelerating government impact with innovation and speed, bringing commercial-grade platforms and mission-ready AI to federal agencies. Headquartered in Tysons, Virginia, LMI serves the defense, space, healthcare, and energy sectors, focusing on agility and collaboration to drive impactful results.

Latin America

  • Design, build, and maintain reliable cloud infrastructure on AWS using CI/CD pipelines and IaC tools.
  • Automate containerized workloads with Docker, implement monitoring and observability solutions, and ensure security best practices.
  • Collaborate with engineering teams to improve system reliability, troubleshoot issues, and drive operational excellence.

GoFasti is a Talent-as-a-Service company that bridges world-class developers and designers from Latin America with first-class companies globally. They are a remote-first organization focused on matching top talent with international opportunities.

Global

  • Own and optimize CI/CD pipelines, Kubernetes deployment, and infrastructure for model serving and inference.
  • Build telemetry, observability, and alerting to catch real problems and reduce noise.
  • Eliminate toil through thoughtful automation and improve developer and agent productivity.

Obvious is building an AI-native workspace that serves as an operating system for work, putting co-intelligence at the center. They are a small, talent-dense team with founders and leaders from top tech companies.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.

$150,000–$175,000/yr
US

  • Design and implement cloud infrastructure using AWS, Azure, and Terraform.
  • Manage Kubernetes clusters for container orchestration and deployment.
  • Collaborate with cross-functional teams to ensure scalability and reliability.

BSC Analytics is a leader in advanced data analytics for highly regulated enterprises, providing technical strategy and teams of exclusively senior talent. The company fosters a culture of senior expertise, tackling the toughest data challenges.

$97,976–$119,748/yr
Europe

  • Design, build, and operate AWS infrastructure across multiple regions using Kubernetes, Terraform, and Helm.
  • Own large-scale object storage environments exceeding 15 petabytes, optimizing reliability, performance, scalability, and cost.
  • Build an internal developer platform for self-service infrastructure, strengthen security practices, and lead incident response and postmortems.

Our partner builds an AI-powered sports media platform handling petabyte-scale media storage and millions of minutes of video monthly. It operates as a remote-first, international, engineering-led team with significant autonomy and a focus on high availability and security.

Global Unlimited PTO

  • Own Primer's internal developer platform end to end, including CI/CD pipelines, deployment workflows, and self-service tooling.
  • Build the human-AI development loop, creating tooling and automation for coding agent workflows.
  • Treat developer productivity as a measurable system, using frameworks like DORA to identify and fix delivery bottlenecks.

Primer provides a unified infrastructure for global payments, enabling finance and payments teams to reduce complexity and capture revenue. Backed by top investors like Sofina and Accel, they operate as a remote-first, async culture with high autonomy and low bureaucracy.