Source Job

Singapore Unlimited PTO

  • Lead the design and rollout of new platforms to minimize incidents and enable customer-facing features.
  • Deploy updates and improvements for both internal and end customer use cases while collaborating with engineering and operations teams.
  • Participate in an on-call rotation evenly distributed across the team in a primary/secondary pattern.

Linux Python Go Kubernetes AWS

20 jobs similar to Infrastructure Operations Engineer (APAC)

Jobs ranked by similarity.

$160,000–$208,000/yr
US Unlimited PTO

  • Design, build, and optimize reliable infrastructure for healthcare technology.
  • Improve scalability, reliability, and performance across distributed systems.
  • Collaborate with engineers and data professionals to shape modern infrastructure practices.

This company provides innovative healthcare technology solutions. It fosters a remote-first culture with a focus on engineering excellence and collaboration.

$165,000–$216,000/yr
US Unlimited PTO

  • Develop internal tools and automate infrastructure using AWS, Kubernetes, and programming languages.
  • Research and design solutions to increase website robustness, availability, and cost efficiency.
  • Collaborate on documentation, code reviews, and rollout of new processes.

Angi powers the future of the home services industry, connecting homeowners with skilled pros. With 9 brands in 8 countries and employees worldwide, Angi has helped homeowners with over 300 million home projects.

APAC Singapore Hong Kong

  • Design and maintain highly available cloud infrastructure across AWS, GCP, and Azure to support blockchain services and distributed systems.
  • Automate infrastructure and improve system reliability using Terraform, Golang, Python, and CI/CD pipelines.
  • Operate Kubernetes clusters and middleware platforms like Kafka, Redis, and NGINX while ensuring observability and disaster recovery.

BNB Chain is a community-first and open-source blockchain ecosystem focused on mass adoption through permissionless and decentralized principles. With a collaborative and dedicated team, it aims to onboard a billion new users to Web3.

United States

  • Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
  • Drive reliability, monitoring, automation, and incident response for AI infrastructure.
  • Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.

Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.

$80,000–$90,000/yr
Global

  • Learn and contribute to production infrastructure automation using Ansible and Terraform.
  • Assist in building and maintaining cloud-native services across VKE, VLB, and VCR.
  • Develop hands-on skills with Kubernetes and container runtimes through guided project work.

Vultr provides high-performance cloud infrastructure solutions including Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage for enterprises and AI innovators worldwide. The company is the world's largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries, and values innovation and employee growth.

Poland

  • Design, write and deliver software to implement and support large web-scale, highly-performant, highly-available infrastructure on GCP/AWS.
  • Monitor infrastructure, respond to incidents, correct and improve systems to prevent incidents, and plan capacity.
  • Tune large-scale clusters for optimal performance and efficiency and support system deployments and product releases.

OpenX develops digital advertising marketplaces and technologies to optimize ad delivery for publishers and advertisers. The company operates a large-scale cloud infrastructure in Poland and values teamwork, customer centricity, and continuous learning.

Global

  • Manage Kubernetes clusters and maintain infrastructure in the cloud.
  • Administer Linux servers and implement configuration management using Puppet or Ansible.
  • Troubleshoot and ensure observability of systems with CI/CD integration.

Xsolla is a global commerce company providing tools and services to help developers solve challenges in the video game industry. They employ over 1,500 developers and cultivate a supportive, collaborative culture focused on creativity and professional growth.

Ireland

  • Diagnose and resolve complex production issues across Linux, Kubernetes, networking, storage, and GPU systems.
  • Act as a senior escalation point for critical incidents, collaborating with engineering teams on root cause analysis.
  • Develop tools and automation in Python, Bash, or Go to improve troubleshooting efficiency and observability.

The partner company provides advanced AI and cloud infrastructure solutions, supporting large-scale distributed computing and AI workloads. They operate in a fast-moving, collaborative environment with highly skilled engineering teams focused on cutting-edge technology and operational excellence.

$150,000–$200,000/yr
US Unlimited PTO

  • Design, build, and maintain scalable cloud infrastructure on AWS using Infrastructure as Code best practices.
  • Own and evolve the Atmos-based IaC framework and manage CI/CD pipelines with GitHub Actions.
  • Collaborate with engineers and scientists to support containerized and geospatial workloads.

Vibrant Planet develops a cloud-based AI platform to manage wildfire risk and modernize land management. They are a small, high-impact team backed by climate leaders, working on pressing climate challenges.

US

  • Troubleshoot and resolve complex technical issues for customers, using debugging, networking, and system administration skills.
  • Own and drive customer technical support experience, collaborating across teams and escalating when necessary.
  • Design and implement automation solutions to scale support offerings, while participating in on-call rotation for after-hours coverage.

Wiz is redefining security for the AI era by connecting code, cloud, and runtime into a single shared context, trusted by over 65% of the Fortune 100. As one of the fastest-growing startups, now powered by Google, we offer a culture that values world-class talent and creative freedom, with a global team scanning over 230 billion files daily.

Asia

  • Design and deploy security solutions to monitor and harden the entire infrastructure in AWS.
  • Implement and update security measures for the protection of Binance.com infrastructure.
  • Utilize log ingestion platforms for security analytics and identification of attacker tactics and patterns.

Binance is a leading global blockchain ecosystem behind the world's largest cryptocurrency exchange by trading volume and registered users. Trusted by 300+ million people in 100+ countries, we offer a results-driven workplace with a flat structure and opportunities for career growth.

$20,258–$25,660/yr
India

  • Manage and maintain cloud infrastructure environments across development, staging, and production to ensure high availability and operational stability.
  • Administer Kubernetes clusters and containerized workloads, including deployment, scaling, and troubleshooting.
  • Develop and maintain automation tools and infrastructure-as-code frameworks to improve efficiency and scalability.

The company is a partner organization focused on building scalable, cloud-native infrastructure. The team is growing and operates in a remote-first, collaborative environment with a focus on continuous learning.

Global Unlimited PTO

  • Lead a high-impact infrastructure team, evolving internal platforms and CI/CD systems to support large-scale engineering operations.
  • Drive automation initiatives and AI-driven practices to reduce operational complexity and improve developer experience.
  • Define and execute strategies for scalable infrastructure, cloud environments, and platform engineering.

The partner company is a technology organization focused on building infrastructure platforms that enable engineering teams to deliver software faster. It is a remote-first company with a collaborative culture and a focus on innovation and scalability.

$90,000–$100,000/yr
Global

  • Build and maintain production-grade automation using Ansible, Terraform, and Go.
  • Engage deeply with Kubernetes internals, including scheduler, kubelet, controllers, and CRDs.
  • Harden platform infrastructure through security best practices like vulnerability scanning, container image signing, and admission controllers.

Vultr makes high-performance cloud infrastructure easy to use, affordable, and locally accessible for enterprises and AI innovators. It is the world's largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries.

US

  • Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.

$160,000–$208,000/yr
US

  • Build systems for declarative application and infrastructure lifecycle management, including CI/CD, Kubernetes, and service inventory.
  • Prioritize and troubleshoot infrastructure issues to minimize downtime and respond to alerts efficiently.
  • Contribute to setting the SRE team's direction and streamline automation of infrastructure processes.

Counterpart Health develops Counterpart Assistant, an AI-enabled primary care tool that supports physicians in chronic disease management. It is a subsidiary of Clover Health, with a remote-first culture and a focus on value-based care through technology.

US

  • Lead deployment and operation of product infrastructure in federal environments within AWS.
  • Build and maintain scalable, secure cloud-native platforms using Kubernetes, Terraform, and GitLab CI.
  • Improve development and deployment processes, create tooling for telemetry, and foster documentation culture.

Horizon3.ai is a fast-growing, remote cybersecurity company that helps organizations proactively find and fix exploitable attack vectors. We are a team of former special ops cyber operators and engineers committed to a culture of respect, collaboration, ownership, and results.

$122,000–$129,200/yr
US Unlimited PTO 16w maternity 4w paternity

  • Build, operate, and scale AWS infrastructure supporting a microservices platform.
  • Develop and maintain CI/CD pipelines to make shipping fast, reliable, and safe.
  • Partner with product engineering teams, translating platform concepts into practical guidance.

Storyblocks is a content company that provides video, audio, images, and creative tools through a subscription model, empowering storytellers. They are a remote-first company with a culture driven by data, empowered communication, and ownership, recognized as a top workplace.

Global

  • Deploy new servers and implement infrastructure changes.
  • Plan and execute infrastructure maintenance, diagnose and resolve issues.
  • Enhance infrastructure lifecycle pipelines.

Gcore is a global provider of infrastructure and software solutions for AI, cloud, network, and security, powering digital experiences worldwide. With a team of over 550 professionals, they collaborate with leading partners like Intel, NVIDIA, Dell, and Equinix to build the foundation for an AI-driven world.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.