Source Job

US

  • Own ScaleOps' infrastructure end-to-end, including self-hosted product, SaaS platform, and AI infrastructure.
  • Manage cloud infrastructure across AWS, GCP, and Azure, covering networking, security, SSO, and compute.
  • Collaborate with customers and internal teams to ensure reliable feature delivery and eliminate operational toil.

Kubernetes Helm Go AWS GCP

20 jobs similar to Infra Engineer Team Lead

Jobs ranked by similarity.

US

  • Lead the strategic direction, engineering, and operational management of enterprise cloud and infrastructure platforms.
  • Own the cloud operating model, including service catalog, SLAs/SLOs, governance, and cost optimization.
  • Build and develop a team of cloud platform engineers, SREs, and security engineers, driving reliability and automation.

Banner Health is one of the largest nonprofit health care systems in the country, delivering hospital services and advanced technology. They offer a comprehensive benefits package and a culture centered on making health care easier for every employee.

India

  • Architect and scale multi-region microservices, APIs, and authentication infrastructure on AWS/GCP.
  • Lead SLOs, observability, incident management, and disaster recovery automation to maintain 99.99% availability.
  • Manage Kubernetes clusters and Terraform IaC while eliminating toil with Python/Go tooling.

JumpCloud is an AI-powered unified IT management platform that secures the modern workforce through identity, device, and access management. The company is remote-first with teams in 15+ countries and values building connections, thinking big, and continuous improvement.

  • Own the reliability, performance, and scalability of Runlayer's infrastructure across AWS and GCP.
  • Manage Kubernetes clusters, database reliability, and CI/CD pipelines for rapid deployments.
  • Lead incident response and partner with product engineers to design resilient systems for enterprise customers.

Runlayer builds a unified platform for MCPs, Skills, and AI Agents, providing enterprises with security, governance, and observability to deploy AI safely and at scale. Founded by engineers who built AI Actions for OpenAI and Zapier Agents, the team has raised $42M from Felicis and Khosla Ventures, serving companies like Gusto, Instacart, and Opendoor.

$200,000–$240,000/yr
US Canada Unlimited PTO

  • Own best practices for managing production infrastructure, including provisioning, scaling, configuration, capacity planning, and monitoring.
  • Build and maintain Kubernetes infrastructure at scale alongside Terraform-provisioned cloud resources.
  • Write custom automation and tooling in Go to reduce manual work and eliminate operational risk.

Turnkey is building the infrastructure for the autonomous economy, providing programmable guardrails that enable organizations to operate with autonomy and control. Founded by the team behind Coinbase Custody, it is a deeply technical, low-ego, high-agency team of experts in cryptography, security, and systems.

UK

  • Lead Cloud Platform and SRE teams to scale securely and efficiently.
  • Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
  • Champion SRE culture with SLOs, error budgets, and observability.

Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.

$150,000–$210,000/yr
Global Unlimited PTO

  • Own the design, development, and operation of infrastructure and build/release pipelines.
  • Deploy IaC and automation using Terraform, Ansible, Helm, and Go to support platform and customer requirements.
  • Collaborate closely with Product to drive roadmap direction and improve how users build and deliver software.

Manifest is on a mission to secure the global software and AI supply chain. Founded by alumni from the Department of Defense, CISA, and Palantir, it is a well-funded early-stage startup backed by leading investors and trusted by government and enterprise organizations.

$126,000–$174,000/yr
US

  • Define DevOps strategy and lead infrastructure architecture across multi-environment, multi-region cloud systems.
  • Architect and own scalable Kubernetes platforms, infrastructure as code, and DevSecOps implementation.
  • Drive platform reliability, performance SLAs, cost optimization, and lead complex migrations and AI/ML platform infrastructure.

Robots & Pencils is an applied AI engineering firm that designs and ships AI co-workers for enterprise operations. Founded in 2009, the company has delivery centers across Canada, the US, Eastern Europe, and Latin America, with teams averaging over 15 years of experience.

North America

  • Design, build, and operate production Kubernetes platforms, owning cluster architecture, networking, reliability, and security.
  • Troubleshoot complex infrastructure issues across Kubernetes, Linux, networking, and cloud environments, and improve observability and automation.
  • Take ownership of critical infrastructure initiatives, incident response, and mentorship while working autonomously in a remote-first environment.

Our partner is a technology company building and operating large-scale Kubernetes platforms and complex production infrastructure. It is a remote-first organization with a focus on autonomy, technical ownership, and collaboration across North America and Latin America.

$185,000–$200,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Coordinate with technical and non-technical staff across departments, including workflow automation that bridges infrastructure and business processes.
  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure in GCP.
  • Maintain incident response process and tooling, and build automation that reduces toil and enables self-healing infrastructure.

Branch empowers workers with financial freedom by helping companies accelerate payments and providing accessible, free financial services. It is a remote-first, award-winning FinTech with employees across the U.S., fostering a culture of transparency, accountability, and trust.

France

  • Evolve an Internal Developer Platform enabling development teams to deploy and operate applications securely with high self-service.
  • Play a crucial role in infrastructure architecture in a multi-cloud environment with a focus on GCP.
  • Build and maintain the platform used by over 50 internal clients, then support their concrete use by teams.

Lifen believes medical data can transform healthcare by reducing administrative burden, improving care coordination, and accelerating scientific discovery. Since 2015, the company has connected 800 hospitals and 150,000 healthcare professionals, with over 150 employees working remotely and from offices to unlock the potential of health data.

Brazil Unlimited PTO

  • Build and maintain the company's internal platform, driving operational excellence.
  • Collaborate with engineering squads to ensure applications are safe and reliable.
  • Take ownership of software infrastructure projects and provide off-hours support.

Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.

US Unlimited PTO

  • Lead the technical vision for platform hardening across AWS and GCP, including tier-0 access and secrets services.
  • Own the hardening roadmap for critical cloud infrastructure and drive operational excellence with rigorous on-call practices.
  • Mentor a global team of engineers and partner cross-functionally to build secure-by-default controls.

DoorDash is a technology and logistics company building the most reliable delivery network for consumers, merchants, and dashers. The company is growing rapidly and fosters a culture of learning, customer obsession, and inclusive leadership.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

$180,000–$205,000/yr
US Unlimited PTO

  • Lead and grow a team of Platform and DevOps engineers while staying hands-on and setting technical direction.
  • Own the AWS platform end-to-end including infrastructure as code, Kubernetes, CI/CD, observability, incident response, and cloud cost.
  • Drive the adoption of AI coding agents and automated testing to make software delivery faster and safer.

HopSkipDrive is a technology company that provides safe, fast, and simple supplemental student transportation through a marketplace connecting kids to highly-vetted caregivers. Founded by three moms, the company has facilitated over five million rides across 20+ states, is a Series D company, and has raised $100M to date.

Global

  • Design and scale cloud-native infrastructure for a fast-growing bandwidth marketplace.
  • Lead platform reliability with deep expertise in Kubernetes, GCP, and container orchestration.
  • Drive infrastructure decisions across the full stack, from virtualization to production APIs.

Share is a venture-backed company building the marketplace for bandwidth, connecting internet capacity suppliers to businesses and individuals. The company is investor-backed with a high-ownership environment and a steep learning curve.

UK Ireland Estonia Netherlands Sweden Israel Eastern Europe Portugal Unlimited PTO

  • Develop and evolve foundational software and services enabling product and development teams.\n- Architect, design, and implement Infrastructure as Code using Terraform.\n- Deploy, manage, and optimize Kubernetes clusters on GCP (GKE) and AWS (EKS).

DoiT is a global technology company that helps cloud-driven organizations leverage cloud for business growth and innovation through data, technology, and human expertise. They work with over 4,000 customers worldwide and foster a remote-first, entrepreneurial culture.

United States Canada Unlimited PTO

  • Lead the Dedicated Infrastructure team, partnering with product management to balance delivery speed and quality, guide architecture decisions, and manage incidents.
  • Maintain at least 99.9% availability for Dedicated infrastructure, strengthen security, and automate operations to support more tenants at scale.
  • Recruit, onboard, and develop team members, providing clear direction and meaningful feedback while building shared ownership of development and operations.

GitLab is an intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and accelerate digital transformation. With over 50 million registered users and trust from more than 50% of the Fortune 100, GitLab fosters a high-performance culture driven by its values and continuous knowledge exchange.

India

  • Design, deploy, and maintain the reliability, availability, and performance of critical systems and APIs across AWS and GCP.
  • Build observability frameworks, define SLIs/SLOs, and implement monitoring using Datadog and Kubernetes.
  • Participate in on-call rotations, incident response, and blameless post-incident reviews to drive systemic improvements.

JumpCloud is an AI-powered unified IT management platform that secures the modern workforce by consolidating identity, device, and access management. The company is remote-first with teams in over 15 countries and values building connections, thinking big, and continuous improvement.

$75,450–$169,700/yr
Global Unlimited PTO

  • Lead Remote's SRE team owning Kubernetes, AWS, PostgreSQL, CI, and observability.
  • Balance 60% hands-on technical work with 40% people leadership and career growth.
  • Drive a maturing reliability practice including SLOs, incident response, and on-call.

Remote is a global employment platform that helps companies recruit, pay, and manage international teams. The company is fully remote with a future-focused, async culture and employees across six continents.

US

  • Innovate and build cutting-edge cloud-native infrastructure using public cloud technologies.
  • Deliver secure, efficient, and highly available frameworks that simplify infrastructure complexity.
  • Optimize and enhance existing systems to boost performance and operational efficiency.

Airbnb is a global community marketplace that connects hosts and guests for stays and experiences. With over 5 million hosts and 2 billion guest arrivals, the company values diversity, inclusion, and belonging.