Source Job

$180,000–$250,000/yr
US Unlimited PTO

  • Own the technical direction and architecture of critical infrastructure domains, establishing scalable patterns and standards.
  • Lead complex, multi-team infrastructure initiatives from design through implementation and production operation.
  • Design and evolve AWS and Kubernetes infrastructure to enable teams to build and deploy systems reliably at scale.

Python AWS Kubernetes Terraform PostgreSQL

20 jobs similar to Staff Infrastructure Engineer

Jobs ranked by similarity.

$185,000–$280,000/yr
US 4w PTO

  • Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
  • Scale single-tenant deployments and build observability, incident response, and compliance practices.
  • Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.

Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.

$133,000–$267,000/yr
US 12w maternity 12w paternity

  • Build, test, and deploy software in a continuous manner, contributing to technical specs and rollout plans.
  • Work with internal customers and stakeholders to ensure we’re solving the right problems.
  • Mentor and sponsor teammates to help them grow.

Pilot provides small businesses with dedicated finance experts and software to manage bookkeeping and financial management. The company has over 3,000 customers and has raised over $170 million, fostering a culture of innovation and growth.

United States

  • Design and build scalable, secure cloud infrastructure and deployment pipelines.
  • Own infrastructure end-to-end, including architecture, provisioning, deployment, and operation.
  • Lead technical design discussions and contribute to infrastructure and platform architecture decisions.

VulnCheck is the Exploit Intelligence Company, delivering structured exploit intelligence for cybersecurity. Founded in 2021, the company has a transparent, collaborative, and supportive culture with a team of experts.

Global

  • Manage and optimize multi-cloud infrastructure (AWS required, GCP optional) with Kubernetes and CI/CD pipelines.
  • Improve observability through monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Coralogix).
  • Drive automation and Infrastructure as Code (IaC) using Terraform and Helm, and provide architectural guidance.

NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior. In 2023, NIQ combined with GfK, bringing together two industry leaders with operations in 100+ markets and covering more than 90% of the world's population.

US

  • Design and implement scalable cloud infrastructure to support growth.
  • Develop monitoring, alerting, and incident response for system reliability.
  • Automate deployment pipelines and ensure high availability and security.

Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.

$186,700–$255,000/yr
US Canada 18w maternity 12w paternity

  • Create and test reliable cloud infrastructure services supporting Webflow's product range.
  • Lead initiatives to reduce triage load, increase reliability, and handle growing customer scale.
  • Collaborate with product engineering teams to deliver new solutions and improve existing services.

Webflow is an agentic web marketing platform that helps modern marketing teams build, manage, and optimize high-performing web experiences. The company values grit, speed, and craft, fostering a culture of ownership and continuous improvement.

Global

  • Own and evolve Quansight's cloud infrastructure across AWS, Azure, and GCP.
  • Lead infrastructure engagements for clients from scoping through delivery.
  • Contribute to open-source projects and participate in upstream communities.

Quansight is rooted in the Python data science community and helps companies build sustainable solutions on open-source software. The team is a small, collaborative, fully distributed group of open-source maintainers and engineers.

$126,290–$190,000/yr
United States 18w maternity 12w paternity

  • Empower engineers on other teams by maintaining monitoring tooling and collaborating on observability best practices.
  • Enhance reliability of Kubernetes applications through resource optimization, streamlined upgrades, and scalability.
  • Participate in on-call and incident response processes, occasionally diving into application code to debug production issues.

Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. It serves over 2 million users worldwide across 190 countries, with tens of thousands of projects launched each month, and fosters a culture of grit, speed, and craft.

$241,000–$270,000/yr
US Unlimited PTO

  • Architect the end-to-end reliability, performance, and resilience of cloud environments, including the SLO framework for critical services.
  • Lead incident response, on-call rotation, root cause analysis, and build a culture of corrective actions.
  • Build observability platforms to detect issues proactively and mentor engineers on reliability standards.

Garner is on a mission to transform the U.S. healthcare system by partnering with employers to steer members to better-performing doctors, resulting in better care and lower costs. With 550+ proprietary clinical metrics, they have helped over 2.5 million people and saved $1B in healthcare costs, recently raising a Series E and doubling five years running.

US

  • Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
  • Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
  • Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.

ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.

$200,000–$240,000/yr
US Unlimited PTO 16w maternity 16w paternity

  • You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.

Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.

US

  • Design and maintain AWS cloud infrastructure using OpenTofu and Terraform.
  • Operate Kubernetes workloads on Amazon EKS, managing GitOps deployments with Argo CD and Helm.
  • Implement observability with Datadog, troubleshoot production incidents, and support on-call rotation.

PAR Technology Corporation provides innovative restaurant technology solutions, including point-of-sale, digital ordering, loyalty, and back-office software, as well as hardware and drive-thru offerings. With over 40 years of experience, the company serves more than 100,000 restaurants globally and fosters a collaborative culture centered on its 'Better Together' ethos.

$97,976–$119,748/yr
Europe

  • Design, build, and operate AWS infrastructure across multiple regions using Kubernetes, Terraform, and Helm.
  • Own large-scale object storage environments exceeding 15 petabytes, optimizing reliability, performance, scalability, and cost.
  • Build an internal developer platform for self-service infrastructure, strengthen security practices, and lead incident response and postmortems.

Our partner builds an AI-powered sports media platform handling petabyte-scale media storage and millions of minutes of video monthly. It operates as a remote-first, international, engineering-led team with significant autonomy and a focus on high availability and security.

North America Canada Latin America

  • Design and operate scalable cloud infrastructure across AWS and GCP.
  • Build and improve Kubernetes, Linux, and cloud networking environments.
  • Strengthen security, disaster recovery, and platform resilience.

Hubstaff provides workforce analytics and time tracking for remote teams, serving over 200,000 global users. The company is a product-led organization with a winning culture and a fully remote team of experienced engineers.

$128,000–$176,000/yr
US

  • Design and implement infrastructure using Terraform, Python, and Kubernetes on AWS.
  • Collaborate with engineering and data science teams to improve cloud infrastructure.
  • Automate CI/CD pipelines and enforce security governance and compliance.

Lyra Health is a mental health care provider serving 20 million people through employer and health plan partnerships. The company has delivered 15 million sessions and published 35 peer-reviewed studies, with a culture focused on clinical effectiveness.

$140,000–$165,000/yr
Global Unlimited PTO

  • Design, build, and optimize cloud infrastructure (AWS/Kubernetes/EKS) and CI/CD pipelines across multiple teams.
  • Troubleshoot and resolve production incidents of varying scope, ensuring reliability and performance.
  • Drive infrastructure projects end-to-end, mentor engineers, and establish standards that improve developer productivity.

Pacvue is a leading Commerce Media OS powering over $12B in advertising spend across 100+ global retail media networks. It enables over 70,000 brands and agencies with an inclusive global community that fosters innovation and career growth.

Canada

  • Define and drive the vision and multi-quarter roadmap for the infrastructure foundations team, tying investments to business outcomes.
  • Lead and mentor a team of infrastructure engineers, fostering ownership, collaboration, and technical excellence.
  • Own the reliability and safety of the foundational AWS layer, including account provisioning, core networking, and IAM access.

Affirm is reinventing credit to make it more honest and friendly, offering consumers the ability to buy now and pay later without hidden fees or compounding interest. As a publicly traded company, Affirm fosters a culture of thorough technical design review, operational excellence, and capable incident response.

US

  • Design and maintain automated test frameworks for AWS infrastructure, infrastructure-as-code, and deployment workflows.
  • Validate infrastructure changes across VPCs, IAM, security groups, and containerized environments.
  • Integrate automated tests into CI/CD pipelines and collaborate with DevOps, Platform, and Compliance teams to strengthen release confidence.

Keeper Security is a cybersecurity software company protecting thousands of organizations and millions of people globally with zero-knowledge and zero-trust security solutions. It is one of the fastest-growing companies in the industry, recognized in the Gartner Magic Quadrant for Privileged Access Management.

US

  • Provide day-to-day technical support to internal software engineering teams on cloud infrastructure and platform-related issues.
  • Troubleshoot deployment, networking, infrastructure, and CI/CD challenges while performing root cause analysis.
  • Identify opportunities to automate manual tasks and improve engineering workflows using Terraform, Python, and Bash.

Software Mind is a team of engineers that provides project ramp-up services for top companies. They are a multicultural, growing company with an excellent work environment certified by Great Place To Work.

Canada

  • Design, implement, maintain, and optimize highly available infrastructure supporting mission-critical applications and services.
  • Monitor production environments, analyze system performance, and proactively identify opportunities to improve stability, scalability, and operational efficiency.
  • Respond to technical escalations, troubleshoot infrastructure, networking, hardware, and software issues, and lead resolution of critical incidents.

Our partner is a technology company focused on high-availability platforms and mission-critical infrastructure. The team is collaborative and works with modern cloud technologies.