Source Job

$230,000–$270,000/yr
US 16w maternity 16w paternity

  • Lead the design, development, and maintenance of highly scalable infrastructure systems.
  • Drive the technical vision and roadmap for infrastructure teams.
  • Mentor and guide engineers, fostering a culture of continuous learning and improvement.

Java Python Go Kubernetes AWS

20 jobs similar to Staff Software Engineer, Infrastructure

Jobs ranked by similarity.

$180,000–$250,000/yr
US Unlimited PTO

  • Own the technical direction and architecture of critical infrastructure domains, establishing scalable patterns and standards.
  • Lead complex, multi-team infrastructure initiatives from design through implementation and production operation.
  • Design and evolve AWS and Kubernetes infrastructure to enable teams to build and deploy systems reliably at scale.

We provide innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. Our company is backed by world-class investors including Craft Ventures and Andreessen Horowitz, with offices across the US and India, and we are growing extremely quickly.

$186,700–$255,000/yr
US Canada 18w maternity 12w paternity

  • Create and test reliable cloud infrastructure services supporting Webflow's product range.
  • Lead initiatives to reduce triage load, increase reliability, and handle growing customer scale.
  • Collaborate with product engineering teams to deliver new solutions and improve existing services.

Webflow is an agentic web marketing platform that helps modern marketing teams build, manage, and optimize high-performing web experiences. The company values grit, speed, and craft, fostering a culture of ownership and continuous improvement.

US Unlimited PTO

  • Build systems that optimize cloud use and scale for expansion, contributing to system architecture and execution.
  • Develop effective partnerships across engineering and product teams, performing design reviews and collaborating on roadmaps.
  • Investigate and understand how to leverage Temporal's own software to power infrastructure at scale, while sharing design principles for reliable systems.

Temporal provides an open source programming model that simplifies code and makes applications more reliable. The company is a growing, values-driven organization with a collaborative and humble culture.

US Unlimited PTO

  • Design, build, and operate shared cloud infrastructure using AWS, Kubernetes, Terraform, Databricks, and Cloudflare.
  • Deliver SRE and DevOps initiatives to improve reliability, scalability, observability, and deployment safety.
  • Build reusable infrastructure modules, automation, and self-service workflows to reduce manual work and improve developer experience.

YipitData is the leading market research and analytics firm for the disruptive economy, recently raising up to $475M from The Carlyle Group at a valuation over $1B. We analyze billions of alternative data points daily and have been recognized as one of Inc’s Best Workplaces, cultivating a people-centric culture focused on mastery, ownership, and transparency.

US

  • Design and implement scalable cloud infrastructure to support growth.
  • Develop monitoring, alerting, and incident response for system reliability.
  • Automate deployment pipelines and ensure high availability and security.

Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.

Brazil

  • Develop, enhance, and maintain software solutions in complex technical environments, applying strong engineering practices and testing standards.
  • Contribute to scalable, resilient, secure architectures and integrate applications, APIs, and distributed systems with modern DevOps and DevSecOps practices.
  • Work with cloud environments (AWS, Azure, GCP) and Infrastructure as Code (Terraform, Docker, Kubernetes) to automate deployment and improve reliability.

$157,600–$267,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Lead a distributed team of engineers, mentoring them in technical and professional growth.
  • Manage technical execution, prioritization, and cross-team collaboration on infrastructure projects.
  • Foster an inclusive, high-performance team environment with continuous feedback and career development.

Pilot provides small businesses with dedicated finance experts and custom software for accurate bookkeeping and financial management. With over 3,000 customers and $170 million in funding, the company values high trust, ownership, and continuous iteration.

Canada

  • Define and drive the vision and multi-quarter roadmap for the infrastructure foundations team, tying investments to business outcomes.
  • Lead and mentor a team of infrastructure engineers, fostering ownership, collaboration, and technical excellence.
  • Own the reliability and safety of the foundational AWS layer, including account provisioning, core networking, and IAM access.

Affirm is reinventing credit to make it more honest and friendly, offering consumers the ability to buy now and pay later without hidden fees or compounding interest. As a publicly traded company, Affirm fosters a culture of thorough technical design review, operational excellence, and capable incident response.

$232,000–$290,000/yr
United States Canada 18w maternity 12w paternity

  • Set the long-term technical vision and architecture for backend systems powering experimentation, personalization, and analytics.
  • Architect and evolve high-throughput distributed systems and APIs for real-time data processing with low latency and high availability.
  • Lead complex, multi-team technical initiatives end-to-end, from problem framing through implementation and long-term operational ownership.

Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. They have a remote-first culture that values craft, quality, and moving fast, with a focus on innovation and collaboration.

US

  • Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
  • Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
  • Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.

ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.

US Unlimited PTO

  • Architect and operate distributed, fault-tolerant systems for a large-scale cloud data platform.
  • Lead optimization initiatives across compute, storage, networking, and infrastructure efficiency.
  • Collaborate with engineering teams and stakeholders to build tooling for cloud cost and resource visibility.

This company builds a highly scalable cloud data platform and focuses on infrastructure efficiency and multi-cloud environments. With operations across more than 20 countries, it fosters a globally distributed, remote-friendly culture that values ownership, collaboration, and innovation.

Global

  • Own and scale cloud infrastructure including compute, networking, storage, and data systems.
  • Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
  • Establish reliability through service-level objectives, observability, and incident response processes.

A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.

$101,167–$204,439/yr
US

  • Provide technical leadership and direction for multiple agile teams to implement software product roadmaps.
  • Design, develop, and maintain complex distributed software systems with high-quality code and test coverage.
  • Mentor junior engineers and collaborate with product managers to estimate effort and prioritize work.

Mercury Insurance is a company that helps people reduce risk and overcome unexpected events through insurance products. They are a midsize employer with a culture that values diverse perspectives and professional growth.

US Unlimited PTO

  • Lead the design and development of systems optimizing network traffic and scaling for global expansion.
  • Drive architectural decisions for high-impact projects, ensuring scalability and reliability.
  • Mentor and guide engineers, sharing best practices for building reliable and scalable networking systems.

Temporal is an open source programming model that simplifies code, making applications more reliable and enabling faster feature delivery. The company is growing, values curiosity, collaboration, and humility, and seeks to be the foundation of every developer's toolbox.

$107,500–$147,500/yr
US Unlimited PTO 24w maternity 24w paternity

  • Build scalable front-end and back-end services and solve distributed systems problems.
  • Collaborate with engineers and product managers in code reviews and architectural discussions.
  • Leverage AI tools to enhance code quality, testing, and team efficiency.

Smartsheet is a cloud-based work management platform that helps teams plan, track, and automate their work. The company fosters a collaborative engineering culture focused on innovation and using AI to improve productivity.

$145,000–$155,000/yr
US 12w maternity 12w paternity

  • Lead the design, implementation, and delivery of scalable, reliable software solutions throughout the software development lifecycle.
  • Utilize AI-assisted engineering tools to improve engineering efficiency, software quality, and developer productivity.
  • Mentor software engineers and QA engineers on engineering best practices, technical problem-solving, and modern development workflows.

Phreesia provides a SaaS platform that digitizes appointment check-in and offers tools to engage patients and improve healthcare efficiency. They are a five-time Modern Healthcare Best Places to Work winner and are committed to inclusion and employee experience.

$53,300–$119,850/yr
Global Unlimited PTO 16w maternity 16w paternity

  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.

Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.

North America Canada Latin America

  • Design and operate scalable cloud infrastructure across AWS and GCP.
  • Build and improve Kubernetes, Linux, and cloud networking environments.
  • Strengthen security, disaster recovery, and platform resilience.

Hubstaff provides workforce analytics and time tracking for remote teams, serving over 200,000 global users. The company is a product-led organization with a winning culture and a fully remote team of experienced engineers.

Global

  • Own and evolve Quansight's cloud infrastructure across AWS, Azure, and GCP.
  • Lead infrastructure engagements for clients from scoping through delivery.
  • Contribute to open-source projects and participate in upstream communities.

Quansight is rooted in the Python data science community and helps companies build sustainable solutions on open-source software. The team is a small, collaborative, fully distributed group of open-source maintainers and engineers.

US

  • Build and operate monitoring, tracing, alerting, and observability infrastructure for system reliability.
  • Drive platform security initiatives with preventative controls and resilient architecture.
  • Lead incident response and recovery, including root-cause analysis and preventative measures.

This role is with a partner company managing AI-powered products. They are a growing technology organization with a fully distributed US-based team and a collaborative culture focused on large-scale infrastructure and AI technology.