Source Job

$140,000–$220,000/yr
North America LATAM Europe

  • Own and scale the cloud infrastructure behind our open-source platform: compute, networking, and the data layer.
  • Lead BYOC: turn customer-cloud deployments into a real product, with provisioning, upgrades, and observability that scale past bespoke work per deal.
  • Make reliability a product feature: meaningful SLOs, and an incident process people trust.

Kubernetes AWS Terraform Postgres TypeScript

20 jobs similar to Staff Engineer, Platform & Infrastructure

Jobs ranked by similarity.

$200,000–$240,000/yr
US Unlimited PTO 16w maternity 16w paternity

  • You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.

Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.

$250,000–$285,000/yr
US Unlimited PTO

  • Define architecture and best practices for the platform and infrastructure layer the product is built on.
  • Own the deploy pipeline and lead the move to a GitOps model (Argo) for fast, safe releases.
  • Design and harden multi-tenant isolation and blast-radius protection for top-tier customers, including dedicated deployments.

We are the Engineering Operations Platform - mission control for the AI software factory, providing visibility, governance, and golden paths. We are a group of 80 passionate individuals, backed by $60M Series C from Sequoia, IVP, and others, with a fully remote culture.

US

  • Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
  • Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
  • Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.

ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

$186,700–$255,000/yr
US Canada 18w maternity 12w paternity

  • Create and test reliable cloud infrastructure services supporting Webflow's product range.
  • Lead initiatives to reduce triage load, increase reliability, and handle growing customer scale.
  • Collaborate with product engineering teams to deliver new solutions and improve existing services.

Webflow is an agentic web marketing platform that helps modern marketing teams build, manage, and optimize high-performing web experiences. The company values grit, speed, and craft, fostering a culture of ownership and continuous improvement.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

$180,000–$250,000/yr
US Unlimited PTO

  • Own the technical direction and architecture of critical infrastructure domains, establishing scalable patterns and standards.
  • Lead complex, multi-team infrastructure initiatives from design through implementation and production operation.
  • Design and evolve AWS and Kubernetes infrastructure to enable teams to build and deploy systems reliably at scale.

We provide innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. Our company is backed by world-class investors including Craft Ventures and Andreessen Horowitz, with offices across the US and India, and we are growing extremely quickly.

$53,300–$119,850/yr
Global Unlimited PTO 16w maternity 16w paternity

  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.

Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.

US

  • Improve system availability, scalability, and resilience across Flowcode's platforms.
  • Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
  • Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.

Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.

APAC

  • Design and deliver significant components and core subsystems of our Kubernetes platform, such as secrets management, workload identity, storage, or cluster networking, from design through production operation.
  • Contribute to the architecture of distributed workloads, working with dependent teams to get runtime and isolation models right, while spending most time hands-on in code.
  • Own operability of built systems including SLOs, failure modes, upgrades, migrations, and on-call, and mentor earlier-career engineers.

ServiceNow is the AI control tower for business reinvention, bringing together any AI, any data, and any workflow to help 85% of the Fortune 500 work smarter, faster, and better. The company fosters an AI-native culture where technology and talent are unstoppable together, with a focus on freeing people from busywork.

US

  • Own, operate and evolve our Cloud and Kubernetes based Platform.
  • Collaborate with Product teams to enable and empower them to build and own their services.
  • Build tooling and AI powered integrations that reduce cognitive load of complex infrastructure operations.

AlphaSense provides AI-driven market intelligence and search to help companies remove uncertainty from decision-making. Founded in 2011, the company has over 2,000 employees globally and is headquartered in New York City.

US

  • You'll contribute to infrastructure scaling to infinitely many apps, improving performance and reliability across backend services.
  • You'll support observability efforts, help implement SLOs, and build foundational services for next-generation cloud infrastructure.
  • You'll participate in triage and on-call processes to diagnose issues and implement changes to prevent recurrence.

Bubble is an AI visual development platform that empowers anyone to create software without code, from first-time entrepreneurs to enterprise teams. With over 6 million users in more than 100 countries and a mission to break down barriers to entrepreneurship, the company fosters a collaborative and inclusive culture focused on empowering builders worldwide.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

Argentina

  • Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
  • Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
  • Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.

Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.

North America Canada Latin America

  • Design and operate scalable cloud infrastructure across AWS and GCP.
  • Build and improve Kubernetes, Linux, and cloud networking environments.
  • Strengthen security, disaster recovery, and platform resilience.

Hubstaff provides workforce analytics and time tracking for remote teams, serving over 200,000 global users. The company is a product-led organization with a winning culture and a fully remote team of experienced engineers.

$126,290–$190,000/yr
United States 18w maternity 12w paternity

  • Empower engineers on other teams by maintaining monitoring tooling and collaborating on observability best practices.
  • Enhance reliability of Kubernetes applications through resource optimization, streamlined upgrades, and scalability.
  • Participate in on-call and incident response processes, occasionally diving into application code to debug production issues.

Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. It serves over 2 million users worldwide across 190 countries, with tens of thousands of projects launched each month, and fosters a culture of grit, speed, and craft.

Global

  • Own and evolve Quansight's cloud infrastructure across AWS, Azure, and GCP.
  • Lead infrastructure engagements for clients from scoping through delivery.
  • Contribute to open-source projects and participate in upstream communities.

Quansight is rooted in the Python data science community and helps companies build sustainable solutions on open-source software. The team is a small, collaborative, fully distributed group of open-source maintainers and engineers.

US Unlimited PTO

  • Building and coaching a high-performing distributed team with a shared operating model.
  • Owning platform capabilities for provisioning, deployment, and operations of infrastructure.
  • Leading infrastructure migration towards a modern SaaS model with incremental delivery.

Totara is a global learning platform trusted by more than 1,500 organisations and 21 million users worldwide, offering flexible learning, compliance, and talent development solutions. With a distributed team across New Zealand, Australia, the UK, and the US, the company values diverse perspectives and offers flexible, hybrid working.

$138,700–$173,400/yr
US

  • Design, build, and operate services and automations to manage Kubernetes clusters at scale, partnering with product management and technical leadership.
  • Drive rigorous code reviews and maintain high testing standards across the platform.
  • Manage cloud configurations across AWS and Azure using Terraform, ensuring deep observability and reliability.

Twilio is shaping the future of communications by delivering innovative solutions to hundreds of thousands of businesses and empowering millions of developers. They are a remote-first company with a strong culture of connection and global inclusion, employing a vibrant and diverse team.

$150,000–$250,000/yr
US Europe Singapore

  • Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
  • Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
  • Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.

A fast-growing AI/ML platform startup building infrastructure for training, evaluating, and aligning AI models within reinforcement learning environments. The engineering team of ~15 includes competitive programming medalists, serial AI startup founders, and researchers published at top venues.