Source Job

$136,750–$222,750/yr
Global 16w maternity 16w paternity

  • Build and operate internal platform services and APIs in Go.
  • Codify infrastructure with Terraform and GitOps practices.
  • Operate and scale multi-tenant EKS clusters and traffic systems.

Go Kubernetes Terraform Linux Networking

20 jobs similar to Software Engineer, Infrastructure Engineering

Jobs ranked by similarity.

US

  • Work as part of a small, cross-functional XP team installing Imogen into client cloud environments, partnering with client infosec, infrastructure, and IT teams.
  • Pair program with other engineers and collaborate closely with product managers and designers.
  • Lead technical discovery efforts for existing customer systems and adapt Imogen to their public cloud estate.

Mechanical Orchard specializes in safely rewriting critical business applications using a unique method that eliminates modernization risks. The company is known for its expertise in Agile practices and has a small, cross-functional team culture focused on collective ownership and continuous improvement.

$74,000–$148,000/yr
Global

  • Develop inter-cloud connectivity solutions for enterprise customers to use Sourcegraph Cloud.
  • Build a control plane to orchestrate a fleet of single-tenant Sourcegraph Cloud instances.
  • Participate in on-call rotation to uphold contractual SLA commitments and expose infrastructure as API.

Sourcegraph is the world's most powerful code intelligence platform that helps developers and agents navigate, understand, and operate on massive codebases. Backed by a16z, Sequoia, and Redpoint, we are a globally distributed team that values high agency, direct communication, and a deep love for developers and their craft.

$100,000–$122,222/yr
Canada 8w PTO

  • End-to-end ownership of internal orchestration platform built on event-driven architecture with Redpanda, including code, architecture, and roadmap.
  • Own infrastructure-as-code using Terraform Cloud, manage Kubernetes workloads with Helm, and provide self-service tooling for engineering teams.
  • Set SLOs, handle production on-call, lead incident response, author design docs, and operate AI-natively using tools like Cursor and Notion AI.

Velora unifies Aplos, Raisely, and Keela into one company with a shared mission to help nonprofit organizations thrive by offering fundraising, donor management, financial tracking, and communications tools. We are a financially solid company with a combined team dedicated to making nonprofit work easier, more impactful, and more sustainable.

UK

  • Lead projects with complete autonomy, planning ahead and thinking globally for the company's benefit.
  • Contribute to operation and deployment of all Dashboard Platform features using Kubernetes, Terraform, and AWS.
  • Share knowledge on software engineering, testing, and deployment best practices to enhance developer experience.

Algolia is a pioneer and market leader in AI Search, empowering 17,000+ businesses to deliver fast, predictive search and browse experiences. We have raised $150 million in Series D funding, value at $2.25 billion, and foster a high-trust, flexible workplace culture.

$185,000–$280,000/yr
US 4w PTO

  • Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
  • Scale single-tenant deployments and build observability, incident response, and compliance practices.
  • Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.

Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.

$160,000–$180,000/yr
US

  • Own the infrastructure and platform powering the marketplace, focusing on reliability, observability, security, and automation.
  • Manage production AWS and EKS clusters, infrastructure as code with Terraform and GitOps, and CI/CD pipelines via GitHub Actions.
  • Build automation and internal tooling in Python, Bash, Go, and Node.js/TypeScript, and operate PostgreSQL, MongoDB, and Temporal.

Office Hours is an on-demand expert network that connects leading organizations with trusted experts across various knowledge domains. The company is hyper-growth, profitable, and expanding quickly, backed by top marketplace investors.

Netherlands

  • Design and scale highly reliable platform systems supporting complex cloud-native workloads across multiple deployment environments.
  • Build and enhance core platform services while contributing to distributed systems, event-driven architectures, and cloud-native infrastructure.
  • Optimize cloud resources, networking, storage, compute, and observability to improve system performance, scalability, reliability, and maintainability.

Jobgether uses an AI-powered matching process to connect candidates with hiring companies. They operate as a job platform, processing applications and sharing top candidates with employers.

$190,000–$230,000/yr
Global Unlimited PTO

  • Design and operate the infrastructure for a high-throughput messaging platform operating at 500K+ events/sec.
  • Build guardrails, runbooks, and validation gates that enable AI agents to safely execute deployments and operations.
  • Lead incident response and encode every fix as a new runbook and regression test.

Postscript is an AI messaging platform trusted by 20,000+ Shopify brands to drive revenue through SMS. The company is fully remote, backed by Greylock and Y Combinator, and has a culture of ownership and innovation.

$165,000–$165,000/yr
US

  • Design, build, and maintain scalable cloud infrastructure services in AWS and Azure, owning automation and deployment pipelines.
  • Develop production-quality code in Go, Python, or similar, and implement Infrastructure as Code practices using Terraform, Packer, or similar tools.
  • Manage production release pipelines, ensure stability, monitor availability, and collaborate with engineering teams on observability and security practices.

Dragos is a global leader in OT cybersecurity, combining technology, threat intelligence, and expert services to protect critical infrastructure systems. The company is a remote-first mission-driven team across North America, Europe, the Middle East, and APAC, built on authenticity, transparency, and trust.

Latin America

  • Build and operate the self-service infrastructure platform where developers and agents can validate changes in minutes.
  • Build golden paths for CI/CD, GitOps, and IaC to enable self-service provisioning and shipping.
  • Own reliability and observability, carrying on-call and turning recurring toil into automation.

Luxury Presence is building the AI growth platform for real estate. Backed by Bessemer Venture Partners, the company is a Series C firm with over 90,000 real estate professionals and has been ranked on the Inc. 5000 fastest-growing companies list three years in a row.

$180,000–$195,000/yr
US

  • Architect, build, and evolve secure, scalable cloud infrastructure on AWS and AWS GovCloud to power the Hypori SaaS platform.
  • Independently own ambiguous, high-impact infrastructure problems, guide technical direction, and act as a senior escalation point during production incidents.
  • Drive strategy and execution of Infrastructure as Code, observability, CI/CD, and operational frameworks while mentoring engineers and raising the technical bar across the organization.

Hypori Inc. is a high-growth cybersecurity SaaS company transforming secure mobility through a virtual workspace platform that enables users to access enterprise apps and data from any mobile device with zero data on the endpoint and total personal privacy. Backed by $55M in funding from investors including UBS, AE Industrial Partners, Hale Capital Partners, and GreatPoint Ventures, the company is expanding into new commercial and regulated markets.

Global Unlimited PTO

  • Build enterprise-scale infrastructure using infrastructure-as-code and Kubernetes-native systems.
  • Sustain platform health and performance by owning critical systems in production.
  • Enable teams and customers to move faster with abstractions and tooling for AI/ML workloads.

Cake makes cutting-edge AI accessible to enterprise teams by removing infrastructure barriers, enabling 10x faster and cheaper AI/ML platform deployment. Backed by top investors, they have a small senior team focused on ownership and operational excellence.

US

  • Drive the definition and adoption of SLIs and SLOs across services, reducing toil through automation and incident response.
  • Design and architect Infrastructure as Code solutions for large-scale environments using Docker, Kubernetes, and cloud-native services.
  • Serve as primary SRE liaison for development teams, influencing architecture and conducting training for clients.

Noctua Technology, LLC is a company that drives digital transformation by treating operations as a software engineering challenge, focusing on cloud native systems. They are a dynamic team seeking a Senior SRE to define strategy and bridge development and operations for clients.

Poland

  • Design, write and deliver software to implement and support large web-scale, highly-performant, highly-available infrastructure on GCP/AWS.
  • Monitor infrastructure, respond to incidents, correct and improve systems to prevent incidents, and plan capacity.
  • Tune large-scale clusters for optimal performance and efficiency and support system deployments and product releases.

OpenX develops digital advertising marketplaces and technologies to optimize ad delivery for publishers and advertisers. The company operates a large-scale cloud infrastructure in Poland and values teamwork, customer centricity, and continuous learning.

$114,000–$142,000/yr
US Unlimited PTO

  • Design and operate a scalable, resilient ingress data plane that directly impacts customer value.
  • Leverage advanced automation and infrastructure-as-code to accelerate development speed across a massive fleet.
  • Collaborate with Product, Design, and partner platform teams to champion initiatives that break down silos and foster cross-team excellence.

New Relic provides an intelligent observability platform that helps companies optimize their digital applications. They are a global team of innovators focused on shaping the future of observability, with a culture of diversity and inclusion.

US

  • Lead deployment and operation of product infrastructure in federal environments within AWS.
  • Build and maintain scalable, secure cloud-native platforms using Kubernetes, Terraform, and GitLab CI.
  • Improve development and deployment processes, create tooling for telemetry, and foster documentation culture.

Horizon3.ai is a fast-growing, remote cybersecurity company that helps organizations proactively find and fix exploitable attack vectors. We are a team of former special ops cyber operators and engineers committed to a culture of respect, collaboration, ownership, and results.

US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

APAC Singapore Hong Kong

  • Design and maintain highly available cloud infrastructure across AWS, GCP, and Azure to support blockchain services and distributed systems.
  • Automate infrastructure and improve system reliability using Terraform, Golang, Python, and CI/CD pipelines.
  • Operate Kubernetes clusters and middleware platforms like Kafka, Redis, and NGINX while ensuring observability and disaster recovery.

BNB Chain is a community-first and open-source blockchain ecosystem focused on mass adoption through permissionless and decentralized principles. With a collaborative and dedicated team, it aims to onboard a billion new users to Web3.

US Europe

  • Design and maintain versioned REST and gRPC APIs for bare-metal server lifecycle, MachineType definitions, and cluster CRUD operations.
  • Develop asynchronous workflows for server enrollment, OS provisioning, and cluster bring-up, exposing durable status to callers.
  • Implement and maintain consoles and interfaces that visualize hardware inventory, provisioning progress, and cluster health.

Mirantis is a Kubernetes-native AI infrastructure company that enables organizations to build scalable, secure, and sovereign infrastructure for AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with a young, high-energy environment where collaboration, risk-taking, and continuous growth are valued.