Source Job

$12,500–$20,800/mo
Turkey

  • Own the technical architecture and evolution of core infrastructure.
  • Engineer for scale and performance through capacity modeling and bottleneck diagnosis.
  • Participate in on-call rotation and drive technical recovery during incidents.

AWS Kubernetes Golang Terraform

20 jobs similar to Principal Infrastructure Engineer

Jobs ranked by similarity.

$180,000–$250,000/yr
US Unlimited PTO

  • Own the technical direction and architecture of critical infrastructure domains, establishing scalable patterns and standards.
  • Lead complex, multi-team infrastructure initiatives from design through implementation and production operation.
  • Design and evolve AWS and Kubernetes infrastructure to enable teams to build and deploy systems reliably at scale.

We provide innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. Our company is backed by world-class investors including Craft Ventures and Andreessen Horowitz, with offices across the US and India, and we are growing extremely quickly.

US

  • Improve system availability, scalability, and resilience across Flowcode's platforms.
  • Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
  • Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.

Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.

$145,000–$260,000/yr
US Canada Unlimited PTO

  • Design, build, and optimize multi-region, high-availability AWS infrastructure.
  • Drive resiliency and automation using GitOps, modern CI/CD, and Infrastructure as Code.
  • Build end-to-end telemetry and own incident management to harden reliability.

VGS is the world's leader in payment tokenization, trusted by the most innovative AI and Fortune 500 companies. They are a remote-first company with a culture of ownership, collaboration, and continuous learning.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

$97,976–$119,748/yr
Europe

  • Design, build, and operate AWS infrastructure across multiple regions using Kubernetes, Terraform, and Helm.
  • Own large-scale object storage environments exceeding 15 petabytes, optimizing reliability, performance, scalability, and cost.
  • Build an internal developer platform for self-service infrastructure, strengthen security practices, and lead incident response and postmortems.

Our partner builds an AI-powered sports media platform handling petabyte-scale media storage and millions of minutes of video monthly. It operates as a remote-first, international, engineering-led team with significant autonomy and a focus on high availability and security.

Indonesia

  • Improve the high availability of our Dockerized microservices to ensure platform stability.
  • Launch and manage services in Kubernetes to scale our global operations.
  • Develop and operate infrastructure via Terraform, promoting automation and standardization.

We are an API-first platform revolutionizing commerce by providing a connected suite of on and off-premise solutions for food and retail industries. We are a rapidly scaling SaaS unicorn with a diverse, international team that values transparency and innovation.

Global

  • Own and scale cloud infrastructure including compute, networking, storage, and data systems.
  • Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
  • Establish reliability through service-level objectives, observability, and incident response processes.

A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.

Romania

  • Design and advance core infrastructure for multi-cloud Kubernetes clusters and developer toolchains.
  • Automate operations and engineering tasks to improve productivity and reliability.
  • Build machine learning infrastructure to enable AI teams to train and deploy large-scale models.

Cresta provides an AI platform that transforms customer conversations into competitive advantages by combining conversational AI, real-time agent augmentation, and conversation intelligence. The company has raised over $270 million from top investors like a16z, Greylock, and Sequoia, and is led by AI industry veterans.

$140,000–$175,000/yr
US

  • Lead and grow a team of platform engineers, coaching them on infrastructure and cloud challenges.
  • Drive the platform roadmap, balancing reliability, cost, security, and developer experience with AWS and Kubernetes.
  • Partner cross-functionally to align platform priorities with business goals and ensure system reliability.

PerfectServe is a leading provider of clinical communication and physician scheduling solutions in the health IT space. The company has 400+ employees and 30,000+ customers, with over $100 million in annual revenue, and has received multiple Best in KLAS awards.

US Unlimited PTO

  • Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
  • Design and maintain infrastructure as code across multiple cloud providers.
  • Provide technical leadership and mentorship across the Systems Engineering team.

Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.

US

  • Drive complex infrastructure migrations and build platform tooling and automation across multiple production environments.
  • Support development teams by consulting on infrastructure needs and improving observability and incident response.
  • Provide operational support and maintain platform reliability through structured debugging and on-call rotations.

PENN Entertainment is North America's leading provider of integrated entertainment, sports content, and casino gaming experiences. We operate across numerous locations in North America and foster a culture that cares about career growth and skill expansion.

Canada

  • Design and operate scalable AWS infrastructure with containerization and orchestration tools.
  • Implement monitoring, logging, and infrastructure as code using Terraform.
  • Improve CI/CD pipelines and troubleshoot production issues in complex SDLC environments.

Sureify builds systems that support millions of users. It is a high-growth, engineering-driven SaaS company with a remote-first culture across the Americas.

Argentina

  • Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
  • Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
  • Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.

Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.

UK

  • Architect and build a robust, scalable, and highly available distributed infrastructure.
  • Build a cutting-edge cloud-native platform on top of the public cloud and automate cloud resource management.
  • Work closely with core database development and security teams to produce the SaaS offering.

ClickHouse is a real-time analytics and data warehousing company recognized on the Forbes Cloud 100 list. With over 4,000 customers and rapid growth, the company is a leader in its space.

$150,000–$250,000/yr
US Europe Singapore

  • Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
  • Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
  • Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.

A fast-growing AI/ML platform startup building infrastructure for training, evaluating, and aligning AI models within reinforcement learning environments. The engineering team of ~15 includes competitive programming medalists, serial AI startup founders, and researchers published at top venues.

$140,000–$220,000/yr
North America LATAM Europe

  • Own and scale the cloud infrastructure behind our open-source platform: compute, networking, and the data layer.
  • Lead BYOC: turn customer-cloud deployments into a real product, with provisioning, upgrades, and observability that scale past bespoke work per deal.
  • Make reliability a product feature: meaningful SLOs, and an incident process people trust.

Nango is a developer infrastructure company that provides API access for agents and apps, enabling AI applications to connect to the real world through integrations. With over 400 paying customers and a team of 14 from top tech companies like AWS, GitHub, and Okta, they are a YC-backed, multi-million ARR company that values ownership and autonomy.

$114,700–$195,000/yr
North America

  • Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
  • Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
  • Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.

Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.

Europe US

  • Design, build, and operate multi-region AWS infrastructure on Kuberneties with Terraform and Helm at 15PB+ scale.
  • Own high-availablity, event-driven architectures and cost optimization across the stack.
  • Drive developer experience, security, and incident response as the second platform team member.

ScorePlay is the AI-powered media infrastructure for sports, automating content operations for the world's biggest sports organizations. We are a 50-person remote-first team based in New York and Paris, growing 2x year over year with 98% retention.

India

  • Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
  • Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
  • Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.

Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.

US

  • Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
  • Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
  • Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.

Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.