Source Job

$190,000–$225,000/yr
US

  • Design, build, and operate core cloud infrastructure on AWS, including compute, networking, and container orchestration.
  • Own the CI/CD platform used across engineering teams, including build pipelines, environment promotion, and progressive rollout.
  • Build and maintain the observability stack across the organization, including logging, metrics, distributed tracing, and alerting.

Python AWS Kubernetes Terraform CI/CD

20 jobs similar to Principal Platform Engineer

Jobs ranked by similarity.

Global

  • Manage and optimize multi-cloud infrastructure (AWS required, GCP optional) with Kubernetes and CI/CD pipelines.
  • Improve observability through monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Coralogix).
  • Drive automation and Infrastructure as Code (IaC) using Terraform and Helm, and provide architectural guidance.

NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior. In 2023, NIQ combined with GfK, bringing together two industry leaders with operations in 100+ markets and covering more than 90% of the world's population.

US

  • Design, implement, and maintain AWS cloud infrastructure using Terraform.
  • Build and optimize CI/CD pipelines to enable rapid, safe deployments across multiple environments.
  • Own observability strategy with comprehensive monitoring, logging, and alerting systems using Datadog.

Avantos is building an AI-native operating system for financial services, transforming fragmented data into a single intelligent system. They are a product-led, fast-moving team at the intersection of AI, fintech, and modern infrastructure.

LATAM

  • Own and evolve Kubernetes and cloud infrastructure on AWS for scalability, reliability, and usability.
  • Design and improve CI/CD pipelines and developer workflows to enable fast, safe, repeatable deployments.
  • Work cross-functionally with product engineers to understand needs and enable them through tooling and best practices.

Artsy is an online platform that connects collectors, artists, and gallerists to make the art world more accessible. The company values an inclusive culture and a diverse workforce, with a team that operates with open-source principles and a focus on impact.

$180,000–$195,000/yr
US

  • Architect, build, and evolve secure, scalable cloud infrastructure on AWS and AWS GovCloud to power the Hypori SaaS platform.
  • Independently own ambiguous, high-impact infrastructure problems, guide technical direction, and act as a senior escalation point during production incidents.
  • Drive strategy and execution of Infrastructure as Code, observability, CI/CD, and operational frameworks while mentoring engineers and raising the technical bar across the organization.

Hypori Inc. is a high-growth cybersecurity SaaS company transforming secure mobility through a virtual workspace platform that enables users to access enterprise apps and data from any mobile device with zero data on the endpoint and total personal privacy. Backed by $55M in funding from investors including UBS, AE Industrial Partners, Hale Capital Partners, and GreatPoint Ventures, the company is expanding into new commercial and regulated markets.

$200,000–$215,000/yr
US

  • Design and deploy a modern observability stack including logging, metrics, and distributed tracing across hundreds of services.
  • Build alerting policies and incident response workflows to reduce manual escalations and improve mean time to detect and resolve.
  • Automate toil and set SRE standards while mentoring engineers on observability tooling.

WeightWatchers is a global digital health company and the world's #1 doctor-recommended behavioral weight health program. They have a cloud architecture supporting hundreds of services and are expanding their platform engineering team.

$118,000–$151,000/yr
US 4w PTO

  • Build and improve platform services, including CI/CD pipelines and cloud infrastructure.
  • Collaborate with senior engineers to design scalable solutions and enhance developer experience.
  • Participate in incident response and retrospectives to drive continuous improvement.

Octopus Energy is a tech-powered energy company focused on renewable energy and customer experience. The company culture emphasizes ownership, collaboration, and making a tangible impact across teams.

$122,000–$129,200/yr
US Unlimited PTO 16w maternity 4w paternity

  • Build, operate, and scale AWS infrastructure supporting a microservices platform.
  • Develop and maintain CI/CD pipelines to make shipping fast, reliable, and safe.
  • Partner with product engineering teams, translating platform concepts into practical guidance.

Storyblocks is a content company that provides video, audio, images, and creative tools through a subscription model, empowering storytellers. They are a remote-first company with a culture driven by data, empowered communication, and ownership, recognized as a top workplace.

US

  • Design and implement monitoring and alerting systems using tools like Prometheus, Grafana, and DataDog to ensure high availability and reliability.
  • Optimize performance and reliability of healthcare payment applications, lead incident response, and develop SLOs/SLIs.
  • Automate CI/CD pipelines, infrastructure provisioning with Terraform, and manage cloud infrastructure on AWS with Kubernetes.

LMI is a digital solutions provider accelerating government impact with innovation and speed, bringing commercial-grade platforms and mission-ready AI to federal agencies. Headquartered in Tysons, Virginia, LMI serves the defense, space, healthcare, and energy sectors, focusing on agility and collaboration to drive impactful results.

$185,000–$280,000/yr
US 4w PTO

  • Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
  • Scale single-tenant deployments and build observability, incident response, and compliance practices.
  • Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.

Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.

$160,000–$180,000/yr
US

  • Own the infrastructure and platform powering the marketplace, focusing on reliability, observability, security, and automation.
  • Manage production AWS and EKS clusters, infrastructure as code with Terraform and GitOps, and CI/CD pipelines via GitHub Actions.
  • Build automation and internal tooling in Python, Bash, Go, and Node.js/TypeScript, and operate PostgreSQL, MongoDB, and Temporal.

Office Hours is an on-demand expert network that connects leading organizations with trusted experts across various knowledge domains. The company is hyper-growth, profitable, and expanding quickly, backed by top marketplace investors.

Global Unlimited PTO

  • Own Primer's internal developer platform end to end, including CI/CD pipelines, deployment workflows, and self-service tooling.
  • Build the human-AI development loop, creating tooling and automation for coding agent workflows.
  • Treat developer productivity as a measurable system, using frameworks like DORA to identify and fix delivery bottlenecks.

Primer provides a unified infrastructure for global payments, enabling finance and payments teams to reduce complexity and capture revenue. Backed by top investors like Sofina and Accel, they operate as a remote-first, async culture with high autonomy and low bureaucracy.

Global

  • Design and maintain AWS infrastructure using Terraform, with a focus on scalability cost and PCI-scoped network segmentation
  • Build and evolve the observability stack and CI/CD pipelines to ensure smooth production operations and rapid deployment
  • Lead incident response define SLOs and run performance tests to optimize payment-critical services

Xplor Technologies provides vertical software, embedded payments, and AI tools for membership-based and service-based industries. With over 130,000 businesses in 72+ countries and processing $47 billion in payments annually, the company values diversity, collaboration, and a people-first culture.

United States

  • Design and build scalable, secure cloud infrastructure and deployment pipelines.
  • Own infrastructure end-to-end, including architecture, provisioning, deployment, and operation.
  • Lead technical design discussions and contribute to infrastructure and platform architecture decisions.

VulnCheck is the Exploit Intelligence Company, delivering structured exploit intelligence for cybersecurity. Founded in 2021, the company has a transparent, collaborative, and supportive culture with a team of experts.

Global

  • Build and manage AWS infrastructure, CI/CD pipelines, and ensure reliability, security, and cost optimization without an infrastructure team above you.
  • Work directly with stakeholders to shape architecture and product direction, shipping fast to learn or slowing down to fix the foundation.
  • Adopt the latest AI improvements to speed up how we build and ship, while collaborating with a remote-first team that meets regularly.

Cello is an AI-powered referral infrastructure company that helps SaaS companies turn users into a sales channel through sharing and recommendations, with APIs and widgets used by 10M+ end users monthly. Backed by $10.3M in funding from top-tier investors, the small team of serial founders and big-tech operators from Twilio, Wise, and Skype values active learning and innovation.

$190,000–$225,000/yr
US

  • Build and maintain end-to-end deployment pipelines for AI-powered applications, including artifact builds, environment promotion, rollback, and observability hooks.
  • Stand up and operate the runtime and lifecycle infrastructure for production agents, including deployment, versioning, monitoring, rate-limiting, and retirement.
  • Design and build the shared developer harness that every AI-powered service uses: prompt management, model routing, retries, tracing, eval hooks, and policy enforcement.

RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company offers innovative solutions with integrated intelligence on a single enterprise platform, connecting the pharmacy ecosystem.

Canada Europe

  • Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
  • Partnering with development teams to establish production readiness and operational readiness.
  • Building tooling to automate observability and operational workflows, eliminating manual toil.

Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.

$109,000–$194,000/yr
US 4w PTO 16w maternity 8w paternity

  • Build, deploy, and maintain scalable, highly available systems on AWS and own CI/CD pipelines and infrastructure as code.
  • Improve system reliability with alerting, runbooks, and observability using Grafana, Loki, Mimir, and Tempo.
  • Mentor junior engineers, collaborate on infrastructure design, and participate in on-call rotation with rare off-hours incidents.

Waymark is a mission-driven team of healthcare providers, technologists, and builders working to transform care for people with Medicaid benefits. We partner with communities to deliver technology-enabled, human-centered support that helps patients stay healthy and thrive.

$164,000–$218,000/yr
US Unlimited PTO

  • Lead design and evolution of secure cloud infrastructure and deployment systems for critical decentralized applications.
  • Drive improvements across CI/CD pipelines, deployment workflows, and engineering productivity practices.
  • Collaborate with developers, security specialists, product leaders, and infrastructure teams in a remote-first environment.

Our partner is building and scaling secure, high-performance infrastructure powering one of the most widely used decentralized technology platforms in the world. They operate as a fully remote, globally distributed team with a focus on DevOps, security, and blockchain technology.

$180,000–$210,000/yr
US Unlimited PTO

  • Lead the design and development of automated, resilient platform technologies including Observability, DevOps, and ITSM. - Manage a team of platform engineers, driving technical roadmaps and ensuring platform reliability and security. - Build and operate OpenTelemetry observability platforms using LGTM stack on Kubernetes.

Flexential is a data center and IT services company building next-gen observability platforms for 40+ data center facilities. They value diversity and offer a collaborative culture focused on innovation.

$190,000–$230,000/yr
Global Unlimited PTO

  • Design and operate the infrastructure for a high-throughput messaging platform operating at 500K+ events/sec.
  • Build guardrails, runbooks, and validation gates that enable AI agents to safely execute deployments and operations.
  • Lead incident response and encode every fix as a new runbook and regression test.

Postscript is an AI messaging platform trusted by 20,000+ Shopify brands to drive revenue through SMS. The company is fully remote, backed by Greylock and Y Combinator, and has a culture of ownership and innovation.