Source Job

North America

  • Design, build, and operate production Kubernetes platforms, owning cluster architecture, networking, reliability, and security.
  • Troubleshoot complex infrastructure issues across Kubernetes, Linux, networking, and cloud environments, and improve observability and automation.
  • Take ownership of critical infrastructure initiatives, incident response, and mentorship while working autonomously in a remote-first environment.

Kubernetes Go Python Terraform AWS

20 jobs similar to Senior/Staff Platform Engineer

Jobs ranked by similarity.

Canada USA Unlimited PTO

  • Own the observability, logging and alerting for Kubernetes clusters and critical workloads.
  • Build and maintain automation for lifecycle management of Kubernetes clusters.
  • Identify and root-fix reliability bottlenecks before they become incidents.

Wrapbook is an AI platform for production finance, built for feature films and TV, trusted by Netflix and Paramount. Backed by top investors, our team of over 350 employees uses AI to transform how finance teams work.

India

  • Architect and scale multi-region microservices, APIs, and authentication infrastructure on AWS/GCP.
  • Lead SLOs, observability, incident management, and disaster recovery automation to maintain 99.99% availability.
  • Manage Kubernetes clusters and Terraform IaC while eliminating toil with Python/Go tooling.

JumpCloud is an AI-powered unified IT management platform that secures the modern workforce through identity, device, and access management. The company is remote-first with teams in 15+ countries and values building connections, thinking big, and continuous improvement.

$200,000–$220,000/yr
US 12w maternity 12w paternity

  • Design and enhance multi-region, multi-cloud infrastructure for reliability and capacity failover.
  • Author custom Kubernetes tools and infrastructure code using Terraform, Go, Python, or Ruby.
  • Diagnose and fix complex production issues and performance bottlenecks across cloud environments.

Huntress is a cybersecurity company founded in 2015 by former NSA cyber operators. They provide enterprise-grade security to businesses of all sizes, securing over 5 million endpoints and 11 million identities. Their remote-first team values resilience and collaboration.

Europe

  • Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
  • Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
  • Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.

Jobgether is a platform that connects job seekers with opportunities through an AI-powered matching process. The company fosters a remote-first culture and emphasizes autonomy and ownership for engineers.

$200,000–$240,000/yr
US Canada Unlimited PTO

  • Own best practices for managing production infrastructure, including provisioning, scaling, configuration, capacity planning, and monitoring.
  • Build and maintain Kubernetes infrastructure at scale alongside Terraform-provisioned cloud resources.
  • Write custom automation and tooling in Go to reduce manual work and eliminate operational risk.

Turnkey is building the infrastructure for the autonomous economy, providing programmable guardrails that enable organizations to operate with autonomy and control. Founded by the team behind Coinbase Custody, it is a deeply technical, low-ego, high-agency team of experts in cryptography, security, and systems.

EU

  • Own and operate production infrastructure across Kubernetes, Linux, networking, and virtualization.
  • Lead incident response and implement observability to improve availability and performance.
  • Define SLOs and automate infrastructure with Ansible, Bash, Python, and GitOps.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective, data-driven processes. They foster a collaborative, international, and fully remote work environment, emphasizing autonomy and ownership for their small to mid-sized team.

$126,000–$174,000/yr
US

  • Define DevOps strategy and lead infrastructure architecture across multi-environment, multi-region cloud systems.
  • Architect and own scalable Kubernetes platforms, infrastructure as code, and DevSecOps implementation.
  • Drive platform reliability, performance SLAs, cost optimization, and lead complex migrations and AI/ML platform infrastructure.

Robots & Pencils is an applied AI engineering firm that designs and ships AI co-workers for enterprise operations. Founded in 2009, the company has delivery centers across Canada, the US, Eastern Europe, and Latin America, with teams averaging over 15 years of experience.

India

  • Design, deploy, and maintain the reliability, availability, and performance of critical systems and APIs across AWS and GCP.
  • Build observability frameworks, define SLIs/SLOs, and implement monitoring using Datadog and Kubernetes.
  • Participate in on-call rotations, incident response, and blameless post-incident reviews to drive systemic improvements.

JumpCloud is an AI-powered unified IT management platform that secures the modern workforce by consolidating identity, device, and access management. The company is remote-first with teams in over 15 countries and values building connections, thinking big, and continuous improvement.

UK

  • Lead Cloud Platform and SRE teams to scale securely and efficiently.
  • Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
  • Champion SRE culture with SLOs, error budgets, and observability.

Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.

Global

  • Design and scale cloud-native infrastructure for a fast-growing bandwidth marketplace.
  • Lead platform reliability with deep expertise in Kubernetes, GCP, and container orchestration.
  • Drive infrastructure decisions across the full stack, from virtualization to production APIs.

Share is a venture-backed company building the marketplace for bandwidth, connecting internet capacity suppliers to businesses and individuals. The company is investor-backed with a high-ownership environment and a steep learning curve.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

$150,000–$165,000/yr
US Unlimited PTO

  • Design and manage high-availability platforms using Kubernetes, Terraform, and Ansible with native-AI capabilities.
  • Develop and operate the observability stack: Grafana, Mimir, Loki, Tempo, and Prometheus on Kubernetes via GitLab CI/CD.
  • Build automation scripts in Python, maintain GitOps pipelines, and mentor mid-level engineers.

Flexential builds and operates critical IT platforms including observability, DevOps, and ITSM technologies. The company fosters a collaborative engineering culture and values diversity.

UK Ireland Estonia Netherlands Sweden Israel Eastern Europe Portugal Unlimited PTO

  • Develop and evolve foundational software and services enabling product and development teams.\n- Architect, design, and implement Infrastructure as Code using Terraform.\n- Deploy, manage, and optimize Kubernetes clusters on GCP (GKE) and AWS (EKS).

DoiT is a global technology company that helps cloud-driven organizations leverage cloud for business growth and innovation through data, technology, and human expertise. They work with over 4,000 customers worldwide and foster a remote-first, entrepreneurial culture.

$170,000–$235,000/yr
US

  • Design and implement backend services for licensing, entitlements, feature access, and usage limits across NodeZero's product and APIs.
  • Build and evolve provisioning, admin experience, MSP/MSSP capabilities, and audit logging for a multi-tenant SaaS platform.
  • Operate production services with monitoring, incident response, and a high bar for design quality and test coverage.

Horizon3 is a fast-growing, remote cybersecurity company that helps organizations proactively find, fix, and verify exploitable attack vectors through its NodeZero autonomous pentesting platform. The team is a fusion of former special operations cyber operators and startup engineers, fostering a culture of respect, collaboration, ownership, and results.

$180,000–$220,000/yr
US

  • Design, implement, and maintain reliable, scalable, and secure infrastructure to support applications and automation systems.
  • Automate infrastructure provisioning, configuration management, and deployment pipelines using tools like Terraform and ArgoCD.
  • Implement observability solutions and enforce security best practices to ensure uptime and system performance.

Bright Machines is a next-generation, AI-enabled manufacturer focused on data center infrastructure production, using proprietary AI-based robotics and software to assemble hardware products for hyperscalers and OEMs. The company is headquartered in San Francisco, California, with an integration center in Guadalajara, Mexico, and has been recognized by Forbes' AI 50 and other leading organizations.

US

  • Design, build, and operate shared platform foundations including GCP, Kubernetes, networking, CI/CD, and observability.
  • Diagnose and troubleshoot complex distributed systems running at high request volume.
  • Raise the reliability bar through dashboards, alerting, on-call readiness, and automation.

Sanity.io builds an AI-powered content operating system that helps teams model, create, and automate content. The company has 200+ employees and a positive, flexible, trust-based culture that supports growth and work-life balance.

$150,000–$210,000/yr
Global Unlimited PTO

  • Own the design, development, and operation of infrastructure and build/release pipelines.
  • Deploy IaC and automation using Terraform, Ansible, Helm, and Go to support platform and customer requirements.
  • Collaborate closely with Product to drive roadmap direction and improve how users build and deliver software.

Manifest is on a mission to secure the global software and AI supply chain. Founded by alumni from the Department of Defense, CISA, and Palantir, it is a well-funded early-stage startup backed by leading investors and trusted by government and enterprise organizations.

Europe

  • Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
  • Establish and manage SLOs and SLAs for ClickHouse Cloud, ensuring monitoring and alerting are in place for all infrastructure.
  • Lead incident response, blameless postmortems, and chaos initiatives to continuously improve reliability and performance.

ClickHouse develops an open-source column-oriented database management system and offers a cloud database service. The company is a rapidly scaling, globally distributed startup with employees in over 25 countries, offering a flexible and collaborative culture.

  • Own the reliability, performance, and scalability of Runlayer's infrastructure across AWS and GCP.
  • Manage Kubernetes clusters, database reliability, and CI/CD pipelines for rapid deployments.
  • Lead incident response and partner with product engineers to design resilient systems for enterprise customers.

Runlayer builds a unified platform for MCPs, Skills, and AI Agents, providing enterprises with security, governance, and observability to deploy AI safely and at scale. Founded by engineers who built AI Actions for OpenAI and Zapier Agents, the team has raised $42M from Felicis and Khosla Ventures, serving companies like Gusto, Instacart, and Opendoor.

France

  • Evolve an Internal Developer Platform enabling development teams to deploy and operate applications securely with high self-service.
  • Play a crucial role in infrastructure architecture in a multi-cloud environment with a focus on GCP.
  • Build and maintain the platform used by over 50 internal clients, then support their concrete use by teams.

Lifen believes medical data can transform healthcare by reducing administrative burden, improving care coordination, and accelerating scientific discovery. Since 2015, the company has connected 800 hospitals and 150,000 healthcare professionals, with over 150 employees working remotely and from offices to unlock the potential of health data.