Source Job

$100,000–$120,000/yr
Europe 6w PTO

  • Manage physical "Metal" environments from bare metal to Kubernetes, including cluster networking and scheduling.
  • Maintain Crossplane compositions and Terraform modules for cloud service provider resources.
  • Work with application teams to understand needs and invest in right capabilities.

Go Python Kubernetes Terraform Networking

20 jobs similar to Senior Platform Engineer - Platform Metal

Jobs ranked by similarity.

$87,480–$110,160/yr
Europe 6w PTO

  • Design, build, and operate reconciliation systems for Grafana Cloud stacks at scale.
  • Collaborate across teams to improve reliability, deployment complexity, and incident response.
  • Contribute to roadmap planning, technical design, and long-term simplification of stack operations.

Grafana Labs is the company behind the open source observability platform Grafana, providing a fully managed observability cloud. With over 1,600 team members across 40+ countries, the company fosters a global, collaborative culture rooted in open source principles.

Ireland 6w PTO

  • Lead automation of release processes including CI/CD, bootstrapping, and configuration management for internal engineering teams.
  • Work with diverse internal teams to implement requirements and maintain the Internal Engineering Platform.
  • Participate in an on-call rotation to support platform tooling and ensure system health.

Grafana Labs is the company behind the open source observability platform Grafana, providing a fully managed observability cloud called Grafana Cloud. With over 1,600 team members across 40+ countries, the company fosters a global collaborative culture rooted in open-source values and transparency.

Europe 6w PTO

  • Own features end-to-end across the stack, from backend Go services to frontend TypeScript/React.
  • Design and evolve systems for scheduling, probe lifecycle, and telemetry ingestion at scale.
  • Build cross-product and AI-assisted workflows that integrate Synthetic Monitoring into the Grafana Cloud platform.

Grafana Labs is the company behind the open source observability cloud, helping organizations see and act on their data. With over 1,600 employees across 40+ countries, the company fosters an open, collaborative culture.

$140,000–$220,000/yr
North America LATAM Europe

  • Own and scale the cloud infrastructure behind our open-source platform: compute, networking, and the data layer.
  • Lead BYOC: turn customer-cloud deployments into a real product, with provisioning, upgrades, and observability that scale past bespoke work per deal.
  • Make reliability a product feature: meaningful SLOs, and an incident process people trust.

Nango is a developer infrastructure company that provides API access for agents and apps, enabling AI applications to connect to the real world through integrations. With over 400 paying customers and a team of 14 from top tech companies like AWS, GitHub, and Okta, they are a YC-backed, multi-million ARR company that values ownership and autonomy.

$174,986–$209,983/yr
US Canada 6w PTO

  • Design and build core backend services for context ingestion, indexing, retrieval APIs, and agent-facing integrations.
  • Create a scalable multi-tenant SaaS foundation with tenant isolation, usage tracking, and reliable service boundaries.
  • Work across product and infrastructure to make practical tradeoffs between fast experimentation and long-term reliability.

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations. We are a 100% remote company with team members across 40+ countries, backed by leading investors, and we value an open-source legacy, global collaboration, and transparency.

UK Unlimited PTO 18w maternity 12w paternity

  • Own the technical strategy for multi-ecosystem scaling, defining architecture for onboarding new language ecosystems.
  • Drive end-to-end remediation automation, leading redesign of CVE workflows to close the loop from detection to verified release.
  • Set platform-wide technical direction spanning package index, build pipelines, and orchestration tooling to serve customers and ecosystem teams.

Chainguard is the trusted source for open source, delivering hardened, secure, and production-ready builds of open source software. They serve Fortune 500 enterprises and global industry leaders, and are venture-backed by leading investors, fostering a culture of customer obsession and intentional action.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

$250,000–$285,000/yr
US Unlimited PTO

  • Define architecture and best practices for the platform and infrastructure layer the product is built on.
  • Own the deploy pipeline and lead the move to a GitOps model (Argo) for fast, safe releases.
  • Design and harden multi-tenant isolation and blast-radius protection for top-tier customers, including dedicated deployments.

We are the Engineering Operations Platform - mission control for the AI software factory, providing visibility, governance, and golden paths. We are a group of 80 passionate individuals, backed by $60M Series C from Sequoia, IVP, and others, with a fully remote culture.

Global

  • Own and scale cloud infrastructure including compute, networking, storage, and data systems.
  • Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
  • Establish reliability through service-level objectives, observability, and incident response processes.

A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.

Ireland

  • Design, build, and deploy production systems with focus on scalability, reliability, and security.
  • Develop and maintain automation to streamline operations and eliminate toil.
  • Proactively monitor systems and implement automated incident response to minimize downtime.

Arista Networks is an industry leader in data-driven networking for large data centers, campus, and routing. With over $8 billion in revenue and a culture valuing diversity, Arista is a Great Place to Work for Best Engineering Team and Best Company for Diversity.

North America

  • Build and operate the control plane for automated cluster deployment from bare metal to customer-ready.
  • Manage machine lifecycle including joining, wiping, verifying, and rejoining between tenants.
  • Operate Kubernetes, Postgres, and custom operators across the fleet, scaling from tens to thousands of nodes.

Andromeda provides scaled AI infrastructure for startups, managing compute across numerous capacity providers. The company operates tens of thousands of GPUs for 80+ customers and fosters an inclusive environment.

$138,700–$173,400/yr
US

  • Design, build, and operate services and automations to manage Kubernetes clusters at scale, partnering with product management and technical leadership.
  • Drive rigorous code reviews and maintain high testing standards across the platform.
  • Manage cloud configurations across AWS and Azure using Terraform, ensuring deep observability and reliability.

Twilio is shaping the future of communications by delivering innovative solutions to hundreds of thousands of businesses and empowering millions of developers. They are a remote-first company with a strong culture of connection and global inclusion, employing a vibrant and diverse team.

US 6w PTO

  • Take an active role in influencing our roadmap and your own career objectives.
  • Design, build, operate, and maintain critical systems, owning reliability, performance, and availability.
  • Collaborate with your team to deliver new features and iterate based on results.

Grafana Labs is the company behind the open-source observability platform Grafana, providing a fully managed observability cloud. With over 1,600 team members across 40+ countries and 35 million users, the company thrives on a transparent, collaborative, and open-source culture.

US Canada 6w PTO

  • Ship features end to end, from the UI to backend services and dashboards.
  • Move between parts of the stack, working with React, Go, and Jsonnet as needed.
  • Own projects from problem statement to production, including user feedback and success measurement.

Grafana Labs is the company behind the open observability cloud, providing a flexible, scalable platform for monitoring and analyzing data. With over 1,600 team members across 40+ countries, the company fosters a remote-first, collaborative culture built on transparency, autonomy, and trust.

$150,000–$175,000/yr
US

  • Design and implement cloud infrastructure using AWS, Azure, and Terraform.
  • Manage Kubernetes clusters for container orchestration and deployment.
  • Collaborate with cross-functional teams to ensure scalability and reliability.

BSC Analytics is a leader in advanced data analytics for highly regulated enterprises, providing technical strategy and teams of exclusively senior talent. The company fosters a culture of senior expertise, tackling the toughest data challenges.

APAC

  • Design and deliver significant components and core subsystems of our Kubernetes platform, such as secrets management, workload identity, storage, or cluster networking, from design through production operation.
  • Contribute to the architecture of distributed workloads, working with dependent teams to get runtime and isolation models right, while spending most time hands-on in code.
  • Own operability of built systems including SLOs, failure modes, upgrades, migrations, and on-call, and mentor earlier-career engineers.

ServiceNow is the AI control tower for business reinvention, bringing together any AI, any data, and any workflow to help 85% of the Fortune 500 work smarter, faster, and better. The company fosters an AI-native culture where technology and talent are unstoppable together, with a focus on freeing people from busywork.

Europe

  • Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
  • Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
  • Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.

Slovakia

  • Own infrastructure as code using Terraform and Terragrunt for scalable, reliable cloud infrastructure.
  • Design and optimize CI/CD pipelines in GitLab and manage Kubernetes workloads for microservices.
  • Drive security, reliability, and collaborate with development teams to ensure compliance and performance.

Deutsche Telekom IT Solutions Slovakia provides innovative information and communication technology services. The company has grown to over 3900 employees and is the second largest employer in eastern Slovakia, promoting a culture of continuous improvement and work-life balance.

Europe

  • Own daily IT and platform operations, resolving access requests, deployments, and infrastructure tasks.
  • Manage cloud infrastructure on GCP and Cloudflare, CI/CD pipelines, and monitoring.
  • Collaborate with DevOps & Security lead to harden systems and scale the platform.

Centrifuge is building open infrastructure for real-world assets on blockchain, partnering with major financial institutions. We are a well-funded, small, high-trust team backed by leading investors, with over $1.7B in TVL.

US

  • Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms.
  • Own the storage layer where Kubernetes meets bare metal, tuning NFS data paths for high-throughput workloads.
  • Automate storage provisioning with infrastructure-as-code and GitOps, ensuring observability and reliability.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams across hybrid, edge, and sovereign environments, fostering a culture of open-source innovation and collaboration among passionate, talented colleagues.