Source Job

Canada USA Unlimited PTO

  • Own the observability, logging and alerting for Kubernetes clusters and critical workloads.
  • Build and maintain automation for lifecycle management of Kubernetes clusters.
  • Identify and root-fix reliability bottlenecks before they become incidents.

Kubernetes Terraform AWS Python Linux

20 jobs similar to Senior Platform Engineer I

Jobs ranked by similarity.

UK Ireland Estonia Netherlands Sweden Israel Eastern Europe Portugal Unlimited PTO

  • Develop and evolve foundational software and services enabling product and development teams.\n- Architect, design, and implement Infrastructure as Code using Terraform.\n- Deploy, manage, and optimize Kubernetes clusters on GCP (GKE) and AWS (EKS).

DoiT is a global technology company that helps cloud-driven organizations leverage cloud for business growth and innovation through data, technology, and human expertise. They work with over 4,000 customers worldwide and foster a remote-first, entrepreneurial culture.

Global

  • Own and scale cloud infrastructure including compute, networking, storage, and data systems.
  • Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
  • Establish reliability through service-level objectives, observability, and incident response processes.

A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.

Brazil 4w PTO

  • Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
  • Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
  • Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.

Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.

US Unlimited PTO

  • Build and operate the Kubernetes platform supporting AI test and evaluation frameworks.
  • Design infrastructure-as-code, GitOps workflows, and automated deployment pipelines.
  • Own platform reliability, observability, capacity planning, and operational readiness.

OpenTeams helps enterprises and governments build AI they control, govern, and evolve themselves. Founded by the creator of NumPy and SciPy, the company is built by people with deep roots across the open-source ecosystem and maintains a remote-first culture.

US

  • Drive complex infrastructure migrations and build platform tooling and automation across multiple production environments.
  • Support development teams by consulting on infrastructure needs and improving observability and incident response.
  • Provide operational support and maintain platform reliability through structured debugging and on-call rotations.

PENN Entertainment is North America's leading provider of integrated entertainment, sports content, and casino gaming experiences. We operate across numerous locations in North America and foster a culture that cares about career growth and skill expansion.

$150,000–$165,000/yr
US Unlimited PTO

  • Design and manage high-availability platforms using Kubernetes, Terraform, and Ansible with native-AI capabilities.
  • Develop and operate the observability stack: Grafana, Mimir, Loki, Tempo, and Prometheus on Kubernetes via GitLab CI/CD.
  • Build automation scripts in Python, maintain GitOps pipelines, and mentor mid-level engineers.

Flexential builds and operates critical IT platforms including observability, DevOps, and ITSM technologies. The company fosters a collaborative engineering culture and values diversity.

$180,000–$220,000/yr
US

  • Design, implement, and maintain reliable, scalable, and secure infrastructure to support applications and automation systems.
  • Automate infrastructure provisioning, configuration management, and deployment pipelines using tools like Terraform and ArgoCD.
  • Implement observability solutions and enforce security best practices to ensure uptime and system performance.

Bright Machines is a next-generation, AI-enabled manufacturer focused on data center infrastructure production, using proprietary AI-based robotics and software to assemble hardware products for hyperscalers and OEMs. The company is headquartered in San Francisco, California, with an integration center in Guadalajara, Mexico, and has been recognized by Forbes' AI 50 and other leading organizations.

$200,000–$220,000/yr
US 12w maternity 12w paternity

  • Design and enhance multi-region, multi-cloud infrastructure for reliability and capacity failover.
  • Author custom Kubernetes tools and infrastructure code using Terraform, Go, Python, or Ruby.
  • Diagnose and fix complex production issues and performance bottlenecks across cloud environments.

Huntress is a cybersecurity company founded in 2015 by former NSA cyber operators. They provide enterprise-grade security to businesses of all sizes, securing over 5 million endpoints and 11 million identities. Their remote-first team values resilience and collaboration.

US Unlimited PTO

  • Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
  • Design and maintain infrastructure as code across multiple cloud providers.
  • Provide technical leadership and mentorship across the Systems Engineering team.

Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.

$168,675–$229,900/yr
US

  • Establish and employ continuous integration and delivery (CI/CD) patterns for successful software solutions.
  • Design secure, operationally sound solutions across AWS, Azure, OpenShift, and IBM Cloud.
  • Implement observability stacks, manage Kubernetes clusters, and automate infrastructure with Terraform.

Conga unifies commercial operations by aligning pricing, quoting, contracting, rebates, and communications so companies run as connected, smarter businesses. With more than 10,000 customers worldwide, including over 50% of the Fortune 100, Conga fosters a collaborative culture where every voice is heard.

Canada

  • Design and operate scalable AWS infrastructure with containerization and orchestration tools.
  • Implement monitoring, logging, and infrastructure as code using Terraform.
  • Improve CI/CD pipelines and troubleshoot production issues in complex SDLC environments.

Sureify builds systems that support millions of users. It is a high-growth, engineering-driven SaaS company with a remote-first culture across the Americas.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

$150,000–$170,000/yr

  • Build and operate AWS infrastructure, ensuring security, reliability, and scalability.
  • Design and maintain deployment pipelines, observability, and incident response.
  • Collaborate with application engineers and contribute to platform architecture decisions.

inKind connects diners with restaurants through an app, offering 20% back on every meal. It partners with thousands of restaurants and has millions of users, fostering a win-win philosophy for diners and restaurants, with a small, growing team focused on hospitality and community.

Spain 5w PTO

  • Define SLIs, SLOs, and reliability targets for the platform.
  • Improve observability, alerting, and production readiness across services.
  • Automate operational work and support cloud/Kubernetes infrastructure.

Lodgify is a fast-growing scale-up in vacation rental technology, backed by $30M in funding. Headquartered in Barcelona, the 380+ person team of 60+ nationalities is passionate about transforming short-term rentals.

UK

  • Lead Cloud Platform and SRE teams to scale securely and efficiently.
  • Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
  • Champion SRE culture with SLOs, error budgets, and observability.

Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.

Romania

  • Design and advance core infrastructure for multi-cloud Kubernetes clusters and developer toolchains.
  • Automate operations and engineering tasks to improve productivity and reliability.
  • Build machine learning infrastructure to enable AI teams to train and deploy large-scale models.

Cresta provides an AI platform that transforms customer conversations into competitive advantages by combining conversational AI, real-time agent augmentation, and conversation intelligence. The company has raised over $270 million from top investors like a16z, Greylock, and Sequoia, and is led by AI industry veterans.

US Unlimited PTO

  • Own the technical direction of the AWS platform, building cost visibility tooling and managing Kubernetes on EKS.
  • Design and maintain Terraform modules for self-service infrastructure provisioning and standardize CI/CD across services.
  • Harden the platform alongside security, lead incident response, and mentor senior engineers through design and code review.

PlayOn powers high school sports ticketing, streaming, and fundraising through platforms like GoFan, NFHS Network, and MaxPreps. Backed by KKR, the company is a growth-stage leader focused on making high school sports more accessible and connected.

$64,021–$92,683/yr
Canada

  • Extend the self-service datastore platform with provisioning automation, guardrails, and paved paths for product engineering teams.
  • Ship observability, alerting, and backup/disaster recovery as built-in defaults for every datastore.
  • Convert recurring pull-in work into platform features or AI tooling that other teams can use directly.

Greenhouse provides a hiring software platform designed to make hiring work for everyone. They have an award-winning culture recognized by Fortune and Inc., and foster inclusivity, transparency, and accountability among their teams.

UK

  • Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
  • Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.

$191,000–$226,000/yr
US Unlimited PTO

  • Own the reliability, performance, and resilience of cloud environments (AWS, Kubernetes) and define SLOs across critical services.
  • Lead incident response, on-call rotation, and drive root cause analysis to ensure high production quality.
  • Build and maintain observability systems and automate operational toil using AI tools.

Garner partners with employers to redesign healthcare by using clinical metrics to identify top doctors and incentivize members to better care. The company has helped over 2.5 million people, saved $1B in costs, and doubled annually for five years, fostering a mission-driven, high-performance culture.