Source Job

North America

  • Own the operational health and reliability of Trino, Lightdash, Coder, and other platform services across development and production.
  • Lead the migration from Hive Metastore to Nessie and drive platform upgrades with zero disruption.
  • Manage and mentor a team of 3-5 platform engineers, setting standards for operational excellence and automation.

Kubernetes Python Observability

20 jobs similar to Senior Software Engineering Manager - FinOps Platform Services

Jobs ranked by similarity.

Canada

  • Influence technical strategy: Define and drive the long-term roadmap for the Lakehouse Platform, balancing scalability, reliability, and cost.
  • Design and develop platform capabilities for secure, discoverable, and easy-to-use analytical data across engineering, analytics, and ML teams.
  • Strengthen governance and access controls, improve analytics engineering foundations, and operate at scale with best practices for performance and cost optimization.

Affirm is reinventing credit to make it more honest and friendly, offering buy now, pay later solutions without hidden fees. With a remote-first culture, the engineering team focuses on building scalable, reliable infrastructure across multiple financial products.

UK

  • Lead Cloud Platform and SRE teams, driving infrastructure strategy and ownership including Kubernetes, GCP, and Terraform.
  • Champion SRE culture, define SLOs, SLAs, and enhance observability and incident management.
  • Own the developer-facing platform as a product, ensuring self-service infrastructure and security compliance (SOC2, ISO-27001).

Prolific builds human data infrastructure for AI development, connecting researchers with a global participant pool. They foster a culture of high performance, ownership, and cross-functional collaboration.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.

$100,000–$122,222/yr
Canada 8w PTO

  • End-to-end ownership of internal orchestration platform built on event-driven architecture with Redpanda, including code, architecture, and roadmap.
  • Own infrastructure-as-code using Terraform Cloud, manage Kubernetes workloads with Helm, and provide self-service tooling for engineering teams.
  • Set SLOs, handle production on-call, lead incident response, author design docs, and operate AI-natively using tools like Cursor and Notion AI.

Velora unifies Aplos, Raisely, and Keela into one company with a shared mission to help nonprofit organizations thrive by offering fundraising, donor management, financial tracking, and communications tools. We are a financially solid company with a combined team dedicated to making nonprofit work easier, more impactful, and more sustainable.

$217,000–$303,900/yr
US

  • Lead and grow the Site Defense teams to scale critical traffic and defense platforms.
  • Drive large migrations including CDN migration from Fastly to Cloudflare and Thrift to gRPC.
  • Maintain high availability and low latency while operating merged on-call responsibilities.

Reddit is a community of communities built on shared interests, passion, and trust, hosting the most open and authentic conversations on the internet. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit is one of the internet's largest sources of information.

Global

  • Own the reliability of Supabase's deployment and release systems against clear SLOs and error budgets.
  • Standardize and instrument pre-production deployment workflows for trustworthy signal.
  • Drive disaster-recovery readiness and reduce mean-time-to-detect and recover for deploy-related incidents.

Supabase is the Postgres development platform providing a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. They are a globally distributed team of ~400 members across 60+ countries, with over $1B raised and 540,000+ community members.

Canada

  • Build and operate core data platform capabilities for automated provisioning, deploy pipelines, and reliability tooling.
  • Develop integrations with CI/CD, identity management, APIs, and event-driven systems for self-service access.
  • Strengthen data access and governance with RBAC, data masking, and cross-database grants.

Affirm is reinventing credit to make it more honest and friendly, offering buy now, pay later solutions without hidden fees. The company is a remote-first organization with a strong focus on employee benefits and a people-first culture.

$134,982–$166,371/yr
US Canada

  • Lead and mentor a team of data and analytics engineers, including hiring, performance management, and career development.
  • Define the data engineering and analytics roadmap, prioritizing data platform investments and cross-functional initiatives.
  • Oversee data infrastructure design, reliability, and governance, ensuring accurate dashboards and analytics delivery.

Serve Robotics is reimagining urban deliveries with personable sidewalk robots. We are a team of tech industry veterans in software, hardware, and design, agile, diverse, and driven, focused on collaborative problem-solving.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.

US Canada Unlimited PTO

  • Lead a unified platform strategy defining and executing a cohesive roadmap aligning Production Engineering with Tenant Experience.
  • Drive operational excellence and reliability, maintaining deep accountability for GitLab’s production outcomes and incident management.
  • Scale engineering leadership by managing a senior management team, coaching on performance, hiring, and team structure.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With over 50 million registered users and a high-performance culture driven by values and continuous knowledge exchange, GitLab is where careers accelerate and innovation flourishes.

Canada Unlimited PTO

  • Lead the Data & Observability Platform team to build scalable data capabilities and improve engineering velocity.
  • Partner with Data & Analytics, SRE, and Cloud Platform Engineering to deliver reliable platform foundations and observability.
  • Drive modernization of legacy data platforms into cloud-native patterns on GCP, Databricks, and dbt.

Kinaxis is a global leader in modern supply chain orchestration, powering complex global supply chains with an AI-infused platform. With over 2000 employees worldwide and multiple Top Employer awards, we foster an innovative and collaborative culture.

$260,000–$280,000/yr
US Unlimited PTO

  • Lead and grow a team of backend and platform engineers migrating post-op processing to a real-time, event-driven architecture.
  • Partner with Product, Attack Engineering, and adjacent Platform teams to sequence the rework for continuous customer value.
  • Invest in operational excellence with SLOs, observability, on-call, and incident response as the system becomes central to every pentest.

Horizon3.ai is a remote cybersecurity company that provides an autonomous pentesting platform called NodeZero. It is a fast-growing company with a culture of respect, collaboration, and ownership.

$137,000–$174,000/yr
Canada

  • Lead and mentor engineering teams to build scalable, high-impact digital experiences for promotional platforms.
  • Collaborate with product and business stakeholders to define roadmaps, prioritize initiatives, and ensure successful delivery.
  • Drive engineering excellence through architectural decisions, process improvements, and a culture of continuous learning.

Jobgether uses AI-powered matching to connect candidates with job opportunities. They leverage technology to streamline recruitment and support a large user base.

$250,000–$285,000/yr
US Unlimited PTO

  • Define architecture and best practices for the platform and infrastructure layer the product is built on.
  • Own the deploy pipeline and lead the move to a GitOps model (Argo) for fast, safe releases.
  • Design and harden multi-tenant isolation and blast-radius protection for top-tier customers, including dedicated deployments.

We are the Engineering Operations Platform - mission control for the AI software factory, providing visibility, governance, and golden paths. We are a group of 80 passionate individuals, backed by $60M Series C from Sequoia, IVP, and others, with a fully remote culture.

UK

  • Help define and mature Engineering Operations by improving application health visibility, service reliability, and operational analytics.
  • Build and implement scalable processes for Incident, Problem, and Change Management that engineers actually want to use.
  • Connect engineering systems, data, and teams to reduce fragmentation and improve operational visibility across the organization.

Turnitin is a recognized innovator in global education, developing learning integrity solutions that help educators and institutions uphold academic integrity. With over 16,000 academic institutions using our services in more than 185 countries, we foster a remote-first culture and a diverse community of colleagues across 35+ countries.

APAC

  • Design and deliver significant components and core subsystems of our Kubernetes platform, such as secrets management, workload identity, storage, or cluster networking, from design through production operation.
  • Contribute to the architecture of distributed workloads, working with dependent teams to get runtime and isolation models right, while spending most time hands-on in code.
  • Own operability of built systems including SLOs, failure modes, upgrades, migrations, and on-call, and mentor earlier-career engineers.

ServiceNow is the AI control tower for business reinvention, bringing together any AI, any data, and any workflow to help 85% of the Fortune 500 work smarter, faster, and better. The company fosters an AI-native culture where technology and talent are unstoppable together, with a focus on freeing people from busywork.

US 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
  • Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.

Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.

US

  • Lead the Analytics Engineering team to build a trusted semantic metrics layer powering decision-making across the company.
  • Partner cross-functionally with Engineering, Product, Data Science, Finance, and Executive Leadership to define data strategy and ensure scalable, accessible data.
  • Drive technical excellence in data modeling, governance, and infrastructure while managing and developing a high-performing team.

HighLevel is an AI-powered business operating system that provides agencies, entrepreneurs, and SMBs with infrastructure to build, automate, and scale their operations. With over 2,000 team members across 10+ countries, we operate as a global, remote-first organization built for speed and ownership, fostering a culture of initiative, clarity, and execution.

$235,000–$275,000/yr

  • Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.

Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.

UK 7w PTO

  • Act as the primary intake point for platform requests from product teams, running workshops to scope initiatives.
  • Own sprint planning, backlog refinement, and quarterly planning across multiple platform squads.
  • Coordinate work across five engineering domains, managing priorities and improving delivery workflows.

Aveni builds production-ready AI systems for financial institutions, automating compliance and advisory workflows. Backed by Puma Private Equity and supported by Lloyds Banking Group, the company is scaling rapidly as a multiple fintech award winner.