Source Job

  • Drive company-wide reliability strategy, standards, and best practices across LinkedIn Engineering.
  • Lead adoption of service criticality models to set reliability expectations based on business impact.
  • Partner with engineering teams to improve system design, reduce incident risk, and strengthen operational readiness.

Distributed Systems Site Reliability Engineering High Availability Observability

20 jobs similar to Principal Staff Software Engineer, Systems Infrastructure

Jobs ranked by similarity.

$235,000–$275,000/yr

  • Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.

Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.

US

  • Establish performance, throughput, latency, and capacity baselines for critical platform workflows.
  • Define and maintain SLOs, error budgets, dashboards, alerts, and reliability thresholds.
  • Lead load, stress, soak, spike, failure, and recovery testing in representative environments.

Tech Holding is a full-service consulting firm that delivers predictable outcomes and high-quality solutions to clients. The company was founded by experienced industry professionals who have held senior positions at startups to Fortune 50 firms, fostering a culture of deep expertise, integrity, transparency, and dependability.

$200,000–$350,000/yr
US

  • Define architecture for complex, high-impact systems.
  • Lead company-critical engineering initiatives.
  • Mentor senior and staff-level engineers.

They build sophisticated infrastructure and software systems for critical business operations. They are a rapidly scaling company with a collaborative, fast-moving culture where ownership is encouraged and decisions are made quickly.

US

  • Architect large-scale infrastructure systems for end-to-end software lifecycle platforms.
  • Design unified CI/CD and deployment platforms for build, test, deploy, and runtime operations.
  • Build observability, monitoring, and cost visibility systems for production infrastructure.

LinkedIn is the world's largest professional network, creating economic opportunity for every member of the global workforce. As a large company, LinkedIn fosters a culture of trust, care, inclusion, and fun, investing in employee growth.

Portugal

  • Work with teams to define SLIs and SLOs, and create systems for observability.
  • Analyze failure scenarios, create runbooks, and reduce work that does not add value.
  • Participate in incident management and facilitate on-call duty to ensure reliable production environments.

Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.

$240,000–$240,000/yr
US

  • Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
  • Drive SLOs, observability, alerting, and on-call processes across teams.
  • Build the platform engineering function from the ground up and influence cross-cutting architecture.

First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.

$200,000–$350,000/yr
US

  • Define technical strategy for critical systems.
  • Lead multi-team architecture and engineering initiatives.
  • Establish engineering standards and technical principles.

They build sophisticated infrastructure and software systems for critical business operations. They are a rapidly scaling company with a collaborative, fast-moving environment where ownership is encouraged.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

$175,000–$195,000/yr

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
  • Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
  • Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.

Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.

  • You own delivery across product, engineering, and QA, leading a team of roughly ten people.
  • You stay close to the work by participating in design reviews, reading PRs, and occasionally writing code.
  • You run the engineering system by unblocking dependencies, killing ambiguity early, and keeping the path from decision to production short and reliable.

Ocra is building the commercial developer platform for parking, acting as the connective layer between sellers, booking channels, and payment systems. They are a small, distributed team that values autonomy, curiosity, and adaptability over rigid processes.

US

  • Lead architectural direction and development of critical systems, services, and infrastructure with a focus on reliability, security, and performance.
  • Collaborate with product and program leaders to align engineering solutions with user needs, regulatory requirements, and business outcomes.
  • Coach and mentor engineers across teams, grow their craft, and strengthen engineering culture through curiosity, learning, and integrity.

Avandra Imaging unlocks clinical data to overcome medical challenges, bringing hope and breakthroughs that transform patient lives. With a national footprint of over 5,000 customer integrations and a market share of approximately 70%, they are poised to revolutionize healthcare research through a large indexed data cloud of medical imaging.

Poland

  • Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
  • Define and drive SRE platform strategy, incident management, and observability engineering.
  • Mentor team members, foster collaboration, and ensure operational excellence.

XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.

$118,800–$237,600/yr
Europe

  • Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
  • Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
  • Manage distributed systems, observability, incident response, and automation with a security-first mindset.

Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.

US Unlimited PTO

  • Lead the organization’s most consequential technical challenges, turning ambiguity into clear direction and measurable outcomes.
  • Simplify complex systems, strengthen engineering practices, and create scalable foundations that are easier to operate and evolve.
  • Influence decisions without formal authority, working closely with senior business and technology leaders.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use technology to ensure fair and efficient application reviews.

$200,000–$350,000/yr
US

  • Lead architecture and implementation of critical platform capabilities.
  • Own complex technical initiatives from conception through production.
  • Design scalable distributed systems and services.

They build sophisticated infrastructure and software systems. They are a rapidly scaling company with a collaborative, fast-moving environment where ownership is encouraged.

$269,500–$385,000/yr
Global

  • Lead the Bridge engineering organization covering Billing, IAM, Data, Operations, and Platform Infrastructure.
  • Architect scalable infrastructure and drive strategic platform evolution for Docker's products.
  • Manage a team of 30+ engineers and collaborate cross-functionally to support company growth.

Docker provides developer tooling trusted by over 20 million monthly users. The company is a globally distributed, remote-first team with offices in Seattle and Paris.

US

  • Ensure Kafka reliability through topic design, governance, and latency optimization.
  • Own Ceph reliability including operations, pool design, and capacity planning.
  • Build operational automation to reduce manual toil and speed up incident response.

PulsePoint sits at the intersection of healthcare and adtech, helping brands and agencies interpret health journey signals and unify digital determinants of health with real-world data. We are a 300+ employee post-acquisition business, one of the leading players in the US healthcare ad market.

US Canada Unlimited PTO

  • Lead a unified platform strategy defining and executing a cohesive roadmap aligning Production Engineering with Tenant Experience.
  • Drive operational excellence and reliability, maintaining deep accountability for GitLab’s production outcomes and incident management.
  • Scale engineering leadership by managing a senior management team, coaching on performance, hiring, and team structure.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With over 50 million registered users and a high-performance culture driven by values and continuous knowledge exchange, GitLab is where careers accelerate and innovation flourishes.

$280,000–$330,000/yr
North America Unlimited PTO

  • Own the health, scalability, and reliability of Pipe's systems end-to-end, including APIs, SDKs, and infrastructure.
  • Lead about 10 engineers through squad leads, setting technical standards and building a culture of ownership.
  • Drive incident response, monitoring, and deployment practices to maintain high velocity and quality.

Pipe builds embedded capital for small businesses, delivering credit products inside the software platforms SMBs already use. The company is remote-first, revenue has accelerated nearly 2.5x year on year, and they foster a high-bar, ego-free culture.

US Unlimited PTO

  • Define and drive the technical vision for Implementation Platform Configuration and Enrollment Platform Core, ensuring scalability and reliability.
  • Lead and mentor a group of Engineering Managers, owning delivery commitments and incident posture across teams.
  • Drive AI adoption and testing strategies to scale platform reliability and reduce manual effort.

Bestow is a vertical technology platform that modernizes life insurance infrastructure for carriers. Backed by leading investors, the company fosters a culture of precision, purpose, and collaboration with flexible work options.