Source Job

Ireland

  • Own the reliability posture of production services, including availability, latency, capacity, and performance.
  • Define and operate against SLIs and SLOs, using error budgets to drive engineering priorities.
  • Lead incident response, write post-mortems, and build automation to measurably improve service reliability.

SRE Distributed Systems Cloud Infrastructure Incident Management

20 jobs similar to Staff Software Engineer (L4)

Jobs ranked by similarity.

$150,000–$185,000/yr
US Unlimited PTO

  • You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
  • You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
  • You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.

Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.

EMEA

  • Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
  • Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
  • Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.

ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.

Global 4w PTO

  • Champion SRE culture and best practices to improve production reliability and system resilience.
  • Communicate with stakeholders at all stages and bring fresh ideas to the table.
  • Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.

Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.

$114,800–$150,000/yr
US 4w PTO

  • Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
  • Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
  • Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.

Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

US

  • Apply expertise in incident management and SRE to evaluate AI-generated documents, spreadsheets, and slide decks for technical accuracy and operational rigor.
  • Assess outputs against real-world reliability practices, identifying factual, technical, and reasoning errors.
  • Provide clear, structured written feedback and collaborate asynchronously with a research team to refine evaluation approaches.

This partner company focuses on AI evaluation and development, seeking experienced professionals to assess AI-generated work products. They offer flexible remote work and independent contractor engagements with weekly payments.

$165,000–$165,000/yr
US Unlimited PTO

  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure.
  • Build and manage CI/CD pipelines, automate operational tasks, and improve deployment processes.
  • Monitor production systems, participate in incident response, and champion DevOps best practices.

First Due provides transformative end-to-end software solutions for fire and EMS agencies, helping them run safer, smarter, and more effective operations. The company offers a fully remote workplace, comprehensive benefits, and opportunities for advancement, with a culture focused on respect, inclusivity, and equal opportunity.

UK

  • Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
  • Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.

US

  • Drive long-term reliability strategy and SRE architecture across Airbnb's infrastructure.
  • Mentor and lead Site Reliability Engineers, fostering a culture of reliability and proactive problem-solving.
  • Collaborate with engineering teams to design reliable systems and automate operational processes.

Airbnb is an online marketplace for lodging and experiences, connecting hosts with guests worldwide. With over 5 million hosts and 2 billion guest arrivals, the company fosters a community of belonging and innovation. They prioritize a culture of inclusion, creativity, and long-term stakeholder success.

US

  • Lead the transformation of a diverse operations-heavy organization into a modern, AI-first Production Engineering function.
  • Own end-to-end reliability, performance, scalability, and security of NICE's global cloud, telecom, and datacenter platforms.
  • Drive adoption of software-first operational practices including automated recovery, infrastructure as code, and observability.

NICE provides software products used by 25,000+ global businesses to deliver extraordinary customer experiences, fight financial crime, and ensure public safety. With over 8,500 employees across 30+ countries, the company fosters a culture of ambition, game-changing innovation, and high standards.

US

  • Design, build, and operate shared platform foundations including GCP, Kubernetes, networking, CI/CD, and observability.
  • Diagnose and troubleshoot complex distributed systems running at high request volume.
  • Raise the reliability bar through dashboards, alerting, on-call readiness, and automation.

Sanity.io builds an AI-powered content operating system that helps teams model, create, and automate content. The company has 200+ employees and a positive, flexible, trust-based culture that supports growth and work-life balance.

$191,000–$226,000/yr
US Unlimited PTO

  • Own the reliability, performance, and resilience of cloud environments (AWS, Kubernetes) and define SLOs across critical services.
  • Lead incident response, on-call rotation, and drive root cause analysis to ensure high production quality.
  • Build and maintain observability systems and automate operational toil using AI tools.

Garner partners with employers to redesign healthcare by using clinical metrics to identify top doctors and incentivize members to better care. The company has helped over 2.5 million people, saved $1B in costs, and doubled annually for five years, fostering a mission-driven, high-performance culture.

Europe

  • Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
  • Establish and manage SLOs and SLAs for ClickHouse Cloud, ensuring monitoring and alerting are in place for all infrastructure.
  • Lead incident response, blameless postmortems, and chaos initiatives to continuously improve reliability and performance.

ClickHouse develops an open-source column-oriented database management system and offers a cloud database service. The company is a rapidly scaling, globally distributed startup with employees in over 25 countries, offering a flexible and collaborative culture.

US

  • Provide technical leadership for reliability across a large-scale advertising technology ecosystem
  • Lead reliability initiatives across ad serving, auctions, targeting, reporting, and billing systems
  • Mentor engineers and influence technical decisions to improve system resilience and developer productivity

The company is a partner organization operating a large-scale advertising technology ecosystem. Its size and culture are not detailed, but the role emphasizes reliability and operational excellence in a high-traffic environment.

Ireland

  • Design and deliver major components of Twilio's carrier test and observability platform.
  • Build high-throughput data and alerting systems that turn telemetry into actionable signals.
  • Extend and harden production LLM systems for automated carrier troubleshooting with safe guardrails.

Twilio is a customer engagement platform that delivers innovative communications solutions to hundreds of thousands of businesses and empowers millions of developers worldwide. The company is remote-first, values connection and global inclusion, and has a vibrant culture where diverse employees make a global impact.

Brazil Unlimited PTO

  • Build and maintain the company's internal platform, driving operational excellence.
  • Collaborate with engineering squads to ensure applications are safe and reliable.
  • Take ownership of software infrastructure projects and provide off-hours support.

Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.

$200,700–$250,900/yr
US Canada

  • Embed with product teams to improve operational maturity through on-call, monitoring, and alerting practices.
  • Run game day exercises and implement reliability techniques in Haskell & TypeScript code.
  • Champion reliability practices through design reviews and advocate for SLOs tied to customer outcomes.

Mercury is a fintech company that provides banking services for startups. They are building a modern banking platform and value reliability and innovation.

India

  • Build and maintain scalable cloud infrastructure for high availability.
  • Enhance observability and monitoring frameworks for accurate alerts.
  • Support on-call rotations and incident response with post-mortems.

GoGuardian is an award-winning learning solutions company purpose-built for K-12, trusted by educators to promote effective teaching and keep students safe. They are a remote, diverse, and committed team of mission-driven employees focused on improving learning environments.

Canada

  • Manage team performance, career development, and project prioritization while driving a culture of automation.
  • Drive initiatives with partner teams to improve infrastructure reliability and act as crisis management.
  • Analyze existing processes to drive continuous improvement and efficiencies.

ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter with an intelligent cloud platform. We are building an AI-native culture where technology and talent are unstoppable together, serving over 8,100 customers.

UK

  • Lead the Incident Operations function, building and managing a team of Incident Commanders for critical incidents.
  • Establish severity models, escalation paths, and 24x7 follow-the-sun coverage across global regions.
  • Drive continuous improvement through retrospectives, KPIs, and AI-powered automation to reduce operational toil.

Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities, focusing on remote and flexible roles. The company values efficiency and objectivity in recruitment, fostering an inclusive environment that emphasizes curiosity, empathy, and accountability.