Source Job

$200,700–$250,900/yr
US Canada

  • Embed with product teams to improve operational maturity through on-call, monitoring, and alerting practices.
  • Run game day exercises and implement reliability techniques in Haskell & TypeScript code.
  • Champion reliability practices through design reviews and advocate for SLOs tied to customer outcomes.

SRE PostgreSQL Temporal Grafana OpenTelemetry

20 jobs similar to Senior Software Engineer - SRE

Jobs ranked by similarity.

$114,800–$150,000/yr
US 4w PTO

  • Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
  • Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
  • Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.

Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.

Europe 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and define per-tenant SLOs and reliability models.
  • Serve as a primary escalation point for incidents, lead response and post-incident reviews, and improve alert quality.

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations to ensure reliability and resolve incidents faster. We are a 100% remote company with team members across 40+ countries, backed by leading investors, and we foster a global collaborative culture and a passion for meaningful work.

Europe US LATAM 4w PTO

  • Own observability for critical product journeys, defining SLIs/SLOs and building metrics, dashboards, and alerts.
  • Act as first responder for production incidents, investigating signals and mitigating issues independently.
  • Work within a cross-functional squad of 6-8 engineers to improve reliability, monitoring, and incident response processes.

Feeld is a dating app creating a safer and more inclusive space for exploring relationships and sexuality. They have a distributed engineering team of around 50 people across Europe and the US, working in small autonomous squads.

$150,000–$185,000/yr
US Unlimited PTO

  • You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
  • You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
  • You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.

Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.

$74,000–$111,000/yr
Canada Unlimited PTO

  • Define and implement observability strategies, standards, and governance across applications and platforms.
  • Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms.
  • Establish and drive SRE best practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.

Valtech is the experience innovation company that helps brands unlock new value in an increasingly digital world by blending crafts, categories, and cultures. They have a workplace culture that fosters creativity, diversity, and autonomy, with a borderless global framework enabling seamless collaboration.

$240,000–$240,000/yr
US

  • Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
  • Drive SLOs, observability, alerting, and on-call processes across teams.
  • Build the platform engineering function from the ground up and influence cross-cutting architecture.

First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.

Spain 5w PTO

  • Define SLIs, SLOs, and reliability targets for the platform.
  • Improve observability, alerting, and production readiness across services.
  • Automate operational work and support cloud/Kubernetes infrastructure.

Lodgify is a fast-growing scale-up in vacation rental technology, backed by $30M in funding. Headquartered in Barcelona, the 380+ person team of 60+ nationalities is passionate about transforming short-term rentals.

$4,538–$5,772/mo
Poland

  • Define and drive reliability of systems at the scale of millions of clients, strengthening SRE practices. - Develop observability platforms and serve as a strategic partner to product engineering teams. - Enhance proactive resilience through early-warning systems, AI/ML, and incident management.

XTB is a global FinTech company specializing in online trading of financial instruments. As the largest FinTech in Poland and a leader in Central and Eastern Europe, we operate across multiple continents and are a certified Great Place to Work, focusing on employee development and training.

EMEA

  • Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
  • Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
  • Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.

ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.

US

  • Improve system availability, scalability, and resilience across Flowcode's platforms.
  • Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
  • Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.

Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.

US

  • You'll contribute to infrastructure scaling to infinitely many apps, improving performance and reliability across backend services.
  • You'll support observability efforts, help implement SLOs, and build foundational services for next-generation cloud infrastructure.
  • You'll participate in triage and on-call processes to diagnose issues and implement changes to prevent recurrence.

Bubble is an AI visual development platform that empowers anyone to create software without code, from first-time entrepreneurs to enterprise teams. With over 6 million users in more than 100 countries and a mission to break down barriers to entrepreneurship, the company fosters a collaborative and inclusive culture focused on empowering builders worldwide.

$119,380–$165,100/yr
Spain UK

  • Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
  • Define SLOs and SLIs to drive architectural decisions and error budget policies.
  • Conduct blameless post-incident reviews and implement long-term preventive measures.

Airalo is the world's first eSIM store, helping travelers access affordable mobile data in 200+ countries. They are a fully remote team of 400+ people across 60+ countries, with a culture of trust, ownership, and freedom.

US UK Ireland Poland Germany Australia

  • Design and implement a scalable observability platform for Whatnot's growing infrastructure.
  • Work with core infrastructure, platform, and developer tools teams to redesign data collection to visualization.
  • Utilize AI agents and open standards to ensure visibility into software stack performance and reliability.

Whatnot is the largest live shopping platform in North America and Europe, enabling sellers to build businesses across hundreds of categories. They are a remote co-located team anchored in hubs across the US, UK, Ireland, Poland, Germany, and Australia, and were recently named the #1 Best Startup Employer in America by Forbes.

Argentina Colombia Costa Rica Mexico

  • Operate and maintain high-availability database systems including Vitess (distributed MySQL) and Cassandra against established runbooks.
  • Develop automation for common operational tasks using tools like Terraform, Ansible, and scripting languages such as Python, Bash, or Go.
  • Participate in on-call rotations, incident response, and post-incident reviews to ensure service reliability and continuous improvement.

Backblaze is the object storage leader in the open cloud movement, helping customers break free from overpriced legacy solutions with cloud storage designed to unlock budgets and unleash innovators. Founded in 2007, the company generates over $136M ARR, manages over three billion gigabytes of data for 500K+ customers in 175+ countries, and values diversity, equity, and inclusion in its workforce.

India

  • Lead, mentor, and grow a team of SRE/DevOps engineers while partnering with engineering leadership to assess team needs and develop talent.
  • Oversee the incident management process end to end, including on-call rotations, escalation paths, incident command, postmortems, and root cause analysis.
  • Define and drive SRE principles like SLIs, SLOs, error budgets, capacity planning, and observability standards, championing a culture of reliability and operational excellence.

Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services, enabling community financial institutions to digitally engage in a secure and compliant way through a world-class digital communications platform. Their platform integrates Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology, bolstered by AI and contact center capabilities, and they value integrity, transparency, and ownership.

$114,700–$195,000/yr
North America

  • Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
  • Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
  • Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.

Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.

Brazil

  • Ensure reliability, scalability, and performance of cloud-based systems using Kubernetes and observability tools.
  • Define and monitor reliability metrics (SLIs, SLOs, MTTR) to continuously improve operational performance.
  • Automate operational tasks and implement Infrastructure as Code to reduce manual work and enhance efficiency.

Our partner is a technology company focused on building and maintaining reliable, scalable digital environments. They promote a culture of continuous improvement, collaboration, and proactive engineering.

Poland

  • Design and develop highly performant backend services for real-time data processing and web APIs.
  • Define and own reliability objectives and error budgets for core API services.
  • Build observability through metrics, logging, tracing, and dashboards, and participate in on-call rotation.

The company provides technology that protects businesses and users from online fraud. It is a fully remote, globally distributed organization.

Poland

  • Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
  • Define and drive SRE platform strategy, incident management, and observability engineering.
  • Mentor team members, foster collaboration, and ensure operational excellence.

XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.

$64,021–$92,683/yr
Canada

  • Extend the self-service datastore platform with provisioning automation, guardrails, and paved paths for product engineering teams.
  • Ship observability, alerting, and backup/disaster recovery as built-in defaults for every datastore.
  • Convert recurring pull-in work into platform features or AI tooling that other teams can use directly.

Greenhouse provides a hiring software platform designed to make hiring work for everyone. They have an award-winning culture recognized by Fortune and Inc., and foster inclusivity, transparency, and accountability among their teams.