Source Job

Canada

  • Manage team performance, career development, and project prioritization while driving a culture of automation.
  • Drive initiatives with partner teams to improve infrastructure reliability and act as crisis management.
  • Analyze existing processes to drive continuous improvement and efficiencies.

Linux Databases Coding Cloud Operations ITIL

20 jobs similar to Manager, Network Reliability and Resiliency

Jobs ranked by similarity.

Canada

  • Manage team, career development, project prioritization, and performance review.
  • Drive a culture of automation and reduce manual activities.
  • Drive initiatives with partner teams to improve infrastructure reliability.

ServiceNow provides an AI platform for business reinvention, used by 85% of the Fortune 500. The company fosters an AI-native culture with a focus on innovation and collaboration.

EMEA

  • Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
  • Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
  • Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.

ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.

$64,021–$92,683/yr
Canada

  • Extend the self-service datastore platform with provisioning automation, guardrails, and paved paths for product engineering teams.
  • Ship observability, alerting, and backup/disaster recovery as built-in defaults for every datastore.
  • Convert recurring pull-in work into platform features or AI tooling that other teams can use directly.

Greenhouse provides a hiring software platform designed to make hiring work for everyone. They have an award-winning culture recognized by Fortune and Inc., and foster inclusivity, transparency, and accountability among their teams.

Canada

  • Lead cross-functional development of large-scale infrastructure programs from concept to delivery.
  • Align stakeholders across engineering, product, and finance to manage risks and drive program success.
  • Optimize engineering processes and drive cost efficiency through automated monitoring and reduction initiatives.

Lime is a global shared micromobility company on a mission to make transportation shared, affordable, and carbon-free. It has powered over one billion rides in nearly 30 countries and is a Time Magazine 100 Most Influential Company, fostering a culture of strategic thinking and operational excellence.

$129,283–$161,622/yr
Canada 20w maternity 20w paternity

  • Lead the Network Operations team through infrastructure modernization and cloud-native transition.
  • Drive automation using Terraform and infrastructure-as-code to reduce manual operations.
  • Maintain 99.999% uptime and manage corporate network infrastructure including Cisco Meraki.

Marqeta is a card issuing platform that enables companies to issue cards and manage payment operations in real time. They are a publicly-traded company powering well-known brands in the new economy, with a culture focused on customer success, innovation, and teamwork.

UK

  • Enable effective execution with Quality and Speed, in partnership with the team's Product Manager.
  • Ensure 3+ 9s availability of Dedicated infrastructure, ensuring security and automating for maximum scalability.
  • Provide clear direction, meaningful feedback and foster an environment where meaningless toil gets ruthlessly automated.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With over 50 million users and 50% of Fortune 100, GitLab fosters a high-performance culture driven by values.

$114,700–$195,000/yr
North America

  • Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
  • Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
  • Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.

Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.

Canada

  • Deliver customer excellence and meet all SLAs.
  • Deploy, upgrade, and support applications, services, and operating systems.
  • Troubleshoot system performance and application health issues.

Kinaxis is a global leader in modern supply chain orchestration, powering complex global supply chains. With over 2000 employees and multiple Top Employer awards, they foster a culture of innovation and collaboration.

India

  • Lead, mentor, and grow a team of SRE/DevOps engineers while partnering with engineering leadership to assess team needs and develop talent.
  • Oversee the incident management process end to end, including on-call rotations, escalation paths, incident command, postmortems, and root cause analysis.
  • Define and drive SRE principles like SLIs, SLOs, error budgets, capacity planning, and observability standards, championing a culture of reliability and operational excellence.

Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services, enabling community financial institutions to digitally engage in a secure and compliant way through a world-class digital communications platform. Their platform integrates Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology, bolstered by AI and contact center capabilities, and they value integrity, transparency, and ownership.

Canada

  • Define and drive the vision and multi-quarter roadmap for the infrastructure foundations team, tying investments to business outcomes.
  • Lead and mentor a team of infrastructure engineers, fostering ownership, collaboration, and technical excellence.
  • Own the reliability and safety of the foundational AWS layer, including account provisioning, core networking, and IAM access.

Affirm is reinventing credit to make it more honest and friendly, offering consumers the ability to buy now and pay later without hidden fees or compounding interest. As a publicly traded company, Affirm fosters a culture of thorough technical design review, operational excellence, and capable incident response.

$145,000–$177,000/yr
US

  • Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
  • Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, and networking.
  • Design and implement solutions that improve availability, scalability, performance, and resilience of the platform.

Everbridge empowers enterprises and government organizations to anticipate, mitigate, respond to, and recover from critical events. The company focuses on building resilient systems and fostering a culture of ownership, continuous improvement, and operational excellence.

$151,000–$242,000/yr
US

  • Ensure reliability, scalability, and operational excellence of analytics and data systems.
  • Provide technical leadership and direction to an offshore contract team.
  • Drive incident response, automation, and data governance initiatives.

Workiva provides an AI-powered platform that unifies finance, risk, and sustainability for complex organizations. It is a large enterprise with a collaborative and innovative culture centered on data integrity and trust.

EMEA

  • Take an active role as co-owner of production services to ensure they are built, maintained, and operated in a reliable and scalable way.
  • Collaborate with Software Engineering to drive operational improvements through metric-driven analysis and help scale AWS and Kubernetes infrastructure.
  • Participate in a weekly on-call rotation to investigate and resolve potential system issues, and automate routine tasks in at least two programming languages.

Zerohash is the leading crypto and stablecoin infrastructure platform, powering the next generation of financial services for banks, brokerages, fintechs, and payment companies. Founded in 2017, the company has raised over $280 million from top venture firms and strategic investors, and is trusted by global brands like Morgan Stanley and Stripe, operating with a compliance-first approach.

Global 6w PTO

  • Operate as the NOC’s first point of escalation for any issues raised within or out of shift that need further assistance or feedback.
  • Ensure NOC shift operations are executed properly and according to pre-defined SLAs.
  • Manage NOC’s shift schedule including PTO requests and act as focal point for administrative issues.

LivePerson is a global leader in enterprise conversations, providing a Conversational Cloud platform for brands like HSBC and Chipotle. The company powers nearly a billion interactions monthly and fosters an inclusive workplace culture that encourages collaboration and innovation.

Global Unlimited PTO

  • Operate the Monad node fleet, including health, sync, upgrades, and incident response for validators, full nodes, and archive nodes.
  • Own infrastructure-as-code with Ansible, Terraform, and Kubernetes, and build observability with Prometheus, Grafana, and Loki.
  • Design and build AI agent tooling for automated operations, including runbooks-as-code and deterministic guardrails.

Category Labs designs and builds decentralized technology, including the Monad blockchain, a high-performance EVM-compatible Layer 1. The team raised $225M in series A funding and is a lean, collaborative group of engineers and researchers with a culture of low ego and high-quality output.

US 4w PTO

  • Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
  • Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
  • Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.

Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.

$175,000–$195,000/yr

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
  • Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
  • Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.

Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.

Poland

  • Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
  • Define and drive SRE platform strategy, incident management, and observability engineering.
  • Mentor team members, foster collaboration, and ensure operational excellence.

XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.

Canada Unlimited PTO

  • Design and operate secure, multi-tenant infrastructure for agentic AI workloads at the Linux and network boundary.
  • Build process isolation and sandboxing mechanisms using namespaces, cgroups, gVisor, Firecracker, and workload identity frameworks.
  • Collaborate with Platform, Security, and Backend teams to own infrastructure from architecture through production operations.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It has a remote-first culture and uses technology to review applications fairly.

$221,200–$387,100/yr
North America Unlimited PTO

  • Drive large-scale, cross-functional strategic programs for the APEX organization, operating at the intersection of strategy and execution.
  • Translate ambiguous business objectives into prioritized milestones and deliverables, ensuring accountability and measurable outcomes.
  • Champion AI-native ways of working, using AI tools to accelerate program delivery and synthesize intelligence across workstreams.

ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter with its AI platform. The company fosters an AI-native culture where technology and talent are unstoppable together.