You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.
Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.
Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.
ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.
Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, and networking.
Design and implement solutions that improve availability, scalability, performance, and resilience of the platform.
Everbridge empowers enterprises and government organizations to anticipate, mitigate, respond to, and recover from critical events. The company focuses on building resilient systems and fostering a culture of ownership, continuous improvement, and operational excellence.
Lead reliability initiatives across multiple Ads domains including ad serving, auctions, targeting, reporting, measurement, and billing.
Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale.
Participate in on-call rotations, lead complex incident investigations and coordinate cross-functional response efforts during major production events.
Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet's largest sources of information.
Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.
Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.
Provide technical leadership for reliability across a large-scale advertising technology ecosystem
Lead reliability initiatives across ad serving, auctions, targeting, reporting, and billing systems
Mentor engineers and influence technical decisions to improve system resilience and developer productivity
The company is a partner organization operating a large-scale advertising technology ecosystem. Its size and culture are not detailed, but the role emphasizes reliability and operational excellence in a high-traffic environment.
Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
Define SLOs and SLIs to drive architectural decisions and error budget policies.
Conduct blameless post-incident reviews and implement long-term preventive measures.
Airalo is the world's first eSIM store, helping travelers access affordable mobile data in 200+ countries. They are a fully remote team of 400+ people across 60+ countries, with a culture of trust, ownership, and freedom.
Lead architecture and development of scalable backend systems for guest and host experiences.
Apply advanced AI technologies to products and engineering workflows.
Mentor engineers and drive technical direction across major initiatives.
The company builds communication and connectivity platforms for millions of users globally. It is a high-growth technology environment that values inclusion, collaboration, and diverse perspectives.
Lead the Incident Operations function, building and managing a team of Incident Commanders for critical incidents.
Establish severity models, escalation paths, and 24x7 follow-the-sun coverage across global regions.
Drive continuous improvement through retrospectives, KPIs, and AI-powered automation to reduce operational toil.
Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities, focusing on remote and flexible roles. The company values efficiency and objectivity in recruitment, fostering an inclusive environment that emphasizes curiosity, empathy, and accountability.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.
Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.
Build and operate monitoring, tracing, alerting, and observability infrastructure for system reliability.
Drive platform security initiatives with preventative controls and resilient architecture.
Lead incident response and recovery, including root-cause analysis and preventative measures.
This role is with a partner company managing AI-powered products. They are a growing technology organization with a fully distributed US-based team and a collaborative culture focused on large-scale infrastructure and AI technology.
Build and maintain the company's internal platform, driving operational excellence.
Collaborate with engineering squads to ensure applications are safe and reliable.
Take ownership of software infrastructure projects and provide off-hours support.
Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.
Lead the design, implementation, and continuous improvement of highly reliable, scalable systems for government environments.
Establish site reliability engineering practices, standards, and architectural approaches across engineering teams.
Develop automation and observability solutions to reduce manual work and improve operational performance.
This company is a research and development organization focused on building and maintaining reliable infrastructure for government-focused technology environments. It operates with a collaborative, remote-first culture and values technical innovation, scalability, and operational excellence.
Apply expertise in incident management and SRE to evaluate AI-generated documents, spreadsheets, and slide decks for technical accuracy and operational rigor.
Assess outputs against real-world reliability practices, identifying factual, technical, and reasoning errors.
Provide clear, structured written feedback and collaborate asynchronously with a research team to refine evaluation approaches.
This partner company focuses on AI evaluation and development, seeking experienced professionals to assess AI-generated work products. They offer flexible remote work and independent contractor engagements with weekly payments.
Champion SRE culture and best practices to improve production reliability and system resilience.
Communicate with stakeholders at all stages and bring fresh ideas to the table.
Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.
Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.
Lead Cloud Platform and SRE teams to scale securely and efficiently.
Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
Champion SRE culture with SLOs, error budgets, and observability.
Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.
Acts as the strategic bridge between Cloud Operations, Product Management, Engineering, and other teams to drive service quality and operational excellence.
Drives large-scale transformation programs, promotes operational best practices, and ensures lessons learned translate into portfolio-wide improvements.
Champions automation, observability, and reliability standards while influencing engineering practices and product roadmaps.
Unit4 is a cloud company redefining ERP for mid-market people-centric organizations with over 40 years of heritage. They are a people-first community focused on trust, accountability, and growth, with a global team and a commitment to sustainability and inclusion.
Define and implement observability strategies, standards, and governance across applications and platforms.
Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms.
Establish and drive SRE best practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.
Valtech is the experience innovation company that helps brands unlock new value in an increasingly digital world by blending crafts, categories, and cultures. They have a workplace culture that fosters creativity, diversity, and autonomy, with a borderless global framework enabling seamless collaboration.