Lead the Incident Operations function, building and managing a team of Incident Commanders for critical incidents.
Establish severity models, escalation paths, and 24x7 follow-the-sun coverage across global regions.
Drive continuous improvement through retrospectives, KPIs, and AI-powered automation to reduce operational toil.
Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities, focusing on remote and flexible roles. The company values efficiency and objectivity in recruitment, fostering an inclusive environment that emphasizes curiosity, empathy, and accountability.
Lead, mentor, and grow a team of SRE/DevOps engineers while partnering with engineering leadership to assess team needs and develop talent.
Oversee the incident management process end to end, including on-call rotations, escalation paths, incident command, postmortems, and root cause analysis.
Define and drive SRE principles like SLIs, SLOs, error budgets, capacity planning, and observability standards, championing a culture of reliability and operational excellence.
Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services, enabling community financial institutions to digitally engage in a secure and compliant way through a world-class digital communications platform. Their platform integrates Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology, bolstered by AI and contact center capabilities, and they value integrity, transparency, and ownership.
Own prioritized constraints end-to-end, taking cross-functional operating problems from zero context through measurable outcome.
Work in any domain as needed, from HR automation to broker dealer constraints, applying systems thinking across unfamiliar areas.
Partner across the company to mobilize resources and drive adoption, creating non-linear leverage through workflow redesign, automation, or org changes.
Alpaca is a US-headquartered global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, and more. We are a dynamic team of 400+ globally distributed members who thrive working from favorite places around the world, backed by $400 million in funding from top-tier investors.
Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
Define and drive SRE platform strategy, incident management, and observability engineering.
Mentor team members, foster collaboration, and ensure operational excellence.
XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.
Lead and modernize Sectigo's global infrastructure organization with a focus on reliability and operational maturity.
Develop a measurable operating model using SLAs, SLOs, and key metrics to drive improvement.
Drive automation, AI-enabled operations, and closer collaboration with engineering teams.
Sectigo is an innovative provider of certificate lifecycle management (CLM) solutions, helping large brands simplify digital trust. With over 700,000 customers including 65% of the Fortune 500, they emphasize a culture of support, excellence, and teamwork.
Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.
Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.
Build and lead a team of incident responders, engineering automation and tooling to make response faster and more scalable.
Guide program maturity through AI-assisted workflows, operational excellence, and cross-functional partnerships.
Serve as incident manager during complex, high-severity events, balancing strategic priorities with team development.
1Password is a cybersecurity company that builds enterprise password management and unified access management solutions, trusted by over 180,000 businesses. With over $400M in ARR and a spot on the Forbes Cloud 100 for four years, the company fosters a human-centric, collaborative culture that values innovation and speed.
Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Lead the transformation of a diverse operations-heavy organization into a modern, AI-first Production Engineering function.
Own end-to-end reliability, performance, scalability, and security of NICE's global cloud, telecom, and datacenter platforms.
Drive adoption of software-first operational practices including automated recovery, infrastructure as code, and observability.
NICE provides software products used by 25,000+ global businesses to deliver extraordinary customer experiences, fight financial crime, and ensure public safety. With over 8,500 employees across 30+ countries, the company fosters a culture of ambition, game-changing innovation, and high standards.
Lead and coach a high-performing Security & IT organization, building team culture and succession depth.
Own detection and incident response end to end, from detection engineering to post-incident reviews.
Partner with Engineering, Product, and other teams to embed security into the development lifecycle and drive IT operations strategy.
Automox provides a cloud-native IT operations platform that replaces traditional tools with autonomous endpoint management. With over 2,500 customers including NASA and Yale, the company fosters a collaborative, one-team culture in a fully remote environment.
Own end-to-end Corporate Actions across U.S. and international markets, including mandatory and voluntary events.
Establish supervisory controls, incident management, and automation to scale operations and reduce manual processing.
Partner with cross-functional teams to eliminate failure modes and strengthen front-to-back controls.
Alpaca is a US-headquartered global leader in brokerage infrastructure for stocks, ETFs, options, crypto, and fixed income. With 400+ globally distributed members and $400 million in funding from top-tier investors, we foster a diverse, open-source community committed to opening financial services to everyone.
Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
Define SLOs and SLIs to drive architectural decisions and error budget policies.
Conduct blameless post-incident reviews and implement long-term preventive measures.
Airalo is the world's first eSIM store, helping travelers access affordable mobile data in 200+ countries. They are a fully remote team of 400+ people across 60+ countries, with a culture of trust, ownership, and freedom.
Enable effective execution with Quality and Speed, in partnership with the team's Product Manager.
Ensure 3+ 9s availability of Dedicated infrastructure, ensuring security and automating for maximum scalability.
Provide clear direction, meaningful feedback and foster an environment where meaningless toil gets ruthlessly automated.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With over 50 million users and 50% of Fortune 100, GitLab fosters a high-performance culture driven by values.
You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.
Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.
Lead security monitoring, incident response, and threat hunting across cloud and AI-enabled environments.
Establish operational priorities, metrics, and playbooks based on organizational risk.
Drive responsible adoption of AI-assisted detection and response capabilities.
Backblaze is a cloud storage and backup provider that helps customers protect their data across over 175 countries. The company fosters a culture centered on fairness, goodness, and work-life balance, with a strong commitment to diversity and inclusion.