Lead the team that commands Alpaca's most critical incidents, building the severity model, escalation paths, and 24x7 follow-the-sun coverage.
Drive severity maturity with Risk on financial and regulatory materiality, and own the escalation path to engineering leadership.
Own the KPIs for response and mitigation, build a blameless review culture, and automate processes with AI agents.
Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, and fixed income. With over 400 employees across the globe and backed by $400 million in funding, the team is diverse and includes experienced engineers, traders, and brokerage professionals.
Build and lead a team of incident responders, engineering automation and tooling to make response faster and more scalable.
Guide program maturity through AI-assisted workflows, operational excellence, and cross-functional partnerships.
Serve as incident manager during complex, high-severity events, balancing strategic priorities with team development.
1Password is a cybersecurity company that builds enterprise password management and unified access management solutions, trusted by over 180,000 businesses. With over $400M in ARR and a spot on the Forbes Cloud 100 for four years, the company fosters a human-centric, collaborative culture that values innovation and speed.
Lead, develop, and support a team of security incident responders, setting clear expectations and driving career development.
Define and drive the security incident response roadmap, including automation and AI-assisted tooling.
Oversee detection, triage, containment, remediation, and post-incident learning for complex high-severity security events.
The company provides password management services to individuals and businesses. It fosters a collaborative, inclusive culture focused on transparency, psychological safety, and continuous learning.
Lead security monitoring, incident response, and threat hunting across cloud and AI-enabled environments.
Establish operational priorities, metrics, and playbooks based on organizational risk.
Drive responsible adoption of AI-assisted detection and response capabilities.
Backblaze is a cloud storage and backup provider that helps customers protect their data across over 175 countries. The company fosters a culture centered on fairness, goodness, and work-life balance, with a strong commitment to diversity and inclusion.
Lead and modernize Sectigo's global infrastructure organization with a focus on reliability and operational maturity.
Develop a measurable operating model using SLAs, SLOs, and key metrics to drive improvement.
Drive automation, AI-enabled operations, and closer collaboration with engineering teams.
Sectigo is an innovative provider of certificate lifecycle management (CLM) solutions, helping large brands simplify digital trust. With over 700,000 customers including 65% of the Fortune 500, they emphasize a culture of support, excellence, and teamwork.
Manage team performance, career development, and project prioritization while driving a culture of automation.
Drive initiatives with partner teams to improve infrastructure reliability and act as crisis management.
Analyze existing processes to drive continuous improvement and efficiencies.
ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter with an intelligent cloud platform. We are building an AI-native culture where technology and talent are unstoppable together, serving over 8,100 customers.
Apply expertise in incident management and SRE to evaluate AI-generated documents, spreadsheets, and slide decks for technical accuracy and operational rigor.
Assess outputs against real-world reliability practices, identifying factual, technical, and reasoning errors.
Provide clear, structured written feedback and collaborate asynchronously with a research team to refine evaluation approaches.
This partner company focuses on AI evaluation and development, seeking experienced professionals to assess AI-generated work products. They offer flexible remote work and independent contractor engagements with weekly payments.
Lead, mentor, and grow a team of SRE/DevOps engineers while partnering with engineering leadership to assess team needs and develop talent.
Oversee the incident management process end to end, including on-call rotations, escalation paths, incident command, postmortems, and root cause analysis.
Define and drive SRE principles like SLIs, SLOs, error budgets, capacity planning, and observability standards, championing a culture of reliability and operational excellence.
Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services, enabling community financial institutions to digitally engage in a secure and compliant way through a world-class digital communications platform. Their platform integrates Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology, bolstered by AI and contact center capabilities, and they value integrity, transparency, and ownership.
Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Lead complex, high-severity security incident responses as incident commander.
Design and build AI-assisted automation for security operations.
Own readiness programs, threat hunting, and insider risk capabilities.
1Password is a cybersecurity company that provides enterprise password management and unified access management, trusted by over 180,000 businesses. They are a remote-first company with a fast-paced, collaborative culture, and have surpassed $400M in ARR.
Own end-to-end management of critical customer escalations for key accounts in EMEA, establishing action plans and communication cadences.
Apply strong judgment in risk management, driving post-mortem analyses and lessons learned to track preventive actions.
Identify recurring escalation themes to drive continuous improvements in product, process, and operational readiness.
Zscaler accelerates digital transformation with its SASE-based Zero Trust Exchange platform, protecting thousands of customers from cyberattacks and data loss. As the world's largest in-line cloud security platform, it operates across 160+ public exchanges globally, with a culture driven by deep customer obsession and a commitment to AI-native enterprise solutions.
Lead cyber incident command during high-pressure events, providing clear updates to senior leaders.
Mature incident management with playbooks, tabletop exercises, and post-incident reviews.
Oversee business continuity planning, impact assessments, and readiness across critical operations.
Pax8 is the global AI and cloud Marketplace for small and medium-sized businesses, connecting service providers with technology solutions. With over 47,000 IT partners and 800,000 SMBs, the company values curiosity, collaboration, and a culture that is personal and driven by passion.
Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
Define and drive SRE platform strategy, incident management, and observability engineering.
Mentor team members, foster collaboration, and ensure operational excellence.
XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.
Orchestrate enterprise-wide Agentic AI and AIOps initiatives, tracking milestones and coordinating deployment of multi-agent systems.
Lead operational delivery for observability, self-healing automation, and security remediation across multi-cloud environments.
Establish executive reporting frameworks communicating delivery health, operational improvements, and measurable business value.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through automated matching. We are committed to an open, respectful, and inclusive environment where employees are empowered to contribute and grow.
Lead and strengthen a resilient security operations capability, overseeing detection engineering, SOC operations, incident response, and exposure remediation.
Build and develop a high-performing Cyber Defense team while establishing strong detection and response practices using threat intelligence.
Coordinate critical incidents across IT, Platform, Product, GRC, Legal, and business teams, driving remediation of high-impact risks.
Jobgether uses an AI-powered matching process to connect candidates with hiring companies. It is a platform that facilitates job applications and shares top-fitting candidate shortlists with employers.
Act as the primary interface for customers during high-visibility incidents, translating technical hurdles into clear business communications.
Lead cross-functional escalation initiatives, managing project timelines, dependencies, and deliverables to high-quality standards.
Mentor your team on advanced strategic thinking, moving them from tactical troubleshooting to proactive risk assessment and executive-level reporting.
Twilio is shaping the future of communications by delivering innovative solutions to hundreds of thousands of businesses and empowering millions of developers worldwide. They are a remote-first company with a strong culture of connection, global inclusion, and diverse experiences.
Lead the transformation of a diverse operations-heavy organization into a modern, AI-first Production Engineering function.
Own end-to-end reliability, performance, scalability, and security of NICE's global cloud, telecom, and datacenter platforms.
Drive adoption of software-first operational practices including automated recovery, infrastructure as code, and observability.
NICE provides software products used by 25,000+ global businesses to deliver extraordinary customer experiences, fight financial crime, and ensure public safety. With over 8,500 employees across 30+ countries, the company fosters a culture of ambition, game-changing innovation, and high standards.
Act as a subject matter expert for service management tools and practices, creating content across various mediums.
Partner with product engineering teams to build demos and coach internal teams on communication.
Interface with open source communities and contribute to the product through feedback, documentation, or code.
Datadog is the leading observability and security platform for the AI era, providing unified visibility across the technology stack. Trusted by Fortune 500 companies and high-growth AI leaders, Datadog fosters a culture of innovation and collaboration.
Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
Define SLOs and SLIs to drive architectural decisions and error budget policies.
Conduct blameless post-incident reviews and implement long-term preventive measures.
Airalo is the world's first eSIM store, helping travelers access affordable mobile data in 200+ countries. They are a fully remote team of 400+ people across 60+ countries, with a culture of trust, ownership, and freedom.
Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.
Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.