Drive company-wide reliability strategy, standards, and best practices across LinkedIn Engineering.
Lead adoption of service criticality models to set reliability expectations based on business impact.
Partner with engineering teams to improve system design, reduce incident risk, and strengthen operational readiness.
LinkedIn is the world's largest professional network, built to create economic opportunity for every member of the global workforce. We foster a culture of trust, care, inclusion, and fun, investing in employee growth to transform the way the world works.
Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.
ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.
Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
Drive SLOs, observability, alerting, and on-call processes across teams.
Build the platform engineering function from the ground up and influence cross-cutting architecture.
First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.
Establish performance, throughput, latency, and capacity baselines for critical platform workflows.
Define and maintain SLOs, error budgets, dashboards, alerts, and reliability thresholds.
Lead load, stress, soak, spike, failure, and recovery testing in representative environments.
Tech Holding is a full-service consulting firm that delivers predictable outcomes and high-quality solutions to clients. The company was founded by experienced industry professionals who have held senior positions at startups to Fortune 50 firms, fostering a culture of deep expertise, integrity, transparency, and dependability.
Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.
Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.
You own delivery across product, engineering, and QA, leading a team of roughly ten people.
You stay close to the work by participating in design reviews, reading PRs, and occasionally writing code.
You run the engineering system by unblocking dependencies, killing ambiguity early, and keeping the path from decision to production short and reliable.
Ocra is building the commercial developer platform for parking, acting as the connective layer between sellers, booking channels, and payment systems. They are a small, distributed team that values autonomy, curiosity, and adaptability over rigid processes.
Enable effective execution with Quality and Speed, in partnership with the team's Product Manager.
Ensure 3+ 9s availability of Dedicated infrastructure, ensuring security and automating for maximum scalability.
Provide clear direction, meaningful feedback and foster an environment where meaningless toil gets ruthlessly automated.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With over 50 million users and 50% of Fortune 100, GitLab fosters a high-performance culture driven by values.
Ensure reliability, scalability, and operational excellence of analytics and data systems.
Provide technical leadership and direction to an offshore contract team.
Drive incident response, automation, and data governance initiatives.
Workiva provides an AI-powered platform that unifies finance, risk, and sustainability for complex organizations. It is a large enterprise with a collaborative and innovative culture centered on data integrity and trust.
Define and drive the technical vision for Implementation Platform Configuration and Enrollment Platform Core, ensuring scalability and reliability.
Lead and mentor a group of Engineering Managers, owning delivery commitments and incident posture across teams.
Drive AI adoption and testing strategies to scale platform reliability and reduce manual effort.
Bestow is a vertical technology platform that modernizes life insurance infrastructure for carriers. Backed by leading investors, the company fosters a culture of precision, purpose, and collaboration with flexible work options.
Deliver the Cells and Organizations roadmap including cell provisioning, routing, data migration, and feature parity.
Lead the engineering organization by hiring, developing managers and senior ICs, and setting the technical bar.
Drive cross-functional leadership across product groups, infrastructure, and security, mainly asynchronously.
GitLab is an intelligent orchestration platform for DevSecOps that enables organizations to increase developer productivity and improve operational efficiency. With over 50 million registered users and a high-performance culture driven by values and continuous knowledge exchange, GitLab embraces AI as a core productivity multiplier.
Lead the transformation of a diverse operations-heavy organization into a modern, AI-first Production Engineering function.
Own end-to-end reliability, performance, scalability, and security of NICE's global cloud, telecom, and datacenter platforms.
Drive adoption of software-first operational practices including automated recovery, infrastructure as code, and observability.
NICE provides software products used by 25,000+ global businesses to deliver extraordinary customer experiences, fight financial crime, and ensure public safety. With over 8,500 employees across 30+ countries, the company fosters a culture of ambition, game-changing innovation, and high standards.
Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
Define and drive SRE platform strategy, incident management, and observability engineering.
Mentor team members, foster collaboration, and ensure operational excellence.
XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.
Lead AI-assisted development as an ongoing experiment, adopting shared standards and building the team's discipline around clear context and human judgment.
Build, grow, and retain a strong engineering team while managing complex cross-team programs and holding tight to launch delivery deadlines.
Own reliability, cost, and change leadership, ensuring operational excellence and smooth transitions during reorganizations or acquisitions.
Sphera provides enterprise software and services that help companies manage and optimize their environmental, health, safety, and sustainability. They are a rapidly expanding team backed by Blackstone, guided by values of customer centricity, accountability, and collaboration.