Establish performance, throughput, latency, and capacity baselines for critical platform workflows.
Define and maintain SLOs, error budgets, dashboards, alerts, and reliability thresholds.
Lead load, stress, soak, spike, failure, and recovery testing in representative environments.
Tech Holding is a full-service consulting firm that delivers predictable outcomes and high-quality solutions to clients. The company was founded by experienced industry professionals who have held senior positions at startups to Fortune 50 firms, fostering a culture of deep expertise, integrity, transparency, and dependability.
Own the full platform stack for Veeam Data Cloud in Government and Sovereign Cloud environments, including incident response, reliability, and observability.
Design and implement high-availability, fault-tolerant infrastructure on Azure (including Azure Government) with SLIs, SLOs, and error budgets.
Drive reliability improvements through automation, chaos engineering, and cross-team collaboration, with a focus on compliance and security.
Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide.
Drive company-wide reliability strategy, standards, and best practices across LinkedIn Engineering.
Lead adoption of service criticality models to set reliability expectations based on business impact.
Partner with engineering teams to improve system design, reduce incident risk, and strengthen operational readiness.
LinkedIn is the world's largest professional network, built to create economic opportunity for every member of the global workforce. We foster a culture of trust, care, inclusion, and fun, investing in employee growth to transform the way the world works.
Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, and networking.
Design and implement solutions that improve availability, scalability, performance, and resilience of the platform.
Everbridge empowers enterprises and government organizations to anticipate, mitigate, respond to, and recover from critical events. The company focuses on building resilient systems and fostering a culture of ownership, continuous improvement, and operational excellence.
Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.
ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.
Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.
Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.
Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
Serve as the on-account technical presence, partnering with customer infrastructure and security teams.
We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.
Design, implement, and maintain scalable and reliable systems.
Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
Develop and maintain automation tools for deployment, monitoring, and system health checks.
LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.
Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
Drive SLOs, observability, alerting, and on-call processes across teams.
Build the platform engineering function from the ground up and influence cross-cutting architecture.
First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.
Define SLIs, SLOs, and reliability targets for the platform.
Improve observability, alerting, and production readiness across services.
Automate operational work and support cloud/Kubernetes infrastructure.
Lodgify is a fast-growing scale-up in vacation rental technology, backed by $30M in funding. Headquartered in Barcelona, the 380+ person team of 60+ nationalities is passionate about transforming short-term rentals.
Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.
They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.
Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
Define and drive SRE platform strategy, incident management, and observability engineering.
Mentor team members, foster collaboration, and ensure operational excellence.
XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.
Define and implement observability strategies, standards, and governance across applications and platforms.
Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms.
Establish and drive SRE best practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.
Valtech is the experience innovation company that helps brands unlock new value in an increasingly digital world by blending crafts, categories, and cultures. They have a workplace culture that fosters creativity, diversity, and autonomy, with a borderless global framework enabling seamless collaboration.
Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.
The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.
Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.
Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.
Provide expert architectural guidance and day-two operational expertise for GCP infrastructure supporting a national aviation safety platform.
Partner with customer operations teams and Google technical advisors to strengthen platform reliability, automation, and DevSecOps practices.
Take proactive ownership of complex issues, from Tier 3 troubleshooting to operational readiness, ensuring high availability and performance.
540 is a forward-thinking consulting firm that partners with government agencies to deliver innovative technology solutions for mission-critical systems, including aviation safety platforms. The company fosters a culture of ownership, collaboration, and technical excellence, with a small-team environment where engineers drive real impact.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Lead the SRE strategy and execution for a high-growth AI company.
Build and scale a high-performing SRE team while defining reliability standards.
Architect secure, scalable cloud infrastructure and implement observability practices.
This company develops advanced AI products and agentic technology. It operates in a high-growth, international environment with a focus on operational excellence and innovation.
Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
Drive automation and eliminate operational toil through self-service tooling and process improvements.
Lead incident response and mentor engineers to improve reliability practices.
Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.
Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.