Define and implement reliability strategy including SLOs, SLIs, error budgets, and incident practices.
Manage cloud infrastructure on AWS using Infrastructure as Code and ensure Kubernetes scalability.
Lead incident response and establish chaos engineering practices to strengthen platform resilience.
This partner company builds a globally scaled, AI-native platform with a focus on reliability and event-driven systems. They offer a collaborative international culture with significant technical ownership and continuous improvement.
Help define and mature Engineering Operations by improving application health visibility, service reliability, and operational analytics.
Build and implement scalable processes for Incident, Problem, and Change Management that engineers actually want to use.
Connect engineering systems, data, and teams to reduce fragmentation and improve operational visibility across the organization.
Turnitin is a recognized innovator in global education, developing learning integrity solutions that help educators and institutions uphold academic integrity. With over 16,000 academic institutions using our services in more than 185 countries, we foster a remote-first culture and a diverse community of colleagues across 35+ countries.
Work with teams to define SLIs and SLOs, and create systems for observability.
Analyze failure scenarios, create runbooks, and reduce work that does not add value.
Participate in incident management and facilitate on-call duty to ensure reliable production environments.
Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.
Establish an SRE function, building a team of 4-6 and professionalizing incident management with clear processes and tooling.
Drive engineering excellence through design reviews, code reviews, and blameless retrospectives to foster a quality culture.
Lead people and projects, balancing incident response with a roadmap of observability and reliability engineering initiatives.
We are a fast-growing, remote cybersecurity company dedicated to enabling organizations to proactively find and fix exploitable attack vectors. Our team is a fusion of former special operations cyber operators and engineers, committed to a culture of respect, collaboration, and ownership.