Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.
Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.
Improve system availability, scalability, and resilience across Flowcode's platforms.
Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.
Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.
Architect the end-to-end reliability, performance, and resilience of cloud environments, including the SLO framework for critical services.
Lead incident response, on-call rotation, root cause analysis, and build a culture of corrective actions.
Build observability platforms to detect issues proactively and mentor engineers on reliability standards.
Garner is on a mission to transform the U.S. healthcare system by partnering with employers to steer members to better-performing doctors, resulting in better care and lower costs. With 550+ proprietary clinical metrics, they have helped over 2.5 million people and saved $1B in healthcare costs, recently raising a Series E and doubling five years running.
Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
Drive automation and eliminate operational toil through self-service tooling and process improvements.
Lead incident response and mentor engineers to improve reliability practices.
Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.
Be on an on-call rotation responding to production incidents and support service engineers.
Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes, making monitoring alert on symptoms.
Design and maintain core infrastructure scaling to hundreds of thousands of concurrent users.
Our client's Cloud Operations team is expanding its SRE function, keeping user-facing services and production systems running smoothly. The team specializes in systems like networking, Linux kernel, and distributed systems, blending pragmatic operations with software engineering.
Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.
Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.
You will ensure the reliability and high availability of Tenable's cloud products in cloud environments.
You will respond to support escalations and troubleshoot complex technical problems.
You will develop software, tools, and scripts to automate deployment and monitoring of production systems.
We are the Exposure Management company, trusted by over 40,000 organizations to understand and reduce cyber risk. Our global team supports 65% of the Fortune 500 and 50% of the Global 2000, with a culture of belonging, respect, and excellence.
Apply SRE principles to improve reliability, scalability, and performance of production systems.
Design and implement automation to reduce operational toil and improve engineering efficiency.
Lead incident response and develop sustainable solutions for complex production issues.
The hiring company is a technology organization focused on reliability and operational excellence. They offer a fully remote, collaborative environment with opportunities for technical leadership and career growth.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.
Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.
Lead, mentor, and grow a team of SRE/DevOps engineers while partnering with engineering leadership to assess team needs and develop talent.
Oversee the incident management process end to end, including on-call rotations, escalation paths, incident command, postmortems, and root cause analysis.
Define and drive SRE principles like SLIs, SLOs, error budgets, capacity planning, and observability standards, championing a culture of reliability and operational excellence.
Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services, enabling community financial institutions to digitally engage in a secure and compliant way through a world-class digital communications platform. Their platform integrates Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology, bolstered by AI and contact center capabilities, and they value integrity, transparency, and ownership.
Design, build, and operate shared cloud infrastructure using AWS, Kubernetes, Terraform, Databricks, and Cloudflare.
Deliver SRE and DevOps initiatives to improve reliability, scalability, observability, and deployment safety.
Build reusable infrastructure modules, automation, and self-service workflows to reduce manual work and improve developer experience.
YipitData is the leading market research and analytics firm for the disruptive economy, recently raising up to $475M from The Carlyle Group at a valuation over $1B. We analyze billions of alternative data points daily and have been recognized as one of Inc’s Best Workplaces, cultivating a people-centric culture focused on mastery, ownership, and transparency.
Design and operate scalable cloud infrastructure across AWS and GCP.
Build and improve Kubernetes, Linux, and cloud networking environments.
Strengthen security, disaster recovery, and platform resilience.
Hubstaff provides workforce analytics and time tracking for remote teams, serving over 200,000 global users. The company is a product-led organization with a winning culture and a fully remote team of experienced engineers.
Own the technical direction and architecture of critical infrastructure domains, establishing scalable patterns and standards.
Lead complex, multi-team infrastructure initiatives from design through implementation and production operation.
Design and evolve AWS and Kubernetes infrastructure to enable teams to build and deploy systems reliably at scale.
We provide innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. Our company is backed by world-class investors including Craft Ventures and Andreessen Horowitz, with offices across the US and India, and we are growing extremely quickly.
Define and implement observability strategies, standards, and governance across applications and platforms.
Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms.
Establish and drive SRE best practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.
Valtech is the experience innovation company that helps brands unlock new value in an increasingly digital world by blending crafts, categories, and cultures. They have a workplace culture that fosters creativity, diversity, and autonomy, with a borderless global framework enabling seamless collaboration.
Own the full platform stack for Veeam Data Cloud in Government and Sovereign Cloud environments, including incident response, reliability, and observability.
Design and implement high-availability, fault-tolerant infrastructure on Azure (including Azure Government) with SLIs, SLOs, and error budgets.
Drive reliability improvements through automation, chaos engineering, and cross-team collaboration, with a focus on compliance and security.
Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide.