Improve system availability, scalability, and resilience across Flowcode's platforms.
Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.
Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.
Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.
Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.
Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.
Be on an on-call rotation responding to production incidents and support service engineers.
Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes, making monitoring alert on symptoms.
Design and maintain core infrastructure scaling to hundreds of thousands of concurrent users.
Our client's Cloud Operations team is expanding its SRE function, keeping user-facing services and production systems running smoothly. The team specializes in systems like networking, Linux kernel, and distributed systems, blending pragmatic operations with software engineering.
Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
Drive automation and eliminate operational toil through self-service tooling and process improvements.
Lead incident response and mentor engineers to improve reliability practices.
Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.
Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
Serve as the on-account technical presence, partnering with customer infrastructure and security teams.
We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.
Drive deployment automation across EKS clusters using GitOps.
Own infrastructure requirements and coordinate maintenance backlog.
Collaborate with engineering teams and integrate AI agents to accelerate workflows.
Ping Identity provides an intelligent cloud identity platform that secures digital experiences. With global offices and serving over half of the Fortune 100, the company fosters a culture of respect and individuality.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Design and advance core infrastructure for multi-cloud Kubernetes clusters and developer toolchains.
Automate operations and engineering tasks to improve productivity and reliability.
Build machine learning infrastructure to enable AI teams to train and deploy large-scale models.
Cresta provides an AI platform that transforms customer conversations into competitive advantages by combining conversational AI, real-time agent augmentation, and conversation intelligence. The company has raised over $270 million from top investors like a16z, Greylock, and Sequoia, and is led by AI industry veterans.
Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.
ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.
Define SLIs, SLOs, and reliability targets for the platform.
Improve observability, alerting, and production readiness across services.
Automate operational work and support cloud/Kubernetes infrastructure.
Lodgify is a fast-growing scale-up in vacation rental technology, backed by $30M in funding. Headquartered in Barcelona, the 380+ person team of 60+ nationalities is passionate about transforming short-term rentals.
Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.
Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.
Design, build, and deploy production systems with focus on scalability, reliability, and security.
Develop and maintain automation to streamline operations and eliminate toil.
Proactively monitor systems and implement automated incident response to minimize downtime.
Arista Networks is an industry leader in data-driven networking for large data centers, campus, and routing. With over $8 billion in revenue and a culture valuing diversity, Arista is a Great Place to Work for Best Engineering Team and Best Company for Diversity.
Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.
Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Design and maintain AWS cloud infrastructure using OpenTofu and Terraform.
Operate Kubernetes workloads on Amazon EKS, managing GitOps deployments with Argo CD and Helm.
Implement observability with Datadog, troubleshoot production incidents, and support on-call rotation.
PAR Technology Corporation provides innovative restaurant technology solutions, including point-of-sale, digital ordering, loyalty, and back-office software, as well as hardware and drive-thru offerings. With over 40 years of experience, the company serves more than 100,000 restaurants globally and fosters a collaborative culture centered on its 'Better Together' ethos.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
Own the technical direction and architecture of critical infrastructure domains, establishing scalable patterns and standards.
Lead complex, multi-team infrastructure initiatives from design through implementation and production operation.
Design and evolve AWS and Kubernetes infrastructure to enable teams to build and deploy systems reliably at scale.
We provide innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. Our company is backed by world-class investors including Craft Ventures and Andreessen Horowitz, with offices across the US and India, and we are growing extremely quickly.
Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
Design and implement automation to scale reliability practices and define per-tenant SLOs and reliability models.
Serve as a primary escalation point for incidents, lead response and post-incident reviews, and improve alert quality.
Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations to ensure reliability and resolve incidents faster. We are a 100% remote company with team members across 40+ countries, backed by leading investors, and we foster a global collaborative culture and a passion for meaningful work.