Design, build, and operate AWS infrastructure across multiple regions using Kubernetes, Terraform, and Helm.
Own large-scale object storage environments exceeding 15 petabytes, optimizing reliability, performance, scalability, and cost.
Build an internal developer platform for self-service infrastructure, strengthen security practices, and lead incident response and postmortems.
Our partner builds an AI-powered sports media platform handling petabyte-scale media storage and millions of minutes of video monthly. It operates as a remote-first, international, engineering-led team with significant autonomy and a focus on high availability and security.
Own infrastructure as code across development, staging, and production environments
Build, maintain, and improve CI/CD pipelines for reliable and efficient deployments
Manage cloud infrastructure, establish scalable engineering practices, and lead incident response
CelebriOS is a software company building B2B SaaS products that help businesses make better decisions and streamline operations. The company has a remote-first working environment and a benefits package designed to support their team.
Take full ownership of the company's platform and infrastructure function, defining and executing the platform engineering strategy in partnership with the CTO.
Lead, mentor, and develop an established team of three DevOps Engineers and one SRE, setting clear responsibilities and development plans.
Own platform-related budgeting, cloud spend optimization, and technology decisions to improve reliability, security, and developer productivity.
Our client is a remote-first digital product company building and scaling SaaS products, AI-powered solutions, and web and mobile applications for global markets. The company operates across more than 20 countries with approximately 40–50 engineers supporting a portfolio of around 15 digital products.
Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.
A fast-growing AI/ML platform startup building infrastructure for training, evaluating, and aligning AI models within reinforcement learning environments. The engineering team of ~15 includes competitive programming medalists, serial AI startup founders, and researchers published at top venues.
Lead and grow a team of platform engineers, coaching them on infrastructure and cloud challenges.
Drive the platform roadmap, balancing reliability, cost, security, and developer experience with AWS and Kubernetes.
Partner cross-functionally to align platform priorities with business goals and ensure system reliability.
PerfectServe is a leading provider of clinical communication and physician scheduling solutions in the health IT space. The company has 400+ employees and 30,000+ customers, with over $100 million in annual revenue, and has received multiple Best in KLAS awards.
Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.
ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.
Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.
Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.
Building and coaching a high-performing distributed team with a shared operating model.
Owning platform capabilities for provisioning, deployment, and operations of infrastructure.
Leading infrastructure migration towards a modern SaaS model with incremental delivery.
Totara is a global learning platform trusted by more than 1,500 organisations and 21 million users worldwide, offering flexible learning, compliance, and talent development solutions. With a distributed team across New Zealand, Australia, the UK, and the US, the company values diverse perspectives and offers flexible, hybrid working.
Own and operate a production agentic AI platform on AWS, ensuring reliability and scaling.
Lead infrastructure automation and release management, driving best practices in security and compliance.
Collaborate with the platform team on agile ceremonies and proactively communicate status to stakeholders.
Inizio Evoke is a healthcare communications company dedicated to making health more human. As part of the larger Inizio network, it emphasizes a collaborative, inclusive culture where employees are encouraged to be their authentic selves.
Design and maintain AWS infrastructure using Terraform, with a focus on scalability cost and PCI-scoped network segmentation
Build and evolve the observability stack and CI/CD pipelines to ensure smooth production operations and rapid deployment
Lead incident response define SLOs and run performance tests to optimize payment-critical services
Xplor Technologies provides vertical software, embedded payments, and AI tools for membership-based and service-based industries. With over 130,000 businesses in 72+ countries and processing $47 billion in payments annually, the company values diversity, collaboration, and a people-first culture.
Design and operate scalable cloud infrastructure across AWS and GCP.
Build and improve Kubernetes, Linux, and cloud networking environments.
Strengthen security, disaster recovery, and platform resilience.
Hubstaff provides workforce analytics and time tracking for remote teams, serving over 200,000 global users. The company is a product-led organization with a winning culture and a fully remote team of experienced engineers.
Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.
Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.
Own the self-service data and storage layer for multimodal biological datasets.
Build platform services that enable self-serve data access and processing.
Engineer security into the data layer with access control and least-privilege.
Bioptimus is building the first universal AI foundation model for biology to accelerate breakthroughs in biomedicine. They are a fast-growing startup with over $75M in funding, headquartered in Paris, with a world-class team redefining AI and life sciences.
Build and operate internal platform services and APIs in Go.
Codify infrastructure with Terraform and GitOps practices.
Operate and scale multi-tenant EKS clusters and traffic systems.
Docker builds tools for developers to build, share, and run applications, trusted by over 20 million monthly users. They are a globally distributed, remote-first team with offices in Seattle and Paris, focused on innovation and inclusion.
Define architecture and best practices for the platform and infrastructure layer the product is built on.
Own the deploy pipeline and lead the move to a GitOps model (Argo) for fast, safe releases.
Design and harden multi-tenant isolation and blast-radius protection for top-tier customers, including dedicated deployments.
We are the Engineering Operations Platform - mission control for the AI software factory, providing visibility, governance, and golden paths. We are a group of 80 passionate individuals, backed by $60M Series C from Sequoia, IVP, and others, with a fully remote culture.
Architect, deliver, and maintain critical cloud platform components on AWS EKS, focusing on production reliability and observability.
Establish SRE standards including SLOs, error budgets, and automated tooling to reduce operational friction.
Provide technical advisory through code reviews and architecture recommendations to maintain high platform standards.
Inflect is a US-based advisory and marketplace revolutionizing digital infrastructure procurement. They are a small team focused on reducing friction in buying datacenter, cloud, and network services through automation and better deal terms.
Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
Scale single-tenant deployments and build observability, incident response, and compliance practices.
Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.
Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.
Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.