Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.
Own infrastructure as code across development, staging, and production environments
Build, maintain, and improve CI/CD pipelines for reliable and efficient deployments
Manage cloud infrastructure, establish scalable engineering practices, and lead incident response
CelebriOS is a software company building B2B SaaS products that help businesses make better decisions and streamline operations. The company has a remote-first working environment and a benefits package designed to support their team.
Design and maintain AWS infrastructure using Terraform, with a focus on scalability cost and PCI-scoped network segmentation
Build and evolve the observability stack and CI/CD pipelines to ensure smooth production operations and rapid deployment
Lead incident response define SLOs and run performance tests to optimize payment-critical services
Xplor Technologies provides vertical software, embedded payments, and AI tools for membership-based and service-based industries. With over 130,000 businesses in 72+ countries and processing $47 billion in payments annually, the company values diversity, collaboration, and a people-first culture.
Own and operate a production agentic AI platform on AWS, ensuring reliability and scaling.
Lead infrastructure automation and release management, driving best practices in security and compliance.
Collaborate with the platform team on agile ceremonies and proactively communicate status to stakeholders.
Inizio Evoke is a healthcare communications company dedicated to making health more human. As part of the larger Inizio network, it emphasizes a collaborative, inclusive culture where employees are encouraged to be their authentic selves.
Design and operate scalable cloud infrastructure across AWS and GCP.
Build and improve Kubernetes, Linux, and cloud networking environments.
Strengthen security, disaster recovery, and platform resilience.
Hubstaff provides workforce analytics and time tracking for remote teams, serving over 200,000 global users. The company is a product-led organization with a winning culture and a fully remote team of experienced engineers.
Own and evolve Kubernetes and cloud infrastructure on AWS for scalability, reliability, and usability.
Design and improve CI/CD pipelines and developer workflows to enable fast, safe, repeatable deployments.
Work cross-functionally with product engineers to understand needs and enable them through tooling and best practices.
Artsy is an online platform that connects collectors, artists, and gallerists to make the art world more accessible. The company values an inclusive culture and a diverse workforce, with a team that operates with open-source principles and a focus on impact.
Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
Drive AI-specific observability, FinOps, and security practices across the platform.
We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.
Design, build, and operate core cloud infrastructure on AWS, including compute, networking, and container orchestration.
Own the CI/CD platform used across engineering teams, including build pipelines, environment promotion, and progressive rollout.
Build and maintain the observability stack across the organization, including logging, metrics, distributed tracing, and alerting.
RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company is an Equal Opportunity and Affirmative Action employer committed to diversity and collaboration.
Own Primer's internal developer platform end to end, including CI/CD pipelines, deployment workflows, and self-service tooling.
Build the human-AI development loop, creating tooling and automation for coding agent workflows.
Treat developer productivity as a measurable system, using frameworks like DORA to identify and fix delivery bottlenecks.
Primer provides a unified infrastructure for global payments, enabling finance and payments teams to reduce complexity and capture revenue. Backed by top investors like Sofina and Accel, they operate as a remote-first, async culture with high autonomy and low bureaucracy.
Design, build, and optimize cloud infrastructure (AWS/Kubernetes/EKS) and CI/CD pipelines across multiple teams.
Troubleshoot and resolve production incidents of varying scope, ensuring reliability and performance.
Drive infrastructure projects end-to-end, mentor engineers, and establish standards that improve developer productivity.
Pacvue is a leading Commerce Media OS powering over $12B in advertising spend across 100+ global retail media networks. It enables over 70,000 brands and agencies with an inclusive global community that fosters innovation and career growth.
Architect, build, and operate secure, multi-account AWS environments using modern Infrastructure as Code.
Design and optimize container orchestration with AWS ECS (Fargate) and Kubernetes (EKS) for specialized workloads.
Build unified CI/CD pipelines with GitHub Actions, embed security by design, and establish observability with Prometheus, Grafana, and CloudWatch.
Berlitz is a global language education company with a history of nearly 150 years, currently undergoing a digital transformation. The company fosters a remote-first, AI-native culture with a focus on autonomy and greenfield development, though employee count is not specified.
Maintain and improve production and staging infrastructure for high availability and scalability. - Design, build, and optimize CI/CD pipelines ensuring reliable deployments. - Automate infrastructure provisioning using Infrastructure as Code and improve monitoring and alerting.
They are an AI-driven payment platform processing millions of transactions across 80+ payment methods including cryptocurrency. Their team of 80+ professionals works in a hybrid format across multiple offices and remotely globally.
Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
Scale single-tenant deployments and build observability, incident response, and compliance practices.
Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.
Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.
Build and operate the self-service infrastructure platform where developers and agents can validate changes in minutes.
Build golden paths for CI/CD, GitOps, and IaC to enable self-service provisioning and shipping.
Own reliability and observability, carrying on-call and turning recurring toil into automation.
Luxury Presence is building the AI growth platform for real estate. Backed by Bessemer Venture Partners, the company is a Series C firm with over 90,000 real estate professionals and has been ranked on the Inc. 5000 fastest-growing companies list three years in a row.
Serve as the technical backbone of cloud infrastructure operations, bridging incident detection and advanced architecture.
Build and maintain CI/CD pipelines, design IaC modules, and optimize cloud resources for performance and cost efficiency.
Lead observability initiatives, integrate DevSecOps practices, and collaborate with cross-functional teams to ensure robust cloud reliability.
CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. They operate with a nearshore model and focus on empowering businesses through staff augmentation, dedicated teams, and software engineering.
Design, build, and operate AWS infrastructure across multiple regions using Kubernetes, Terraform, and Helm.
Own large-scale object storage environments exceeding 15 petabytes, optimizing reliability, performance, scalability, and cost.
Build an internal developer platform for self-service infrastructure, strengthen security practices, and lead incident response and postmortems.
Our partner builds an AI-powered sports media platform handling petabyte-scale media storage and millions of minutes of video monthly. It operates as a remote-first, international, engineering-led team with significant autonomy and a focus on high availability and security.
Take an active role as co-owner of production services to ensure they are built, maintained, and operated in a reliable and scalable way.
Collaborate with Software Engineering to drive operational improvements through metric-driven analysis and help scale AWS and Kubernetes infrastructure.
Participate in a weekly on-call rotation to investigate and resolve potential system issues, and automate routine tasks in at least two programming languages.
Zerohash is the leading crypto and stablecoin infrastructure platform, powering the next generation of financial services for banks, brokerages, fintechs, and payment companies. Founded in 2017, the company has raised over $280 million from top venture firms and strategic investors, and is trusted by global brands like Morgan Stanley and Stripe, operating with a compliance-first approach.
Design and maintain AWS cloud infrastructure using OpenTofu and Terraform.
Operate Kubernetes workloads on Amazon EKS, managing GitOps deployments with Argo CD and Helm.
Implement observability with Datadog, troubleshoot production incidents, and support on-call rotation.
PAR Technology Corporation provides innovative restaurant technology solutions, including point-of-sale, digital ordering, loyalty, and back-office software, as well as hardware and drive-thru offerings. With over 40 years of experience, the company serves more than 100,000 restaurants globally and fosters a collaborative culture centered on its 'Better Together' ethos.
Help engineering teams ship code faster, safely, and on reliable infrastructure.
Split time between embedded work in domain squads, central platform work building shared tools, and AI integration.
Key responsibilities include internal tooling, core infrastructure management, automation, and production support.
Bloom & Wild Group is Europe's largest direct-to-consumer flower and gifting business. They have a 60+ person Tech team that builds the software powering their online shops, warehouses, and delivery logistics.
Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.
ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.
Design and secure scalable AWS infrastructure for high-volume browser rendering and enterprise workloads.
Manage infrastructure-as-code with Terraform and improve CI/CD workflows for engineering efficiency.
Collaborate with engineering teams on security reviews, incident response, and long-term cloud architecture strategy.
The company builds a widely used developer platform supporting millions of browser sessions. They offer a remote-first culture with collaborative teams focused on quality and innovation.