Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
Drive AI-specific observability, FinOps, and security practices across the platform.
We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.
Design and operate the infrastructure for a high-throughput messaging platform operating at 500K+ events/sec.
Build guardrails, runbooks, and validation gates that enable AI agents to safely execute deployments and operations.
Lead incident response and encode every fix as a new runbook and regression test.
Postscript is an AI messaging platform trusted by 20,000+ Shopify brands to drive revenue through SMS. The company is fully remote, backed by Greylock and Y Combinator, and has a culture of ownership and innovation.
Operate the Monad node fleet, including health, sync, upgrades, and incident response for validators, full nodes, and archive nodes.
Own infrastructure-as-code with Ansible, Terraform, and Kubernetes, and build observability with Prometheus, Grafana, and Loki.
Design and build AI agent tooling for automated operations, including runbooks-as-code and deterministic guardrails.
Category Labs designs and builds decentralized technology, including the Monad blockchain, a high-performance EVM-compatible Layer 1. The team raised $225M in series A funding and is a lean, collaborative group of engineers and researchers with a culture of low ego and high-quality output.
Design and maintain AWS infrastructure using Terraform, with a focus on scalability cost and PCI-scoped network segmentation
Build and evolve the observability stack and CI/CD pipelines to ensure smooth production operations and rapid deployment
Lead incident response define SLOs and run performance tests to optimize payment-critical services
Xplor Technologies provides vertical software, embedded payments, and AI tools for membership-based and service-based industries. With over 130,000 businesses in 72+ countries and processing $47 billion in payments annually, the company values diversity, collaboration, and a people-first culture.
Own the technical operations domain covering IT, security, DevOps, and infrastructure for an AWS-native platform.
Build and lead a lean team, leveraging AI agents and tooling to automate DevOps, security, and compliance tasks.
Partner with engineering leadership to ensure reliability, scalability, and HIPAA/CMS compliance in a regulated environment.
Spark Advisors builds healthcare technology for Medicare advisors, helping seniors navigate complex coverage. They are a fast-growing platform backed by Primary Ventures and Viewpoint Ventures, recognized as one of Inc. Magazine's Best Workplaces of 2025.
Evolving the AI knowledge platform with retrieval, indexing, and synthesis for organization-wide use.
Architecting and operating agentic infrastructure on AWS with cost guardrails and observability.
Partnering with product engineering to define the AI platform API surface and building reference agent implementations.
ShiftKey is a healthcare workforce marketplace that connects facilities with licensed professionals to fill shifts, addressing staffing shortages. The company fosters an inclusive and collaborative culture, valuing diverse perspectives.
Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
Use AI agents as force multipliers to automate manual processes and improve developer experience.
Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.
Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.
Design, build, and operate core cloud infrastructure on AWS, including compute, networking, and container orchestration.
Own the CI/CD platform used across engineering teams, including build pipelines, environment promotion, and progressive rollout.
Build and maintain the observability stack across the organization, including logging, metrics, distributed tracing, and alerting.
RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company is an Equal Opportunity and Affirmative Action employer committed to diversity and collaboration.
Architect and implement infrastructure for seamless migrations of critical systems, keeping uptime and reliability front and center.
Design, build, and maintain the tooling and processes that let our SaaS products scale and run themselves on AWS.
Partner closely with developers to strengthen the reliability, performance, security, and scalability of our application architecture.
Duetto is the hospitality industry's leading revenue management platform, providing a suite of tools for hotels, resorts, and casinos. Backed by GrowthCurve Capital since 2024, the company has been named the #1 Best Place to Work in Hotel Tech in 2025 and is passionate about the industry it serves.
Architect and scale high-availability backend systems across AWS.
Build CI/CD, developer environments, and AI workloads infrastructure.
Develop advanced autoscaling, caching, and storage strategies.
Source.dev builds software tools to simplify device software development. The company has a high-trust, passionate engineering culture with a focus on end-to-end ownership.
Own Primer's internal developer platform end to end, including CI/CD pipelines, deployment workflows, and self-service tooling.
Build the human-AI development loop, creating tooling and automation for coding agent workflows.
Treat developer productivity as a measurable system, using frameworks like DORA to identify and fix delivery bottlenecks.
Primer provides a unified infrastructure for global payments, enabling finance and payments teams to reduce complexity and capture revenue. Backed by top investors like Sofina and Accel, they operate as a remote-first, async culture with high autonomy and low bureaucracy.
Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
Manage distributed systems, observability, incident response, and automation with a security-first mindset.
Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.
Take full ownership of the company's platform and infrastructure function, defining and executing the platform engineering strategy in partnership with the CTO.
Lead, mentor, and develop an established team of three DevOps Engineers and one SRE, setting clear responsibilities and development plans.
Own platform-related budgeting, cloud spend optimization, and technology decisions to improve reliability, security, and developer productivity.
Our client is a remote-first digital product company building and scaling SaaS products, AI-powered solutions, and web and mobile applications for global markets. The company operates across more than 20 countries with approximately 40–50 engineers supporting a portfolio of around 15 digital products.
Engineer and maintain optimal performance and availability of AWS environments using IaC and automation.
Diagnose and resolve complex cloud infrastructure issues, providing support for cloud systems.
Create and maintain technical documentation for architecture decisions, automation processes, and compliance artifacts.
Millennium is a cybersecurity company providing mission-critical support for national security, with a focus on Red Team Operations and Defensive Cyber Operations. The team of over 300 professionals supports the Department of Defense and federal civilian customers, fostering a culture of expertise and equal opportunity.
Own the operational backbone of an AI-native development workflow, including environment setup, CI/CD, testing, documentation, and API prototyping.
Prepare dev environments, run test suites, review code, manage deployments, and handle lower-complexity coding tasks using AI tools.
Research and prototype third-party API integrations and maintain technical documentation for regulatory compliance.
SPHERE Technology is a fully capitalized HealthTech SaaS company building an AI-native platform from scratch, backed by a Miami-based family office with a long runway and zero burn-rate pressure. The company is a startup with a small, focused team, ambitious product roadmap, and a direct access culture with no layers or committees.
Design, implement, and maintain AWS cloud infrastructure using Terraform.
Build and optimize CI/CD pipelines to enable rapid, safe deployments across multiple environments.
Own observability strategy with comprehensive monitoring, logging, and alerting systems using Datadog.
Avantos is building an AI-native operating system for financial services, transforming fragmented data into a single intelligent system. They are a product-led, fast-moving team at the intersection of AI, fintech, and modern infrastructure.
Drive deployment automation across EKS clusters using GitOps.
Own infrastructure requirements and coordinate maintenance backlog.
Collaborate with engineering teams and integrate AI agents to accelerate workflows.
Ping Identity provides an intelligent cloud identity platform that secures digital experiences. With global offices and serving over half of the Fortune 100, the company fosters a culture of respect and individuality.
Design, build, and optimize cloud infrastructure (AWS/Kubernetes/EKS) and CI/CD pipelines across multiple teams.
Troubleshoot and resolve production incidents of varying scope, ensuring reliability and performance.
Drive infrastructure projects end-to-end, mentor engineers, and establish standards that improve developer productivity.
Pacvue is a leading Commerce Media OS powering over $12B in advertising spend across 100+ global retail media networks. It enables over 70,000 brands and agencies with an inclusive global community that fosters innovation and career growth.
Set technical direction and design durable platform abstractions to enable rapid innovation.
Lead complex cross-team initiatives and establish best-in-class platform standards.
Pioneer the use of AI as a code collaborator to ship production code faster while maintaining quality.
Fanatics is building a leading global digital sports platform, offering products and services across Commerce, Collectibles, and Betting & Gaming. With over 22,000 employees, they are committed to enhancing the fan experience and delighting sports fans globally.
Design and maintain AWS cloud infrastructure using OpenTofu and Terraform.
Operate Kubernetes workloads on Amazon EKS, managing GitOps deployments with Argo CD and Helm.
Implement observability with Datadog, troubleshoot production incidents, and support on-call rotation.
PAR Technology Corporation provides innovative restaurant technology solutions, including point-of-sale, digital ordering, loyalty, and back-office software, as well as hardware and drive-thru offerings. With over 40 years of experience, the company serves more than 100,000 restaurants globally and fosters a collaborative culture centered on its 'Better Together' ethos.