Design and build self-service platform capabilities for engineering teams.
Develop and maintain global hybrid infrastructure across bare-metal, Linux, and Kubernetes.
Automate operational tasks and improve observability and reliability.
The partner company builds globally distributed infrastructure and platform capabilities. It is a small, highly autonomous team with a strong focus on reliability and developer experience.
Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.
Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.
Own and scale cloud infrastructure including compute, networking, storage, and data systems.
Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
Establish reliability through service-level objectives, observability, and incident response processes.
A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.
Create and test reliable cloud infrastructure services supporting Webflow's product range.
Lead initiatives to reduce triage load, increase reliability, and handle growing customer scale.
Collaborate with product engineering teams to deliver new solutions and improve existing services.
Webflow is an agentic web marketing platform that helps modern marketing teams build, manage, and optimize high-performing web experiences. The company values grit, speed, and craft, fostering a culture of ownership and continuous improvement.
Build and operate the Kubernetes platform supporting AI test and evaluation frameworks.
Design infrastructure-as-code, GitOps workflows, and automated deployment pipelines.
Own platform reliability, observability, capacity planning, and operational readiness.
OpenTeams helps enterprises and governments build AI they control, govern, and evolve themselves. Founded by the creator of NumPy and SciPy, the company is built by people with deep roots across the open-source ecosystem and maintains a remote-first culture.
Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
Design and maintain infrastructure as code across multiple cloud providers.
Provide technical leadership and mentorship across the Systems Engineering team.
Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.
Drive the performance, stability, security, and reliability of production environments with a focus on automation and proactive improvements.
Design and maintain infrastructure using Infrastructure as Code tools like Terraform, and manage Kubernetes and cloud environments.
Lead vulnerability management, incident response, and secure CI/CD practices to ensure resilience and operational excellence.
Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. It processes applications and shares shortlists with employers, offering a remote-first and inclusive work environment.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Own the technical direction and architecture of critical infrastructure domains, establishing scalable patterns and standards.
Lead complex, multi-team infrastructure initiatives from design through implementation and production operation.
Design and evolve AWS and Kubernetes infrastructure to enable teams to build and deploy systems reliably at scale.
We provide innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. Our company is backed by world-class investors including Craft Ventures and Andreessen Horowitz, with offices across the US and India, and we are growing extremely quickly.
Own the design, development, and operation of infrastructure and build/release pipelines.
Deploy IaC and automation using Terraform, Ansible, Helm, and Go to support platform and customer requirements.
Collaborate closely with Product to drive roadmap direction and improve how users build and deliver software.
Manifest is on a mission to secure the global software and AI supply chain. Founded by alumni from the Department of Defense, CISA, and Palantir, it is a well-funded early-stage startup backed by leading investors and trusted by government and enterprise organizations.
Own and improve production infrastructure reliability and stability.
Prepare, execute, and support deployments and infrastructure changes.
Build and maintain Infrastructure-as-Code solutions using Ansible and Terraform.
Social Discovery Group (SDG) is a group of social discovery companies that solve problems of loneliness, isolation, and disconnection by transforming virtual intimacy into the new normal. Our international team of digital nomads works remotely from all over the world and we are proud to be a two-time 'Great Place to Work' winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025).
Design and advance core infrastructure for multi-cloud Kubernetes clusters and developer toolchains.
Automate operations and engineering tasks to improve productivity and reliability.
Build machine learning infrastructure to enable AI teams to train and deploy large-scale models.
Cresta provides an AI platform that transforms customer conversations into competitive advantages by combining conversational AI, real-time agent augmentation, and conversation intelligence. The company has raised over $270 million from top investors like a16z, Greylock, and Sequoia, and is led by AI industry veterans.
Develop and evolve foundational software and services enabling product and development teams.\n- Architect, design, and implement Infrastructure as Code using Terraform.\n- Deploy, manage, and optimize Kubernetes clusters on GCP (GKE) and AWS (EKS).
DoiT is a global technology company that helps cloud-driven organizations leverage cloud for business growth and innovation through data, technology, and human expertise. They work with over 4,000 customers worldwide and foster a remote-first, entrepreneurial culture.
Automate and build runtime production environments, serving as the link between application development and platform teams.
Validate solutions and implementations to ensure alignment with business requirements and maintain platform integrity.
Independently solve business problems and develop small components to address challenges without explicit architecture diagrams.
Defense Unicorns delivers mission value by streamlining software delivery so our customers can focus on the most important challenges. Our team is composed of innovators, software engineers, and veterans with decades of experience delivering technology programs across the federal market.
Build and run monitoring, tracing, and alerting infrastructure to ensure platform reliability and security.
Lead incident response and recovery, including root cause analysis, and improve deployment processes for fast, safe code changes.
Collaborate with engineering teams to deliver a stable, scalable platform and handle load for resource-intensive applications.
WellSaid Labs is the leading AI voiceover studio for enterprise and professional use, providing ultra-realistic voices that the world’s biggest brands trust. We are a fully distributed team across the U.S. with a focus on responsible AI and an inclusive culture.
Support the architecture, deployment, and testing of bare-metal infrastructure including GPU compute, storage, and networking.
Deploy, configure, maintain, and troubleshoot UDS and Kubernetes-based platform services.
Develop and manage infrastructure as code (IaC) to enable consistent, repeatable deployments.
Defense Unicorns delivers mission value by streamlining software delivery for mission-focused customers. The team is composed of innovators, software engineers, and veterans with decades of experience.
Apply SRE principles to improve reliability, scalability, and performance of production systems.
Design and implement automation to reduce operational toil and improve engineering efficiency.
Lead incident response and develop sustainable solutions for complex production issues.
The hiring company is a technology organization focused on reliability and operational excellence. They offer a fully remote, collaborative environment with opportunities for technical leadership and career growth.
Own and operate production infrastructure across Kubernetes, Linux, networking, and virtualization.
Lead incident response and implement observability to improve availability and performance.
Define SLOs and automate infrastructure with Ansible, Bash, Python, and GitOps.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective, data-driven processes. They foster a collaborative, international, and fully remote work environment, emphasizing autonomy and ownership for their small to mid-sized team.
Own and scale the cloud infrastructure behind our open-source platform: compute, networking, and the data layer.
Lead BYOC: turn customer-cloud deployments into a real product, with provisioning, upgrades, and observability that scale past bespoke work per deal.
Make reliability a product feature: meaningful SLOs, and an incident process people trust.
Nango is a developer infrastructure company that provides API access for agents and apps, enabling AI applications to connect to the real world through integrations. With over 400 paying customers and a team of 14 from top tech companies like AWS, GitHub, and Okta, they are a YC-backed, multi-million ARR company that values ownership and autonomy.
Build and operate monitoring, tracing, alerting, and observability infrastructure for system reliability.
Drive platform security initiatives with preventative controls and resilient architecture.
Lead incident response and recovery, including root-cause analysis and preventative measures.
This role is with a partner company managing AI-powered products. They are a growing technology organization with a fully distributed US-based team and a collaborative culture focused on large-scale infrastructure and AI technology.