Design, build, and operate production Kubernetes platforms, owning cluster architecture, networking, reliability, and security.
Troubleshoot complex infrastructure issues across Kubernetes, Linux, networking, and cloud environments, and improve observability and automation.
Take ownership of critical infrastructure initiatives, incident response, and mentorship while working autonomously in a remote-first environment.
Our partner is a technology company building and operating large-scale Kubernetes platforms and complex production infrastructure. It is a remote-first organization with a focus on autonomy, technical ownership, and collaboration across North America and Latin America.
Own and operate production infrastructure across Kubernetes, Linux, networking, and virtualization.
Lead incident response and implement observability to improve availability and performance.
Define SLOs and automate infrastructure with Ansible, Bash, Python, and GitOps.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective, data-driven processes. They foster a collaborative, international, and fully remote work environment, emphasizing autonomy and ownership for their small to mid-sized team.
Design and manage high-availability platforms using Kubernetes, Terraform, and Ansible with native-AI capabilities.
Develop and operate the observability stack: Grafana, Mimir, Loki, Tempo, and Prometheus on Kubernetes via GitLab CI/CD.
Build automation scripts in Python, maintain GitOps pipelines, and mentor mid-level engineers.
Flexential builds and operates critical IT platforms including observability, DevOps, and ITSM technologies. The company fosters a collaborative engineering culture and values diversity.
Design, build, and operate secure, resilient Kubernetes platforms for DoD environments.
Automate infrastructure and platform configuration using Terraform, Ansible, Helm, and Infrastructure as Code.
Implement Kubernetes networking, RBAC, security controls, and service mesh capabilities such as Istio.
Rackner builds cloud-native platforms and software for complex federal missions, including DevSecOps, AI/ML, data, and modern software engineering. Its culture emphasizes hands-on engineering, career development, and remote work support.
Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.
Jobgether is a platform that connects job seekers with opportunities through an AI-powered matching process. The company fosters a remote-first culture and emphasizes autonomy and ownership for engineers.
Own the observability, logging and alerting for Kubernetes clusters and critical workloads.
Build and maintain automation for lifecycle management of Kubernetes clusters.
Identify and root-fix reliability bottlenecks before they become incidents.
Wrapbook is an AI platform for production finance, built for feature films and TV, trusted by Netflix and Paramount. Backed by top investors, our team of over 350 employees uses AI to transform how finance teams work.
Own the design, development, and operation of infrastructure and build/release pipelines.
Deploy IaC and automation using Terraform, Ansible, Helm, and Go to support platform and customer requirements.
Collaborate closely with Product to drive roadmap direction and improve how users build and deliver software.
Manifest is on a mission to secure the global software and AI supply chain. Founded by alumni from the Department of Defense, CISA, and Palantir, it is a well-funded early-stage startup backed by leading investors and trusted by government and enterprise organizations.
Build and operate reliable, scalable cloud infrastructure on AWS and Kubernetes.
Own production infrastructure, containerized applications, deployment workflows, and monitoring.
Collaborate with development teams to streamline CI/CD and drive high availability.
Our partner is a fast-growing AdTech and e-commerce platform. They offer a flexible, remote-first culture that values ownership, proactive problem-solving, and continuous improvement.
Build, maintain, and release C8 distribution artifacts for self-managed customers, including Helm Charts, C8Run, and Docker Compose, along with documentation.
Improve reliability and usability of deployment artifacts across the full customer lifecycle: installation, upgrades, and day-2 operations.
Engage directly with customer-facing teams to understand real deployment problems and shape prioritization, collaborating with Development, Product Management, QA, and Documentation.
Camunda is the enterprise platform for agentic orchestration, enabling organizations to coordinate AI agents, people, and systems across complex business processes. Trusted by over 700 organizations worldwide, including 9 of top 10 US banks, Camunda is a fully remote and global company transforming into an AI-first organization.
Lead the design, implementation, and ongoing improvement of reliable, scalable, and secure production platforms and services.
Work closely with cross-functional teams to build and maintain resilient infrastructure and deployment patterns.
Provide technical leadership and mentorship, promoting strong engineering standards and operational best practices.
Cision is a global leader in PR, marketing and social media management technology and intelligence, helping brands connect with customers and stakeholders. They have offices in 24 countries, a network of over 1.1 billion influencers, and a culture that champions diversity, equity, and inclusion.
Design, implement, and maintain reliable, scalable, and secure infrastructure to support applications and automation systems.
Automate infrastructure provisioning, configuration management, and deployment pipelines using tools like Terraform and ArgoCD.
Implement observability solutions and enforce security best practices to ensure uptime and system performance.
Bright Machines is a next-generation, AI-enabled manufacturer focused on data center infrastructure production, using proprietary AI-based robotics and software to assemble hardware products for hyperscalers and OEMs. The company is headquartered in San Francisco, California, with an integration center in Guadalajara, Mexico, and has been recognized by Forbes' AI 50 and other leading organizations.
Design and implement security-focused tooling and guardrails for infrastructure.
Collaborate with security, SRE, and platform teams to integrate trust principles.
Build internal automation to reduce developer cognitive load while enforcing security standards.
BlueCloud is a Snowflake Elite Partner that helps enterprises migrate to AI-ready data platforms. With 450+ consultants and 200+ successful transformations, they combine expertise with AI accelerators to deliver results 40-50% faster.
Own and evolve production infrastructure, leading the migration from Docker Swarm to Kubernetes on premises.
Drive observability, enforce IaC practices, and ensure CI/CD reliability across ~50 services.
Participate in on-call rotation, resolve incidents, and build platform tooling to reduce infrastructure toil.
Webshare is an enterprise-grade proxy platform providing access to over 80 million global IPs across 195 countries. With 99.97% uptime, it serves tens of thousands of businesses, and the team emphasizes mentorship, knowledge-sharing, and team events.
Architect and automate scalable cloud environments across AWS and Azure using Terraform, Ansible, Helm, and CDK.
Serve as Linux subject matter expert, managing system builds, core services, and performance from kernel up.
Lead CI/CD pipelines, observability, security, and incident response to ensure platform reliability.
Fueled is a leading digital strategy, design, and engineering agency. The 300+ person team has designed and built hundreds of digital products for major brands like Google, Apple, and The New York Times, and thrives in a culture that values flexibility, creativity, and cutting-edge technology.
Design, develop, and maintain cloud and containerized platforms on AWS and Kubernetes.
Automate deployments, infrastructure, and monitoring using Terraform, GitHub Actions, and GitOps.
Mentor junior engineers and lead technical projects for medium to large-scale initiatives.
New Era Technology securely connects people, places, and information with end-to-end technology solutions at scale. With a global team of over 3,000 professionals, we foster a team-oriented culture that prioritizes personal and professional development.
Design and implement scalable Kubernetes infrastructure for high-throughput event processing.
Build cloud-agnostic environments and implement GitOps workflows using Terraform.
Manage databases, monitoring, and security in production Kubernetes environments.
Miratech is a global IT services and consulting company that helps visionaries change the world. The company retains nearly 1000 full-time professionals and has a culture of relentless performance with a 99% project success rate.
Contribute to platform and harness engineering, including CI/CD and developer tooling.
Build systems to reduce toil and maintain production infrastructure under conversational AI traffic.
Participate in on-call rotation and incident management to ensure platform uptime.
Replicant builds an AI-powered customer service platform that helps contact centers resolve requests and improve agent performance. The company is distributed, with a focus on ownership and collaboration, and serves Fortune 500 companies.
Architect and scale multi-region microservices, APIs, and authentication infrastructure on AWS/GCP.
Lead SLOs, observability, incident management, and disaster recovery automation to maintain 99.99% availability.
Manage Kubernetes clusters and Terraform IaC while eliminating toil with Python/Go tooling.
JumpCloud is an AI-powered unified IT management platform that secures the modern workforce through identity, device, and access management. The company is remote-first with teams in 15+ countries and values building connections, thinking big, and continuous improvement.
Own infrastructure end to end, including Kubernetes clusters, Docker builds, and DigitalOcean services.
Move production deploys toward safer, more frequent releases and manage capacity, autoscaling, and load balancing.
Operate PostgreSQL, Redis, and NATS under real traffic, keep Cloudflare tight, and document systems for the whole team.
Colonist is building the biggest digital board game platform on the internet, where players have spent over 3,000 years playing 60 million games on their Catan alternative. The company is a fully remote, asynchronous team spread across multiple continents, focused on crafting polished digital board games.
Design, deploy, and sustain AWS-based platform services for mission-critical Department of War applications.
Administer Kubernetes clusters including lifecycle management, security, and observability.
Build CI/CD pipelines and infrastructure-as-code using Terraform, GitOps, and automation tools.
LMI is a digital solutions provider accelerating government impact with innovation and speed. Headquartered in Tysons, Virginia, the company serves defense, space, healthcare, and energy sectors with a focus on agility and collaboration.