Design and build a managed Slurm service on Kubernetes
Write clean, reliable, and maintainable Go code, developing scheduling and orchestration for GPU workloads
Build observability and automated remediation for GPU, node, network, and control-plane failures
Gcore provides infrastructure and software solutions for AI, cloud, network, and security, powering digital experiences worldwide. With over 550 professionals, they build and support the global digital ecosystem.
Design, develop, and maintain internal software, services, and automation using Go and Rust.
Build and operate Kubernetes-based infrastructure and improve developer workflows and CI/CD.
Collaborate across teams to solve ambiguous technical challenges and improve system reliability.
Our partner is a high-growth technology organization building internal platforms to enable efficient engineering. They operate with small, autonomous teams in a high-trust, collaborative remote environment with a focus on technical excellence.
Own the technical strategy for multi-ecosystem scaling, defining architecture for onboarding new language ecosystems.
Drive end-to-end remediation automation, leading redesign of CVE workflows to close the loop from detection to verified release.
Set platform-wide technical direction spanning package index, build pipelines, and orchestration tooling to serve customers and ecosystem teams.
Chainguard is the trusted source for open source, delivering hardened, secure, and production-ready builds of open source software. They serve Fortune 500 enterprises and global industry leaders, and are venture-backed by leading investors, fostering a culture of customer obsession and intentional action.
Design, code, and debug scalable software applications and APIs in Golang with attention to performance and security.
Collaborate with product managers and cross-functional teams to translate business requirements into technical specifications.
Provide technical leadership, drive code reviews, and guide the team in AI tooling best practices.
EasyPost is a YC unicorn that makes shipping simple for businesses with the first developer-friendly REST API for shipping. Our team is rapidly growing, and we're builders and problem-solvers who move fast and innovate in an industry that needs it.
Design and build control-plane services and drivers for storage integration with Kubernetes-based AI workloads.
Write production-quality Go code with strong testing and operational rigor.
Deliver storage integration for k0s-based Kubernetes via Cluster API and K0rdent topologies.
Mirantis is a Kubernetes-native AI infrastructure company that builds scalable, secure infrastructure for AI and data-intensive applications. It is committed to open standards and freedom from lock-in, empowering platform engineering teams.
Architect and build a robust, scalable, and highly available distributed infrastructure.
Build a cutting-edge cloud-native platform on top of the public cloud and automate cloud resource management.
Work closely with core database development and security teams to produce the SaaS offering.
ClickHouse is a real-time analytics and data warehousing company recognized on the Forbes Cloud 100 list. With over 4,000 customers and rapid growth, the company is a leader in its space.
Own features end-to-end, from design through implementation and iteration.
Design and build backend systems in Go with REST/gRPC APIs for entitlements and policies.
Provide technical leadership and mentorship to senior and mid-level engineers.
Chainguard provides hardened, secure, and production-ready builds of open source software. It is venture-backed by leading investors and serves Fortune 500 enterprises like Anduril, Canva, and OpenAI.
Design and build self-service platform capabilities for engineering teams.
Develop and maintain global hybrid infrastructure across bare-metal, Linux, and Kubernetes.
Automate operational tasks and improve observability and reliability.
The partner company builds globally distributed infrastructure and platform capabilities. It is a small, highly autonomous team with a strong focus on reliability and developer experience.
Develop, test, and maintain distributed systems and microservices using Golang, Node.js, Java, and RESTful/gRPC APIs.
Deploy scalable web applications inside containerized environments using Kubernetes, Docker, and AWS.
Utilize automated CI/CD pipelines and leverage AI in the SDLC to ship code safely.
JumpCloud is the AI-powered unified IT management platform designed to secure the modern workforce. The company is remote-first with teams in 15+ countries, fostering a collaborative and fast-moving environment.
Build and run monitoring, tracing, and alerting infrastructure to ensure platform reliability and security.
Lead incident response and recovery, including root cause analysis, and improve deployment processes for fast, safe code changes.
Collaborate with engineering teams to deliver a stable, scalable platform and handle load for resource-intensive applications.
WellSaid Labs is the leading AI voiceover studio for enterprise and professional use, providing ultra-realistic voices that the world’s biggest brands trust. We are a fully distributed team across the U.S. with a focus on responsible AI and an inclusive culture.
Architect and operate distributed, fault-tolerant systems for a large-scale cloud data platform.
Lead optimization initiatives across compute, storage, networking, and infrastructure efficiency.
Collaborate with engineering teams and stakeholders to build tooling for cloud cost and resource visibility.
This company builds a highly scalable cloud data platform and focuses on infrastructure efficiency and multi-cloud environments. With operations across more than 20 countries, it fosters a globally distributed, remote-friendly culture that values ownership, collaboration, and innovation.
Collaborate with engineers and community contributors on features, bug fixes, reviews, and releases through clear communication and responsible stewardship.
Engage with open source communities and industry partners to move projects forward, including facilitating community meetings with internal and external users.
Own architecture decisions and ensure technical quality across open source codebases, APIs, and contributor workflows.
Defense Unicorns delivers mission value by streamlining software delivery so customers can focus on critical challenges. Our team is composed of innovators, software engineers, and veterans with decades of experience in federal market technology programs.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Independently resolve complex CI/CD, software, and infrastructure issues for enterprise customers.
Support customers via Slack, Zoom, email, and community forums while leading retrospectives.
Design proactive tools, publish docs, and mentor peers to evolve the support function.
Buildkite is a CI/CD platform trusted by leading engineering teams at Canva, Meta, and NVIDIA, helping ship software used by over one billion people daily. The company fosters an inclusive remote culture with support for learning, growth, and work-life balance.
Reproduce, debug, and escalate technical issues in partnership with Engineering and Product.
Act as a technical expert and initial escalation point for our SDK and Cloud products.
Work directly with developers to provide guidance and resolve complex issues.
Ditto is a peer-to-peer sync engine that enables developers to build real-time applications that stay connected without internet. With over $145 million in funding and a globally distributed team, we are a fast-growing startup committed to diversity and inclusion.
Own and ship high-impact features end-to-end, from technical design through production.
Work directly with customers and internal stakeholders to understand pain points and build solutions.
Set a high technical bar through design reviews, code review, and hands-on execution.
Float Health builds a marketplace for specialty pharmacy to bring care home. The Y Combinator-backed startup has a small engineering team, facilitated over 100,000 patient visits with 1,400 nurses, and emphasizes high autonomy.
Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.
Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.
Design and develop Go-based backend services and CLI tooling to improve internal developer experience.
Build scalable automation and governance-as-code systems for engineering workflows.
Operate and optimize systems on Kubernetes and AWS, focusing on performance and reliability.
Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud and data centers. With a crew of over 600 people across Asia-Pacific, Europe, and the Americas, the company fosters a collaborative, supportive, and fun culture.
Develop software in Go and Python, and review teammates' code
Write design docs to communicate your ideas
Build shared platforms and services that support critical government systems
Corbalt is a technology company that partners with federal agencies to modernize and operate complex technology ecosystems. They are a remote-first team that values curiosity, kindness, ownership, and continuous learning.
Build and operate monitoring, tracing, alerting, and observability infrastructure for system reliability.
Drive platform security initiatives with preventative controls and resilient architecture.
Lead incident response and recovery, including root-cause analysis and preventative measures.
This role is with a partner company managing AI-powered products. They are a growing technology organization with a fully distributed US-based team and a collaborative culture focused on large-scale infrastructure and AI technology.