Design and develop foundational components and frameworks for our Agentic AI platform.
Collaborate with cross-functional teams to create platform solutions that empower developers.
Provide production support and ensure platform stability, working closely with ML engineers.
Legion builds secure, reliable AI systems that integrate with complex platforms, optimizing workflows and enhancing human capability. They work with partners like Palantir, Nvidia, HPE, and Oracle, and are looking for bold thinkers to shape the future of grounded AI.
Build and maintain backend services for our LLM gateway, including routing, rate limiting, and observability.
Contribute to sandboxing and isolation infrastructure for safe agent-generated code execution.
Write high-performance backend code in Go, Rust, or async Python, supporting Kubernetes-based platform services.
AZX accelerates positive impact in critical industries through AI transformation. Founded in 2024 and profitable from the start, we work with category leaders in real estate, energy, logistics, and utilities.
You own the platform runtime, voice infrastructure, enterprise security, and evaluation layer for all clients.
Work directly with the founder on architecture, from design through production and incident response.
Build self-serve infrastructure, lead incident response, and keep the platform fast and reliable on Kubernetes.
Rifa AI builds an AI agents platform for contact centers in regulated industries, helping enterprises deploy trustworthy voicebots. They are a small, passionate engineering team with paying enterprise clients and growing revenue, backed by Seaborne Capital.
Design and build LLM serving infrastructure on Kubernetes, including deployment, GPU scheduling, and model lifecycle management.
Package the platform for enterprise environments with Helm-based installs and support for restricted or offline networks.
Integrate the serving layer with API gateway, identity, and metering services, and build observability for GPU inference in production.
Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build and operate scalable infrastructure for AI and data-intensive applications. It is part of an IREN company and empowers platform engineering teams with open-source innovation and deep expertise in Kubernetes orchestration.
Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
Together AI is a research-driven artificial intelligence company that builds open and transparent AI systems. The company has contributed to leading open-source research like FlashAttention and RedPajama, and aims to lower the cost of modern AI through co-designed software, hardware, algorithms, and models.
Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.
Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
Drive AI-specific observability, FinOps, and security practices across the platform.
We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.
Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.
Define architecture and best practices for the platform and infrastructure layer the product is built on.
Own the deploy pipeline and lead the move to a GitOps model (Argo) for fast, safe releases.
Design and harden multi-tenant isolation and blast-radius protection for top-tier customers, including dedicated deployments.
We are the Engineering Operations Platform - mission control for the AI software factory, providing visibility, governance, and golden paths. We are a group of 80 passionate individuals, backed by $60M Series C from Sequoia, IVP, and others, with a fully remote culture.
Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.
Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.
Design and build sandboxed evaluation environments for AI models to safely execute code and interact with tools.
Build backend services and infrastructure supporting large-scale AI and agentic evaluations.
Develop agent scaffolding, evaluation harnesses, and systems for provisioning isolated environments using Docker, Kubernetes, and cloud infrastructure.
10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. The company focuses on adversarial red teaming and model evaluations to help teams deploy AI systems safely.
Design, build, and maintain core AI capabilities that power intelligent experiences across the platform.
Develop reusable capabilities for model integration, reasoning, and evaluation.
Optimize AI systems through experimentation, quantitative evaluation, and iterative refinement.
Legion embeds intelligence inside complex systems, unlocking data and accelerating human workflows for government and enterprise. They collaborate with world-class partners like Palantir and Nvidia, building intelligent infrastructure with a team of bold thinkers.
Own the technical path from customer interest to working deployment, integrating the platform into production AI environments.
Build and operate AI/MLOps pipelines, debug complex environments, and create prototypes and demos.
Translate customer needs into product improvements, partnering with Sales, Product, and Engineering.
Neuromorphic Labs is a Seed-stage AI startup building a trust layer for production AI, making security, governance, and control intrinsic to every model and deployment. Backed by top-tier VCs, the team is small and fast-paced, emphasizing ownership, high standards, and collaboration with exceptional builders.
Operate the Monad node fleet, including health, sync, upgrades, and incident response for validators, full nodes, and archive nodes.
Own infrastructure-as-code with Ansible, Terraform, and Kubernetes, and build observability with Prometheus, Grafana, and Loki.
Design and build AI agent tooling for automated operations, including runbooks-as-code and deterministic guardrails.
Category Labs designs and builds decentralized technology, including the Monad blockchain, a high-performance EVM-compatible Layer 1. The team raised $225M in series A funding and is a lean, collaborative group of engineers and researchers with a culture of low ego and high-quality output.
Design, build, and own shared infrastructure for Interpretability research environments, data systems, and compute tooling.
Lead cross-team efforts with agentic engineering, security, compute, and storage platform teams.
Discover and resolve major organization-wide developer experience issues and help take interpretability methods from research code to dependable audit pipelines.
Anthropic is an AI safety company focused on building reliable, interpretable, and steerable AI systems. They are a quickly growing team of researchers, engineers, policy experts, and business leaders committed to beneficial AI.
Architect, design, build, deploy, and maintain Model Serving infrastructure using industry-standard AI tools.
Own projects that scale model serving and data processing services to handle 10x traffic.
Collaborate closely with MLE and Data Science teams to distill feedback and execute on strategy.
Abnormal AI protects the humans behind the world's most critical organizations from AI-powered cybercrime. Over 4,500 enterprises trust their behavioral AI platform, fostering a culture of innovation and security.
Design and build production AI agent systems for diagnosing and remediating infrastructure issues in large-scale GPU environments.
Develop distributed services, orchestration frameworks, knowledge graphs, and retrieval systems to power AI agents.
Own services end-to-end from architecture through production, collaborating with infrastructure and engineering teams.
The company builds AI agents that operate and automate large-scale GPU infrastructure. The engineering team is highly collaborative and remote, fostering ownership and autonomy.
Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
Collaborate with GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.
Bitdeer provides comprehensive Bitcoin mining solutions and AI computational infrastructure. The company operates globally with data centers in multiple countries and focuses on AI and blockchain technology.
Build and deploy production-grade LLM inference systems from scratch, owning the pipeline from query to response.
Optimize inference workloads for latency, throughput, and cost using tools like vLLM, SGLang, and TensorRT-LLM.
Collaborate with the CTO and Product to define the technical roadmap for inference infrastructure as the organization scales.
The company is an open-source-oriented startup building production-grade inference infrastructure for large language models. It is a remote-first, globally distributed team that values engineering ownership, speed, and customer impact.