Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.
Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.
Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
Together AI is a research-driven artificial intelligence company that builds open and transparent AI systems. The company has contributed to leading open-source research like FlashAttention and RedPajama, and aims to lower the cost of modern AI through co-designed software, hardware, algorithms, and models.
Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.
Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.
TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.
Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
Collaborate with GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.
Bitdeer provides comprehensive Bitcoin mining solutions and AI computational infrastructure. The company operates globally with data centers in multiple countries and focuses on AI and blockchain technology.
Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
Act as a customer-facing SRE to ensure Kubernetes clusters remain healthy and stable.
Become a product expert in GPU Cluster service, serving as the last line of technical defense before escalation.
Together AI is a research-driven artificial intelligence company focused on open and transparent AI systems. The team has contributed to leading open-source research and aims to build the next generation AI infrastructure.
Build the technical product marketing function for the Provider business, creating collateral like white papers, reference architectures, and demo environments.
Directly support pipeline development by partnering with sales and solution architects through technical storytelling and proof-points.
Develop competitive intelligence and represent Mirantis as a credible technical voice on GPU infrastructure and sovereign AI.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It empowers platform engineering teams across any environment with a strong benefits plan and professional development.
Design and build a managed Slurm service on Kubernetes
Write clean, reliable, and maintainable Go code, developing scheduling and orchestration for GPU workloads
Build observability and automated remediation for GPU, node, network, and control-plane failures
Gcore provides infrastructure and software solutions for AI, cloud, network, and security, powering digital experiences worldwide. With over 550 professionals, they build and support the global digital ecosystem.
Deploy and scale MCP-based AI agents on Kubernetes for enterprise customers across the US-East and EMEA regions.
Lead complex technical engagements, build reusable deployment patterns, and mentor engineers on the team.
Shape product roadmap by feeding back field insights from regulated industries and defining regional engagement standards.
Stacklok builds the control plane for enterprise AI agents, enabling organizations to run, govern, and secure them on Kubernetes and private cloud. Founded by two Kubernetes creators, the company is already adopted by leading tech and regulated industries, fostering a collaborative, AI-maximalist culture with deep open-source roots.
Build and operate the control plane for automated cluster deployment from bare metal to customer-ready.
Manage machine lifecycle including joining, wiping, verifying, and rejoining between tenants.
Operate Kubernetes, Postgres, and custom operators across the fleet, scaling from tens to thousands of nodes.
Andromeda provides scaled AI infrastructure for startups, managing compute across numerous capacity providers. The company operates tens of thousands of GPUs for 80+ customers and fosters an inclusive environment.
Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.
Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.
Design and maintain CI/CD and MLOps pipelines for AI and software applications, ensuring seamless deployment and automation.
Build and scale cloud-native infrastructure using Kubernetes, Docker, and GPU clusters to support high-performance AI workloads.
Champion Infrastructure as Code and observability practices to ensure high availability, security, and compliance across multi-cloud environments.
Bitdeer is a world-leading technology company providing AI and Bitcoin mining infrastructure. Headquartered in Singapore, the company has a global presence with data centers in multiple countries and a culture focused on innovation and reliability.
Design and build control-plane services and drivers for storage integration with Kubernetes-based AI workloads.
Write production-quality Go code with strong testing and operational rigor.
Deliver storage integration for k0s-based Kubernetes via Cluster API and K0rdent topologies.
Mirantis is a Kubernetes-native AI infrastructure company that builds scalable, secure infrastructure for AI and data-intensive applications. It is committed to open standards and freedom from lock-in, empowering platform engineering teams.
Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms.
Own the storage layer where Kubernetes meets bare metal, tuning NFS data paths for high-throughput workloads.
Automate storage provisioning with infrastructure-as-code and GitOps, ensuring observability and reliability.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams across hybrid, edge, and sovereign environments, fostering a culture of open-source innovation and collaboration among passionate, talented colleagues.
Design and build production AI agent systems for diagnosing and remediating infrastructure issues in large-scale GPU environments.
Develop distributed services, orchestration frameworks, knowledge graphs, and retrieval systems to power AI agents.
Own services end-to-end from architecture through production, collaborating with infrastructure and engineering teams.
The company builds AI agents that operate and automate large-scale GPU infrastructure. The engineering team is highly collaborative and remote, fostering ownership and autonomy.
Lead and build the SE / Solutions Architect team, defining the pre-sales operating model as the org scales.
Own the technical win in large, complex deals, architecting solutions across compute, networking, storage, and orchestration.
Be the technical voice of the customer internally, feeding structured product requirements back to product and platform teams.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The company is committed to open standards and freedom from lock-in, empowering platform engineering teams to deliver composable, production-ready developer platforms across any environment.
Own the technical path from customer interest to working deployment, integrating the platform into production AI environments.
Build and operate AI/MLOps pipelines, debug complex environments, and create prototypes and demos.
Translate customer needs into product improvements, partnering with Sales, Product, and Engineering.
Neuromorphic Labs is a Seed-stage AI startup building a trust layer for production AI, making security, governance, and control intrinsic to every model and deployment. Backed by top-tier VCs, the team is small and fast-paced, emphasizing ownership, high standards, and collaboration with exceptional builders.
Work directly with customers to understand technical requirements and build production solutions.
Develop integrations, APIs, and internal tooling to accelerate platform adoption.
Own projects from technical discovery through deployment and iteration.
A fast-growing technology company building advanced AI-powered infrastructure and products for enterprise customers. The company operates in a fast-paced, ambiguous environment and values high ownership and customer focus.
Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.
Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.