Build and deploy production-grade LLM inference systems from scratch, owning the pipeline from query to response.
Optimize inference workloads for latency, throughput, and cost using tools like vLLM, SGLang, and TensorRT-LLM.
Collaborate with the CTO and Product to define the technical roadmap for inference infrastructure as the organization scales.
The company is an open-source-oriented startup building production-grade inference infrastructure for large language models. It is a remote-first, globally distributed team that values engineering ownership, speed, and customer impact.
Deploy LLMs into production across GPU infrastructure, owning the full pipeline from customer query to served response.
Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
Apply quantization, batching, caching, and routing to optimize latency and cost at scale.
vCluster Labs is a venture-backed tech startup pioneering Kubernetes virtualization for the AI era, enabling AI Cloud providers and AI factories to operate GPU infrastructure with hyperscaler-like experiences. We raised over $30M from top-tier VCs like Khosla Ventures, are in a hyper-growth phase, and maintain a remote-first, distributed global team with headquarters in San Francisco.
Lead in-depth technical discovery with engineering teams and customer stakeholders to understand AI inference requirements.
Translate customer objectives into production-ready architectures and define PoC success criteria.
Identify recurring workload patterns and communicate insights to Product and Engineering for platform evolution.
They are a partner company focused on AI infrastructure and performance-sensitive AI inference workloads. They have an international, engineering-led team solving complex challenges at the forefront of AI.
Develop and maintain RESTful APIs for managing GPU clusters, virtual machines, and dedicated servers.
Utilize modern technologies including FastAPI, Python, OpenStack, Kubernetes, and more.
Participate in lifecycle management of compute products focusing on scalability and customer needs.
We are a global provider of infrastructure and software solutions for AI, cloud, network, and security, powering everything from real-time communication to enterprise AI. With over 550 professionals, we build the infrastructure and software that supports the entire digital ecosystem.
Own the cost and performance of the inference stack, improving throughput and latency without compromising reliability.
Optimize through KV-cache management, continuous batching, speculative decoding, and quantization.
Work within serving engines like vLLM, SGLang, and TensorRT-LLM, profiling performance down to kernel level.
Adaption builds efficient AI that evolves in real-time, making intelligence flexible, personalized, and accessible to everyone. They focus on talent density, bringing together driven individuals to push the boundaries of continual adaptation.
Build and scale ML infrastructure and tools to support AI-driven genetic medicine research.
Partner with scientists and engineers to translate research prototypes into robust systems.
Optimize model training and inference performance using advanced computing frameworks.
Dyno Therapeutics is building high-performance genetic technologies to transform patient lives. Their team unites world-class experts at the intersection of AI and genetic medicine.
Build reference architectures, benchmarks, and documentation that engineers trust.
Own the developer community, answering hard questions and setting the tone.
Write production-grade code for integrations and tooling that lower the barrier for new users.
Andromeda Cluster gives early-stage startups access to scaled AI infrastructure once reserved for hyperscalers. We are a unicorn with a small senior team, building the liquidity layer for global AI compute.
Develop and deploy machine learning and AI systems.
Work with LLMs, generative AI, and modern ML frameworks.
Optimize model performance, latency, and cost.
A fast-growing technology company building critical infrastructure that powers high-volume, real-time business operations across multiple systems and platforms. It is a collaborative, fast-moving environment where engineers have meaningful influence on architecture and product direction.
Design and optimize training and post-training pipelines for large language models.
Improve model quality through supervised fine-tuning, reinforcement learning, and evaluation.
Build PyTorch-based training infrastructure and optimize distributed training across multi-GPU environments.
Lightning AI builds an end-to-end platform for developing, training, and deploying AI systems, founded in 2019. They are a global company with offices in New York, San Francisco, Seattle, and London, backed by major venture capital firms, and foster a builder culture that values urgency, ownership, and open communication.
Own the infrastructure layer for AI workloads including inference serving, Kubernetes, and agent-sandboxing platforms.
Manage the serving tier for open-weight models, Kubernetes operators, and stateful data planes.
Oversee the sandbox runtime, control-plane services, and observability tooling.
AZX accelerates positive impact in critical industries through AI transformation, specializing in physics-informed ML and enterprise AI solutions for climate and sustainability. Founded in 2024, the company is a profitable public benefit corporation with a growing team working with category leaders in real estate, energy, logistics, and utilities.
Design and write high-performing scalable software for training models.
Develop new tools to support and accelerate research and LLM training.
Collaborate with engineering teams and scientific teams to implement experiments on cluster and data infrastructure.
Cohere is a security-first enterprise AI company building cutting-edge foundation AI models and end-to-end products for real-world business problems. The company is a global team of researchers, engineers, and designers passionate about AI, headquartered in Toronto with offices worldwide.
Design and build sandboxed evaluation environments for AI models to safely execute code and interact with tools.
Build backend services and infrastructure supporting large-scale AI and agentic evaluations.
Develop agent scaffolding, evaluation harnesses, and systems for provisioning isolated environments using Docker, Kubernetes, and cloud infrastructure.
10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. The company focuses on adversarial red teaming and model evaluations to help teams deploy AI systems safely.
Design and build LLM serving infrastructure on Kubernetes, including deployment, GPU scheduling, and model lifecycle management.
Package the platform for enterprise environments with Helm-based installs and support for restricted or offline networks.
Integrate the serving layer with API gateway, identity, and metering services, and build observability for GPU inference in production.
Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build and operate scalable infrastructure for AI and data-intensive applications. It is part of an IREN company and empowers platform engineering teams with open-source innovation and deep expertise in Kubernetes orchestration.
Design, train, and improve large-scale machine learning models for recommendation and personalization, leveraging modern deep learning architectures.
Own end-to-end delivery of major ML system components, from problem framing to production rollout, with cross-functional partners.
Optimize distributed training, model efficiency, and online inference performance while ensuring low-latency, high-throughput production systems.
Reddit is a community of communities built on shared interests, passion, and trust, hosting authentic conversations. With over 100,000 active communities and about 130 million daily active users, Reddit is one of the internet's largest sources of information.
Build the environments and verifiers our models train against
Own the synthetic data pipeline, from generation through quality gates
Ship models into production and keep improving them on real usage
LiveKit builds the infrastructure layer for the voice-driven era of computing, enabling developers to build, deploy, and scale voice AI applications. Founded in 2021, they power voice AI for major companies and have a small, senior team that values craft and creativity.
Design, build, and own shared infrastructure for Interpretability research environments, data systems, and compute tooling.
Lead cross-team efforts with agentic engineering, security, compute, and storage platform teams.
Discover and resolve major organization-wide developer experience issues and help take interpretability methods from research code to dependable audit pipelines.
Anthropic is an AI safety company focused on building reliable, interpretable, and steerable AI systems. They are a quickly growing team of researchers, engineers, policy experts, and business leaders committed to beneficial AI.
Design, develop, and deploy AI/ML and GenAI solutions from concept through production.
Build and optimize models for NLP, forecasting, classification, and anomaly detection.
Collaborate with stakeholders to translate business needs into scalable AI systems.
Our partner company builds and deploys high-impact AI and machine learning solutions across various industries. They offer a remote-first culture and emphasize collaboration and technical excellence.
Design, develop, and maintain automated diagnostic, validation, and remediation frameworks for production GPU hardware using Python and infrastructure automation tools.
Engineer and support Python-based agents, APIs, and Ansible automation for hardware provisioning, telemetry, and health monitoring.
Analyze workload performance, thermals, and diagnostic output to identify hardware issues and improve validation methodologies.
Vultr is on a mission to make high-performance cloud infrastructure easy to use, affordable, and locally accessible for enterprises and AI innovators worldwide. We are the world's largest privately-held cloud infrastructure company, self-funded for over a decade, with 33 global data centers and hundreds of thousands of active customers.
Develop resilient, high-scale, low-latency distributed microservices running on hundreds of cloud nodes.
Solve complex scaling challenges for ML inference on vast geospatial data.
Collaborate with ML and GIS engineers to maintain cutting-edge GIS software products.
Blackshark.ai provides a scalable AI platform for geospatial insights, enabling custom model training on global features with real-time accuracy. Founded in 2020, they have secured $35 million in series A funding from top US and European investors and work with customers in intelligence, government, and enterprise sectors.
You will work directly within client environments, translating real operational needs into production-ready AI systems rather than proof-of-concept solutions.
The role spans technical discovery, architecture, software development, deployment, optimization, and ongoing delivery ownership.
You will design and build RAG systems, agentic AI workflows, backend services, and cloud-native architectures that operate at enterprise scale.
Jobgether uses an AI-powered matching process to connect candidates with roles at partner companies. They operate remotely and focus on efficient, fair hiring through automated shortlisting, with a commitment to data privacy and human oversight in final decisions.