Own the cost and performance of the inference stack, improving throughput and latency without compromising reliability.
Optimize through KV-cache management, continuous batching, speculative decoding, and quantization.
Work within serving engines like vLLM, SGLang, and TensorRT-LLM, profiling performance down to kernel level.
Adaption builds efficient AI that evolves in real-time, making intelligence flexible, personalized, and accessible to everyone. They focus on talent density, bringing together driven individuals to push the boundaries of continual adaptation.
Build and deploy production-grade LLM inference systems from scratch, owning the pipeline from query to response.
Optimize inference workloads for latency, throughput, and cost using tools like vLLM, SGLang, and TensorRT-LLM.
Collaborate with the CTO and Product to define the technical roadmap for inference infrastructure as the organization scales.
The company is an open-source-oriented startup building production-grade inference infrastructure for large language models. It is a remote-first, globally distributed team that values engineering ownership, speed, and customer impact.
Deploy LLMs into production across GPU infrastructure, owning the full pipeline from customer query to served response.
Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
Apply quantization, batching, caching, and routing to optimize latency and cost at scale.
vCluster Labs is a venture-backed tech startup pioneering Kubernetes virtualization for the AI era, enabling AI Cloud providers and AI factories to operate GPU infrastructure with hyperscaler-like experiences. We raised over $30M from top-tier VCs like Khosla Ventures, are in a hyper-growth phase, and maintain a remote-first, distributed global team with headquarters in San Francisco.