Deploy LLMs into production across GPU infrastructure, owning the full pipeline from customer query to served response.
Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
Apply quantization, batching, caching, and routing to optimize latency and cost at scale.
vCluster Labs is a venture-backed tech startup pioneering Kubernetes virtualization for the AI era, enabling AI Cloud providers and AI factories to operate GPU infrastructure with hyperscaler-like experiences. We raised over $30M from top-tier VCs like Khosla Ventures, are in a hyper-growth phase, and maintain a remote-first, distributed global team with headquarters in San Francisco.
Build and deploy production-grade LLM inference systems from scratch, owning the pipeline from query to response.
Optimize inference workloads for latency, throughput, and cost using tools like vLLM, SGLang, and TensorRT-LLM.
Collaborate with the CTO and Product to define the technical roadmap for inference infrastructure as the organization scales.
The company is an open-source-oriented startup building production-grade inference infrastructure for large language models. It is a remote-first, globally distributed team that values engineering ownership, speed, and customer impact.
Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium using CUDA, Triton, ROCm/HIP, or Neuron SDK.
Profile and improve inference performance in vLLM, SGLang, and custom runtimes through kernel fusion, scheduling, and memory optimizations.
Ship code upstream to open-source AI infrastructure projects with tests and documentation, working on a well-scoped project from design to production.
Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. They are a remote-first team with a focus on high-performance AI computing and offer a flexible, collaborative work environment.
Partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten's platform.
Own the journey from initial exploration to production deployment, translating ambiguous goals into reliable services.
Work across product, software development, performance engineering, and customer-facing implementations.
Baseten powers mission-critical inference for dynamic AI companies like Cursor and Notion. They are rapidly growing, recently raised a $1.5B Series F, and foster a collaborative, forward-thinking culture.
Develop, prototype, and deploy techniques to improve LLM inference efficiency in production.
Explore and ship breakthroughs across model architecture, decoding, and software/hardware co-design.
Optimize performance without compromising model quality.
Cohere is a security-first enterprise AI company building cutting-edge foundation models and end-to-end products for real-world business problems. We are a global team of researchers, engineers, and designers passionate about our craft, with offices in Toronto, San Francisco, London, New York City, Montreal, Seoul, and Paris.
Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.