Similar Jobs

See all

Responsibilities:

  • Deploying models to production: Take LLMs and put them into production across one or more machines on GPU infrastructure.
  • Serving frameworks: Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
  • Optimizing for scale: Apply quantization, batching, caching, and routing to keep latency and cost in check.

Requirements:

  • Production LLM serving experience with vLLM, SGLang, or TensorRT-LLM.
  • Hands-on inference optimization with quantization, batching, caching, and routing.
  • Strong engineering skills in Python or Golang with real production code.

Culture & Values:

  • Make it Happen: Relentless bias for action and grit to push through obstacles.
  • Own the Outcome: Responsibility ends when value is delivered, connecting daily actions to broader success.
  • Create Wow: Measure success by the experience generated for customers and team.

vCluster Labs

vCluster Labs is a venture-backed tech startup pioneering Kubernetes virtualization for the AI era, enabling AI Cloud providers and AI factories to operate GPU infrastructure with hyperscaler-like experiences. We raised over $30M from top-tier VCs like Khosla Ventures, are in a hyper-growth phase, and maintain a remote-first, distributed global team with headquarters in San Francisco.

Apply for This Position