Similar Jobs

See all

Accountabilities:

  • Build and deploy production-grade LLM inference systems across one or more GPU machines, owning the complete pipeline from customer query to served response.
  • Design, implement, and operate model-serving infrastructure using technologies such as vLLM, SGLang, and TensorRT-LLM.
  • Optimize inference workloads for scale, balancing latency, throughput, reliability, and infrastructure costs.

Requirements:

  • Significant professional experience building and operating production software or infrastructure systems, with strong hands-on engineering capabilities.
  • Demonstrated experience deploying and serving large language models in production, ideally using vLLM, SGLang, TensorRT-LLM, or comparable inference frameworks.
  • Strong programming skills in Python or Golang, with a track record of writing and maintaining production-quality code.

Benefits:

  • Competitive compensation package including equity, health, dental, vision, and life insurance.
  • Flexible working schedule focused on outcomes rather than fixed hours, with high workplace flexibility supporting remote work.
  • Significant ownership over the architecture, implementation, and long-term roadmap of the inference platform, with direct collaboration with senior technical leadership and Product teams.

Undisclosed

The company is an open-source-oriented startup building production-grade inference infrastructure for large language models. It is a remote-first, globally distributed team that values engineering ownership, speed, and customer impact.

Apply for This Position