Inference Infrastructure Architect

Telnyx

Remote regions

China

Benefits

Similar Jobs

See all

About the Role:

  • Telnyx runs its own B300 GPU fleet and needs an architect to turn it into efficient inference throughput.
  • You will own the platform from bare metal to customer endpoint, with two mandates: operate the fleet efficiently and make the product run on it.

What You'll Build:

  • Serverless serving pools using vLLM/SGLang with advanced techniques like prefix caching and MoE parallelism.
  • The fleet layer with llm-d/NVIDIA Dynamo for KV-cache-aware routing and prefill/decode disaggregation.
  • Kubernetes on bare metal with GPU Operator and topology-aware scheduling, plus bare-metal lifecycle management.

What We Look For:

  • Deep experience running production LLM serving at scale, diagnosing bottlenecks and improving latency or cost.
  • Strong Kubernetes on GPU fleets, with operational command of vLLM or SGLang and system-level performance engineering.
  • Python/Go skills, plus real Linux, networking, and storage depth; community presence in CNCF/OpenInfra is a plus.

Telnyx

Telnyx is an industry leader building the future of global connectivity through a private, multi-cloud IP network and edge APIs. The company is financially stable and profitable, with a global team and a focus on innovation and continuous learning.

Apply for This Position