Similar Jobs

See all

Responsibilities:

  • Optimize ML inference systems to improve latency, throughput, scalability, and cost.
  • Profile and identify bottlenecks across GPU and CPU inference pipelines.
  • Implement advanced optimization techniques such as quantization, KV-cache optimization, and speculative decoding.

Requirements:

  • Strong experience in ML inference optimization and high-performance systems.
  • Deep understanding of neural network architectures, attention mechanisms, and compute graphs.
  • Hands-on experience with PyTorch, CUDA, and inference frameworks like TensorRT or vLLM.

Benefits:

  • Competitive compensation package with meaningful equity participation.
  • Opportunity to work on performance-critical AI systems with direct product impact.
  • Flexible remote work environment and a culture that values technical excellence.

Partner Company

The company is an AI-focused organization that develops advanced machine learning systems for production environments. It values technical excellence and experimentation, offering a flexible remote work environment.

Apply for This Position