Source Job

Global

  • Own the cost and performance of the inference stack, improving throughput and latency without compromising reliability.
  • Optimize through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Work within serving engines like vLLM, SGLang, and TensorRT-LLM, profiling performance down to kernel level.

Python C++ Rust CUDA VLLM

7 jobs similar to Inference Performance Engineer

Jobs ranked by similarity.

Global

  • Deploy LLMs into production across GPU infrastructure, owning the full pipeline from customer query to served response.
  • Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
  • Apply quantization, batching, caching, and routing to optimize latency and cost at scale.

vCluster Labs is a venture-backed tech startup pioneering Kubernetes virtualization for the AI era, enabling AI Cloud providers and AI factories to operate GPU infrastructure with hyperscaler-like experiences. We raised over $30M from top-tier VCs like Khosla Ventures, are in a hyper-growth phase, and maintain a remote-first, distributed global team with headquarters in San Francisco.

Brazil

  • Build and deploy production-grade LLM inference systems from scratch, owning the pipeline from query to response.
  • Optimize inference workloads for latency, throughput, and cost using tools like vLLM, SGLang, and TensorRT-LLM.
  • Collaborate with the CTO and Product to define the technical roadmap for inference infrastructure as the organization scales.

The company is an open-source-oriented startup building production-grade inference infrastructure for large language models. It is a remote-first, globally distributed team that values engineering ownership, speed, and customer impact.

Global

  • Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium using CUDA, Triton, ROCm/HIP, or Neuron SDK.
  • Profile and improve inference performance in vLLM, SGLang, and custom runtimes through kernel fusion, scheduling, and memory optimizations.
  • Ship code upstream to open-source AI infrastructure projects with tests and documentation, working on a well-scoped project from design to production.

Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. They are a remote-first team with a focus on high-performance AI computing and offer a flexible, collaborative work environment.

$165,000–$330,000/yr
US Unlimited PTO

  • Partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten's platform.
  • Own the journey from initial exploration to production deployment, translating ambiguous goals into reliable services.
  • Work across product, software development, performance engineering, and customer-facing implementations.

Baseten powers mission-critical inference for dynamic AI companies like Cursor and Notion. They are rapidly growing, recently raised a $1.5B Series F, and foster a collaborative, forward-thinking culture.

Global 6w PTO 26w maternity 26w paternity

  • Develop, prototype, and deploy techniques to improve LLM inference efficiency in production.
  • Explore and ship breakthroughs across model architecture, decoding, and software/hardware co-design.
  • Optimize performance without compromising model quality.

Cohere is a security-first enterprise AI company building cutting-edge foundation models and end-to-end products for real-world business problems. We are a global team of researchers, engineers, and designers passionate about our craft, with offices in Toronto, San Francisco, London, New York City, Montreal, Seoul, and Paris.

North America

  • Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
  • Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
  • Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.

Europe

  • Lead and mentor a small team of engineers focused on AI platform and product engineering.
  • Drive architectural direction for AI product with emphasis on inference performance and cost efficiency.
  • Contribute hands-on to design and optimization of core AI systems including model serving and inference pipelines.

They operate a cloud and infrastructure platform. The team is small and high-leverage, focusing on engineering excellence and ownership.