Source Job

Global

  • Deploy LLMs into production across GPU infrastructure, owning the full pipeline from customer query to served response.
  • Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
  • Apply quantization, batching, caching, and routing to optimize latency and cost at scale.

Python Golang VLLM

20 jobs similar to Senior Inference Engineer

Jobs ranked by similarity.

Switzerland

  • Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
  • Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
  • Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

$165,000–$330,000/yr
US Unlimited PTO

  • Partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten's platform.
  • Own the journey from initial exploration to production deployment, translating ambiguous goals into reliable services.
  • Work across product, software development, performance engineering, and customer-facing implementations.

Baseten powers mission-critical inference for dynamic AI companies like Cursor and Notion. They are rapidly growing, recently raised a $1.5B Series F, and foster a collaborative, forward-thinking culture.

$200,000–$350,000/yr
US

  • Develop and deploy machine learning and AI systems.
  • Work with LLMs, generative AI, and modern ML frameworks.
  • Optimize model performance, latency, and cost.

A fast-growing technology company building critical infrastructure that powers high-volume, real-time business operations across multiple systems and platforms. It is a collaborative, fast-moving environment where engineers have meaningful influence on architecture and product direction.

Canada

  • Optimize machine learning inference systems for latency, throughput, and cost-efficiency.
  • Profile and troubleshoot GPU/CPU bottlenecks, implement advanced techniques like quantization and speculative decoding.
  • Collaborate with research and engineering teams to productionize new models and improve inference infrastructure.

The company is an AI-focused organization that develops advanced machine learning systems for production environments. It values technical excellence and experimentation, offering a flexible remote work environment.

  • Build and ship AI features end-to-end, from model to system to user experience.
  • Design and iterate on prompts, tools, memory, and agent workflows for real-world reliability.
  • Debug full-stack issues and optimize for latency, cost, and production performance.

A1 builds a proactive smart assistant for everyday users, bringing intelligence to conversations, errands, organizing, and workflows with minimal prompting. The team is small, world-class, and focuses on rapid iteration and shipping high-quality AI products.

Global

  • Design, build, and deploy production ML and LLM-based systems (RAG, agentic workflows, fine-tuning, embeddings) for enterprise clients.
  • Own technical delivery end-to-end: from architecture and prototyping to deployment, monitoring, and iteration.
  • Mentor and support other ML engineers on the team with code reviews, technical guidance, and knowledge sharing.

TensorOps is a boutique AI consultancy that bridges strategy and execution, designing and shipping production-grade AI systems for enterprise clients. We are a 100% remote team of 11+ people, partnering with unicorns and NASDAQ-listed companies, and have a culture of autonomy, open communication, and continuous learning.

Brazil

  • Lead the end-to-end lifecycle of language models and AI solutions, from research to production.
  • Navigate between closed and open-source ecosystems to maximize quality, optimize latency/cost, and ensure data governance.
  • Conduct applied research, experimentation, and data curation to create new AI capabilities and continuously improve existing ones.

Blip is a technology company that develops conversational AI and customer service platforms. The company fosters a culture of innovation and technical excellence, with a team of engineers and researchers working on cutting-edge AI solutions.

US

  • Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
  • Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability.
  • Build performance tooling, optimization playbooks, and efficiency primitives that benefit multiple teams.

Reddit is a community of communities built on shared interests and authentic conversations. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit has a flexible workforce and values collaboration.

Global

  • Own end-to-end delivery quality for major engagements, translating ambiguous client needs into practical execution plans.
  • Lead solution architecture and technical decision-making, making pragmatic tradeoffs between speed, quality, and client value.
  • Build and ship production AI/ML systems using Python, ML frameworks, and cloud-native infrastructure while mentoring other engineers.

Eliza is a technology services company and Advanced-tier OpenAI partner that helps organizations build and deploy AI solutions, from generative AI to predictive analytics. They are a collaborative, mission-driven team focused on real-world AI impact.

Europe

  • Lead and mentor a small team of engineers focused on AI platform and product engineering.
  • Drive architectural direction for AI product with emphasis on inference performance and cost efficiency.
  • Contribute hands-on to design and optimization of core AI systems including model serving and inference pipelines.

They operate a cloud and infrastructure platform. The team is small and high-leverage, focusing on engineering excellence and ownership.

Global 6w PTO 26w maternity 26w paternity

  • Design and write high-performing scalable software for training models.
  • Develop new tools to support and accelerate research and LLM training.
  • Collaborate with engineering teams and scientific teams to implement experiments on cluster and data infrastructure.

Cohere is a security-first enterprise AI company building cutting-edge foundation AI models and end-to-end products for real-world business problems. The company is a global team of researchers, engineers, and designers passionate about AI, headquartered in Toronto with offices worldwide.

Global 6w PTO 26w maternity 26w paternity

  • Develop, prototype, and deploy techniques to improve LLM inference efficiency in production.
  • Explore and ship breakthroughs across model architecture, decoding, and software/hardware co-design.
  • Optimize performance without compromising model quality.

Cohere is a security-first enterprise AI company building cutting-edge foundation models and end-to-end products for real-world business problems. We are a global team of researchers, engineers, and designers passionate about our craft, with offices in Toronto, San Francisco, London, New York City, Montreal, Seoul, and Paris.

Mexico

  • Design and develop AI-powered applications and enterprise solutions using Large Language Models.
  • Build scalable Python backend services and integrate AI with cloud platforms and databases.
  • Optimize AI model performance and collaborate with multidisciplinary teams.

The company is a partner organization that develops AI-powered solutions for enterprise clients. They have a diverse, international team and foster a culture of innovation and continuous learning.

Global

  • Design and scale production ML systems for LLM-based applications.
  • Build training and evaluation pipelines for continuous model improvement.
  • Fine-tune foundation models using modern adaptation techniques such as LoRA, QLoRA, SFT and DPO.

A1 is a new AI venture building the next generation of AI-native productivity applications, starting with an email agent that uses autonomous AI. Backed by an initial $100M investment, the company is a small, high-talent founding engineering team focused on solving challenging AI infrastructure problems.

Global

  • Design and maintain LLM-powered backend services using Python and FastAPI.
  • Implement retrieval-augmented generation (RAG) for structured and unstructured fleet data.
  • Optimize retrieval accuracy, latency, and hallucination rates through automated evaluation pipelines.

Datakrew revolutionizes EV fleet intelligence with IoT and AI solutions. They aim to serve one million EVs within 5 years and cultivate a mission-driven culture.

Global 6w PTO 26w maternity 26w paternity

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.

North America

  • Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
  • Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
  • Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.

UK Germany Netherlands Ireland Spain Poland Bulgaria Unlimited PTO

  • Build and own the model serving infrastructure, real-time inference, feature retrieval, and the latency budget that governs both.
  • Build the deployment path for data scientists to ship models, including bring-your-own-model support.
  • Own models in production: monitoring, drift detection, retraining, incident response, and the on-call rotation.

Sardine is the leading agentic risk platform for fighting financial crime. We are a remote-first company with hubs in the Bay Area, NYC, Austin, Toronto, and São Paulo, hiring talented individuals with extreme ownership and high growth orientation.

Global

  • Taking ML or LLM proof-of-concept to production for large enterprises.
  • Designing and hardening data and training pipelines for enterprise ML systems.
  • Building LLM and RAG systems with retrieval quality, evaluation and cost control.

Janea Systems (USA) is a dynamic team of the best & brightest software engineering specialists and solutions innovators from around the world. From kernel to cloud, we provide high-impact software development services to Fortune 500 companies.

Global

  • Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium using CUDA, Triton, ROCm/HIP, or Neuron SDK.
  • Profile and improve inference performance in vLLM, SGLang, and custom runtimes through kernel fusion, scheduling, and memory optimizations.
  • Ship code upstream to open-source AI infrastructure projects with tests and documentation, working on a well-scoped project from design to production.

Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. They are a remote-first team with a focus on high-performance AI computing and offer a flexible, collaborative work environment.