Source Job

Global 6w PTO 26w maternity 26w paternity

  • Develop, prototype, and deploy techniques to improve LLM inference efficiency in production.
  • Explore and ship breakthroughs across model architecture, decoding, and software/hardware co-design.
  • Optimize performance without compromising model quality.

Machine Learning Large Language Models Software Engineering

20 jobs similar to Staff Research Engineer, Model Efficiency

Jobs ranked by similarity.

US

  • Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
  • Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability.
  • Build performance tooling, optimization playbooks, and efficiency primitives that benefit multiple teams.

Reddit is a community of communities built on shared interests and authentic conversations. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit has a flexible workforce and values collaboration.

Global 6w PTO 26w maternity 26w paternity

  • Conduct cutting-edge machine learning research, building and training large language models.
  • Focus on research projects aimed at expanding the frontier of knowledge in language modelling and associated areas.
  • Disseminate your research results through publications, datasets, and code.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation AI models and end-to-end products for enterprise AI systems. We are a global team of researchers, engineers, and designers passionate about our craft, with offices in Toronto, San Francisco, London, and more.

US

  • Fine-tune and optimize Large Language Models for healthcare and life science applications.
  • Develop and enhance RAG pipelines to improve AI-powered information retrieval.
  • Collaborate with cross-functional teams to deploy deep learning solutions for complex healthcare challenges.

The company develops AI-powered healthcare solutions using Large Language Models. They foster a highly collaborative and innovation-driven environment with a diverse international team.

Canada

  • Optimize machine learning inference systems for latency, throughput, and cost-efficiency.
  • Profile and troubleshoot GPU/CPU bottlenecks, implement advanced techniques like quantization and speculative decoding.
  • Collaborate with research and engineering teams to productionize new models and improve inference infrastructure.

The company is an AI-focused organization that develops advanced machine learning systems for production environments. It values technical excellence and experimentation, offering a flexible remote work environment.

US

  • Own end-to-end ML system execution including data pipelines, training workflows, evaluation systems, inference architecture, and deployment.
  • Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation.
  • Architect scalable inference systems, balance latency, cost, and reliability, and deploy production-grade ML solutions.

Gina's Tech Jobs is a recruiting and staffing company that helps firms hire technical talent. They are a small agency focused on IT roles, fostering a high-trust, collaborative environment.

US

  • Own end-to-end Machine Learning (ML) system execution including data pipelines, training, and deployment.
  • Fine-tune and adapt models using state-of-the-art methods like LoRA and DPO.
  • Architect scalable inference systems and collaborate closely with application engineering.

This company develops advanced production-grade machine learning systems. The team is small and high-trust, with a culture of ownership and pragmatism.

France

  • Design and implement advanced knowledge distillation pipelines, including teacher-student approaches and multi-teacher architectures.
  • Run large-scale machine learning experiments to optimize model quality, latency, efficiency, and cost trade-offs.
  • Collaborate with research teams to transform emerging distillation techniques into reliable production-ready implementations.

Our partner is an innovative company focused on advancing the efficiency and scalability of next-generation machine learning systems. They offer a remote-friendly work environment with an async-first culture and a small, senior team combining research expertise and engineering excellence.

US Unlimited PTO

  • Lead the development and optimization of Large Language Models and Mixture of Experts models.
  • Collaborate with cross-functional teams to integrate ML models into our platform and conduct cutting-edge research in machine learning.
  • Mentor junior engineers and contribute to the team’s knowledge sharing and best practices.

webAI is an end-to-end private AI platform that enables enterprises and governments to bring AI to their data, powering specialized intelligence trained on their own knowledge. The company is a dynamic, fast-growing team fostering an exciting and growth-oriented work culture, committed to truth, ownership, tenacity, and humility.

United States

  • Architect and build large-scale ML systems spanning data, training, evaluation, inference, and deployment.
  • Implement evaluation pipelines covering performance, robustness, safety, and bias.
  • Own production deployment including GPU optimization, memory efficiency, latency reduction, and scaling policies.

Global 6w PTO 26w maternity 26w paternity

  • Manage the program portfolio covering inference, efficiency, serving, and endpoints to scale Cohere's infrastructure.
  • Lead cross-functional coordination with Modeling and customer-facing teams for end-to-end execution.
  • Identify pain points, establish processes, and improve engineering best practices.

Cohere is the leading security-first enterprise AI company building cutting-edge foundation AI models and end-to-end products. We are a global team of researchers, engineers, and designers passionate about our craft, with offices across North America and Europe.

$107,360–$152,900/yr
Canada United States

  • Build and iterate on consumer-facing AI features powered by large language models (LLMs) and generative AI systems
  • Collaborate with engineers across the AI stack including prompt engineering and agentic workflow optimization
  • Run structured experiments and monitor production AI systems to optimize latency, cost, and scalability

Quora is a global knowledge sharing platform with over 300M monthly unique visitors, connecting people to share insights and learn. Poe provides a platform for users to chat and build with AI language models. They are a remote-first company with a culture rooted in transparency and experimentation.

$190,000–$230,000/yr
US

  • Lead the design and development of our production inference platform, defining the technical roadmap for inference infrastructure, model serving, and runtime optimization.
  • Build and operate scalable, cost-effective systems for serving large language models in production, optimizing latency, throughput, GPU utilization, and memory efficiency.
  • Partner with ML engineers to productionize new models and inference techniques, establish benchmarking methodologies, and make key architectural decisions.

Syllo is on a mission to transform litigation with a unified platform that enables lawyers to safely harness AI. Since going to market, they have gained diverse enterprise customers including big law firms and corporations, and are quickly expanding.

$165,000–$330,000/yr
US Unlimited PTO

  • Partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten's platform.
  • Own the journey from initial exploration to production deployment, translating ambiguous goals into reliable services.
  • Work across product, software development, performance engineering, and customer-facing implementations.

Baseten powers mission-critical inference for dynamic AI companies like Cursor and Notion. They are rapidly growing, recently raised a $1.5B Series F, and foster a collaborative, forward-thinking culture.

Global

  • Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium using CUDA, Triton, ROCm/HIP, or Neuron SDK.
  • Profile and improve inference performance in vLLM, SGLang, and custom runtimes through kernel fusion, scheduling, and memory optimizations.
  • Ship code upstream to open-source AI infrastructure projects with tests and documentation, working on a well-scoped project from design to production.

Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. They are a remote-first team with a focus on high-performance AI computing and offer a flexible, collaborative work environment.

Switzerland

  • Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
  • Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
  • Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

$138,500–$225,500/yr
US 16w maternity 16w paternity

  • Design, train, and ship ML systems for governance and security like anomaly detection and trust scoring.
  • Build data pipelines, model serving, evaluation frameworks, and feedback loops.
  • Set technical direction, own architecture, and help recruit and mentor as the team grows.

Docker provides developer tooling trusted by over 20 million monthly users and billions of container pulls. They are a globally distributed, remote-first team building tools for software delivery.

Global Unlimited PTO

  • Define the multi-quarter vision and strategy across Machine Learning & Recall and Query teams.
  • Drive query intelligence and advance recall systems to improve search-attributed conversion and revenue.
  • Build evaluation frameworks and optimize ML performance trade-offs for low-latency production.

Constructor provides product search and discovery solutions for large retailers, serving billions of requests weekly. The company is a remote-first, growing team of technologists that values empathy, openness, and continuous improvement.

Global

  • Design and scale production ML systems for LLM-based applications.
  • Build training and evaluation pipelines for continuous model improvement.
  • Fine-tune foundation models using modern adaptation techniques such as LoRA, QLoRA, SFT and DPO.

A1 is a new AI venture building the next generation of AI-native productivity applications, starting with an email agent that uses autonomous AI. Backed by an initial $100M investment, the company is a small, high-talent founding engineering team focused on solving challenging AI infrastructure problems.

$216,700–$303,400/yr
US

  • Design, develop, and deploy ML models, including large language models, for various NLP tasks.
  • Collaborate with cross-functional teams to gather requirements, define architectures, and iterate on model development.
  • Stay up-to-date with latest research and contribute to best practices for responsible ML development.

Reddit is a community of communities built on shared interests, passion, and trust, hosting the most open and authentic conversations on the internet. With over 100,000 active communities and approximately 126 million daily active users, it is one of the internet's largest sources of information, fostering a culture of authenticity and community.

$250,000–$265,000/yr
US Unlimited PTO 16w maternity 8w paternity

  • Design and build core architecture for AI systems spanning multiple product areas, ensuring coherence and reuse.
  • Set the technical bar for the org in evaluation frameworks, testing philosophy, CI/CD for model behavior, and shared infrastructure.
  • Take direct ownership of the hardest, most ambiguous technical problems and build solutions end-to-end.

Twin Health empowers people to prevent and reverse chronic metabolic diseases using AI Digital Twin technology. The company is rapidly scaling across the U.S. and globally with over $100 million raised, and has been recognized as an innovator and top workplace.