Staff Research Engineer, Model Efficiency

Cohere

Remote regions

Global

Benefits

6w PTO 26w maternity 26w paternity

Similar Jobs

See all

Role Overview:

  • Develop, prototype, and deploy techniques to improve LLM inference efficiency.
  • Explore and ship breakthroughs across model architecture, decoding, and hardware co-design.
  • Optimize performance without compromising model quality.

What We're Looking For:

  • PhD in Machine Learning or related field with deep understanding of LLM architecture.
  • Significant experience with model efficiency techniques and strong software engineering skills.
  • Publications at top-tier conferences and passion for mentoring.

Why Join Us:

  • Full health and dental benefits, plus mental health budget.
  • 6 weeks of paid vacation and 100% parental leave top-up for up to 6 months.
  • Remote-friendly environment with home office stipend and learning budget.

Cohere

Cohere is a security-first enterprise AI company building cutting-edge foundation models and end-to-end products for real-world business problems. We are a global team of researchers, engineers, and designers passionate about our craft, with offices in Toronto, San Francisco, London, New York City, Montreal, Seoul, and Paris.

Apply for This Position