Similar Jobs

See all

Model Strategy & Selection:

  • Evaluate and select models from major and emerging providers using rigorous benchmarks and LLM-as-judge frameworks.
  • Apply statistical analysis to make model decisions defensible.

AI Evaluation & Development:

  • Design and maintain offline evaluation sets, judge calibration, and regression benchmarks.
  • Build classification and fine-tuned models for routing, categorization, and detection.
  • Develop agentic AI products and prompt/context engineering strategies.

Technical Leadership & Research:

  • Partner with data engineering, software engineering, and product teams.
  • Investigate model hallucinations, drift, and prompt sensitivity from first principles.
  • Maintain high engineering standards for code quality, testing, and documentation.

Jobgether

A production agentic AI platform that builds and deploys advanced machine learning systems. The team is collaborative and values innovation, offering a remote work environment with high autonomy.

Apply for This Position