Similar Jobs

See all

About the Role:

  • Senior Software Engineer focused on LLM evaluation and AI-assisted software engineering.
  • Collaborate with research and technical professionals on evaluation strategies.

Responsibilities:

  • Curate high-quality code examples and datasets for model training and benchmarking.
  • Evaluate AI-generated code for correctness, maintainability, and scalability.
  • Build agents and automated mechanisms to verify code quality across the SDLC.

Requirements:

  • 3+ years of professional software-engineering experience.
  • Strong full-stack development and software architecture skills.
  • Proficiency in Python, JavaScript, Java, C++, Rust, or related languages.

Benefits:

  • Fully remote, part-time independent contractor engagement.
  • Flexible workload from 10 to 40 hours per week.

Partner Company

This partner company specializes in evaluating large language models and improving AI systems through rigorous engineering benchmarks. It offers a remote, collaborative culture where engineers and researchers advance AI evaluation workflows together.

Apply for This Position