Source Job

US

  • Evaluate AI quality across the advisor stack, including pre-call briefs, in-call guidance, and post-call outputs.
  • Iterate inside ORA by refining prompts, updating knowledge base entries, and tweaking skills to close the loop on issues.
  • Surface trends and drive continuous improvement by tagging conversations, logging issues, and recommending prioritized improvements.

Quality Assurance AI Evaluation SQL Customer Success

20 jobs similar to Quality Analyst II

Jobs ranked by similarity.

US Unlimited PTO

  • Own the performance of customer-facing AI experiences across Ashley Digital's brands, combining customer insights, conversation design, analytics, and knowledge management.
  • Design, test, and continuously improve conversational experiences through prompts, response logic, and system instructions to ensure natural, helpful interactions.
  • Analyze AI performance using KPIs like CSAT, resolution time, and conversion to identify optimization opportunities and drive measurable business results.

Ashley Digital is the e-commerce engine behind Ashley Furniture, driving online retail for home furnishings across all platforms. They are a fast-moving, highly collaborative team with a remote-first culture, combining the energy of a tech company with the stability of an industry leader.

$120,000–$170,000/yr
US

  • Design, build, and maintain automated AI evaluation pipelines for production LLM applications.
  • Develop prompt engineering strategies and evaluate model performance using quantitative methods.
  • Analyze production AI behavior with Python, SQL, and statistical techniques to identify improvement opportunities.

GovWorx provides an AI-powered platform, CommsCoach, that supports 9-1-1 and emergency communications centers by automating quality assurance, training, and real-time call evaluation. The company is a growing technology team focused on public safety, collaborating across AI, engineering, product, and data science.

US Canada

  • Lead end-to-end development for the AI platform and customer experiences, from initial roadmap to rollout.
  • Partner with engineers to design prompt strategies, evaluation frameworks, and guardrails balancing latency, cost, and accuracy.
  • Serve as the technical translator between engineering and the broader organization, establishing AI best practices and platform standards.

Jerry.ai is building the first AI agent to manage all your physical assets, starting with car insurance and expanding into home, motorcycle, and more. The company has 5M+ customers, raised $240M+, and has been profitable since 2024 with a fully remote team and offices in Palo Alto, New York, Chicago, and Toronto.

AI Builder

Best10
Global

  • Build internal tools and coach-facing features on top of LLMs to turn unstructured conversational data into structured signals.
  • Design and maintain prompt chains, RAG pipelines, and agent workflows, rapidly prototyping from idea to demo in days.
  • Wire AI into existing systems, own full stack, and sit with coaches to build for their reality.

We deliver highly personalised diet and lifestyle coaching at scale through WhatsApp and our own technology platform. We partner with leading health insurers and employers, and our culture is about using AI to remove administrative work and enhance the personal coach-member relationship.

AI Builder

Best10

  • Build internal tools and coach-facing features on top of LLMs like Claude and OpenAI.
  • Turn unstructured conversational data into structured, reliable signals for daily prioritisation.
  • Own the full stack of what you ship: frontend, backend, model integration, and deployment.

Best10 delivers online, highly personalised diet and lifestyle coaching through WhatsApp and its own platform, working with medical insurers to reduce costs from high-risk members. The company fosters a builder culture where teams ship working AI solutions rapidly and take pride in systems that outlast them.

Europe

  • Own and extend the offline eval suite across AI products, including datasets, judges, and metrics.
  • Build and maintain online quality dashboards tracking resolution rate, CSAT, LLM-as-judge signals, and more.
  • Close the production feedback loop by mining failure patterns and translating data into product decisions.

Finom is a European tech startup headquartered in Amsterdam, developing an all-in-one financial B2B platform integrating banking, accounting, and invoicing. With over $346 million in total funding and a team of hundreds, they foster a start-up culture focused on innovation, swift implementation, and user impact.

$64,350–$113,500/yr
US Canada

  • Drive quality consistency and operational integrity: Review customer interactions to assess compliance, risk, and quality standards.
  • Deliver behavior-based quality insights: Provide actionable feedback on customer interactions to support coaching and improvement.
  • Leverage AI-assisted tools to uncover patterns and translate insights into operational improvements.

Mercury is a fintech company that provides banking services through partner banks. They are an equal opportunity employer and value diversity and belonging.

United States

  • Design and deliver high-quality AI evaluation data initiatives, from proposals through pilot execution and production readiness.
  • Recruit and manage subject-matter experts across technical domains, ensuring rigorous quality control frameworks.
  • Act as key interface with AI lab partners, converting pilots into scaled production engagements.

Jobgether uses AI-powered matching to connect candidates with roles quickly and fairly. They are a remote-first company that shares top-fitting candidates with hiring partners.

UK

  • You will design and implement AI agent deployments, crafting conversational flows and prompts using LLMs, NLU, and integration platforms.
  • You will build and configure integrations with enterprise systems like CRMs and ERPs, ensuring quality and performance.
  • You will collaborate with customers and internal teams to analyze requirements, prototype solutions, and document best practices.

Parloa empowers effortless customer conversations using agentic AI. With over one billion interactions handled for global brands like Booking.com and Allianz, Parloa fosters a collaborative culture focused on innovation and scale.

Global

  • Research data collection strategies and design high-impact data slices that uncover model failure modes.
  • Model annotator behavior and design experiments to optimize instruction clarity and reward signal reliability.
  • Develop metrics and frameworks for evaluating dataset quality, diversity, and impact on downstream model alignment.

Surge AI builds a platform that powers the most powerful AI models in partnership with companies like Anthropic, Google, Microsoft, and Meta. They are a profitable, bootstrapped company focused on human intelligence and data quality.

US

  • Own and optimize AI-powered support tools including chatbots, copilots, and automation workflows.
  • Analyze customer interactions to improve AI response accuracy and reduce manual workload.
  • Collaborate with Product and Engineering teams to design intelligent workflows and enhance customer experience.

The partner company is a technology-driven organization specializing in AI-powered support solutions for financial services. It operates as a remote-first, global team with a culture focused on innovation and collaboration.

Global

  • Build AI-powered product features integrated into real insurance workflows.
  • Optimize LLM-based interactions and develop prompts for production use cases.
  • Evaluate model performance and improve reliability through guardrails and validation.

US Unlimited PTO 16w maternity 16w paternity

  • Leverage and test AI-driven knowledge to ensure chatbot and internal Copilot draw from accurate content, identify performance gaps, and provide feedback loops for content updates.
  • Build and maintain internal knowledge resources, optimize chatbot workflows in Intercom, and use data to continuously improve AI response accuracy and resolution rates.
  • Stay close to customer challenges through a small volume of support tickets, deepen product expertise, and partner with cross-functional teams to align AI strategies.

Vanta helps businesses earn and prove trust through continuous security monitoring and compliance automation. Founded in 2018, Vanta has a kind and talented team and is the world's leading Trust Management Platform used by thousands of companies.

AI Developer

Aphex
Global 4w PTO

  • Build and implement AI features by selecting the right model and approach for each use case.
  • Evaluate and monitor AI feature performance in production to ensure accuracy and reliability.
  • Improve and iterate on prompts and implementations based on how features behave in the wild.

Aphex is a construction execution platform that replaces traditional spreadsheets with collaborative tools for delivery teams. They are a remote-first company with a growing engineering team in the Philippines, serving major contractors on multi-billion dollar projects.

$160,000–$240,000/yr
US

  • You will build AI-assisted tools, workflow automations, agents, prompts, and integrations to reduce manual effort and improve productivity.
  • You will partner with business stakeholders to understand high-friction workflows and deliver fit-for-purpose AI solutions.
  • You will implement engineering controls for data handling, access management, prompt safety, and output validation.

Shield AI is a venture-backed defense-tech company that develops intelligent systems, including Hivemind autonomy software and V-BAT and X-BAT aircraft, to protect service members and civilians. With offices across the U.S., Europe, the Middle East, and Asia-Pacific, the company's technology supports operations worldwide.

Global

  • Build AI-powered product features integrated into real insurance workflows.
  • Design and optimize LLM-based interactions for customer and internal systems.
  • Improve reliability of AI outputs through guardrails, fallback logic, and validation layers.

BJAK is Southeast Asia's largest digital insurance platform, using AI to simplify insurance and financial services for millions of users. The company values technical excellence, speed of execution, and practical decision-making, with a global engineering team that works closely across product, design, and AI teams.

$4–$5/hr
Global

  • Review AI-generated responses against source images and quality guidelines.
  • Identify issues like hallucinations, missing details, or policy violations.
  • Provide structured feedback to improve model performance and output quality.

Jobgether uses AI-powered matching to connect candidates with partner companies. They focus on efficient, objective hiring processes and operate as a platform for remote opportunities.

EMEA 5w PTO

  • Own problem spaces end to end: write specs, acceptance criteria, and own the architecture.
  • Build AI into the product, e.g., turning free-text email replies into bookable quotes.
  • Make AI trustworthy with structured outputs, evals, and confidence-gated human review.

Cargo.one operates an AI-native operating system for freight, serving 30,000+ users across 172 countries with customers like Lufthansa Cargo and Kuehne+Nagel, backed by Index and Bessemer. The culture is positive, diverse, hard-working, feedback-heavy, and playful.

Eastern Europe

  • Design, develop, and maintain automated test suites for web applications, APIs, and AI-powered features.
  • Evaluate LLM outputs for response accuracy, relevance, consistency, and hallucination risks using frameworks like Promptfoo.
  • Build end-to-end UI automation with Playwright and develop reusable Python scripts for API testing and LLM quality evaluation.

Nagarro is a digital product engineering company that builds products, services, and experiences across all devices and digital mediums. With over 18,000 experts in 39 countries, we have a dynamic and non-hierarchical work culture.

United States

  • Design and own the quality of clinical datasets for training and evaluating health AI models.
  • Collaborate with cross-functional teams to translate customer goals into dataset specifications and evaluation plans.
  • Ensure clinical realism, statistical defensibility, and compliance in health AI evaluation workflows.

Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence. With over 36 years of legacy, they deliver high-quality data and outstanding outcomes for AI builders and adopters.