Innodata is a global data engineering company that enables responsible AI advancement by providing data, evaluation frameworks, and human expertise. With over 36 years of experience, the company delivers high-quality data and outcomes for Generative AI builders.
Lead technical discovery with foundation model labs, frontier AI teams, and large enterprises to understand model objectives and constraints.
Design end-to-end solutions across the post-training stack including SFT data curation, RLHF/DPO pipelines, custom benchmarks, and LLM-as-judge systems.
Author technical proposals, run workshops and POCs, and serve as ongoing technical advisor during delivery.
Design and own the datasets and evaluation specifications for financial-domain LLMs, vision-language models, and AI agents, focusing on unstructured and multimodal financial data.
Translate customer goals into concrete dataset specifications, taxonomies, rubrics, and acceptance criteria, ensuring domain validity and statistical defensibility.
Develop evaluation methodology beyond surface accuracy, covering numerical consistency, hallucination rates, refusal appropriateness, and fairness across customer segments.