Applied Research Scientist, LLM Evaluation & Post-Training

External

Innodata · Remote

Full-timeRemote2d ago

DocumentationGenerative AILeadershipLLMsMachine LearningPython

Cover Letter Connect

Prepare for this interview

Elite

AI-generated questions, company research, and talking points tailored to this role

Responsibilities

This is a highly collaborative role that sits at the intersection of research, engineering, and language/data operations. Additional responsibilities include (but are not limited to):
Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement
Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes
Develop and validate evaluation frameworks for LLM and multimodal systems, including:
benchmark/task design
scoring methods
judge/model-assisted evaluation
human evaluation protocols
robustness/stress testing
Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations
Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs
Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign
Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines
Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs
Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations
Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets
Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations
Contribute to thought leadership and best practices in LLM evaluation, post-training, and GenAI quality measurement
You'll Thrive in This Role If You Have:
MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field (PhD strongly preferred)
5+ years of relevant experience in applied research / research science in ML/AI, with substantial work in LLMs or foundation models
Demonstrated experience with LLM evaluation, benchmarking, alignment, post-training, or model quality research
Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems
Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization)
Experience working with modern ML tooling/frameworks (e.g., PyTorch, Hu

Additional Information

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers. Scope of the Role: Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement. This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains. The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.

Your Match

How well this role fits your profile.

Company Intel

What employees say

Worked at Innodata Inc.? Share your experience

Interested in this role?

Apply on the company's website.

Cover Letter Connect