Open role
Data Evaluation Associate
The hard part of model evaluation isn't producing volume — it's designing tasks that expose real failure modes, applying a quality bar consistently, and trusting the data that comes out the other end.
As a Data Evaluation Associate, you will run evaluation workflows, build evaluation datasets, develop rubrics and perform quality review across work produced by the contributor network. Your judgment decides what ships to a model team.
You will pass a capability-specific assessment and project calibration before working on production data. This role is a fit for people who think rigorously about ambiguous problems, hold a consistent standard without supervision, and are genuinely curious about how modern AI systems work.
Responsibilities
- Run structured evaluation workflows and apply rubrics consistently
- Build evaluation datasets with careful attention to quality
- Perform quality review, adjudication and rework loops on contributor output
- Identify model failure patterns and document them systematically
- Verify expert-generated datasets against domain-specific standards
- Write clear structured analysis and technical summaries
- Support research teams with structured, reliable data
- Escalate edge cases and drive process improvements
Requirements
- Strong analytical ability and rigorous thinking
- Strong written and verbal communication
- Ability to reason carefully about ambiguous problems and edge cases
- Strong curiosity about AI and how modern models work
- Comfort working with technical material and data
- Programming experience is preferred (any language)
Nice to have
- Python
- PyTorch
- Machine learning coursework
- NLP or LLM experience
- Statistics
- Research experience (lab, thesis, or independent)
- Competitive programming
- Open-source contributions
What you'll learn
- How evaluation methodology works in frontier AI development
- How to design tasks and rubrics that reliably measure capability
- How to run quality control and adjudication at production scale
- The full loop from expert data to model post-training
- Hands-on experience with LLMs, datasets and evaluation tooling
Ready to work on real AI problems?
Applications take about 15 minutes. If you're selected, we'll send a confirmation email and get back to you shortly.
Questions? hello@synthiumlabs.com