Skip to content
Synthium Labs
Synthium Labs

Open role

Data Evaluation Associate

Remote — global (any timezone)Part-time / Full-time (flexible)Selective — by capability assessmentCompetitive, discussed in the first call

The hard part of model evaluation isn't producing volume — it's designing tasks that expose real failure modes, applying a quality bar consistently, and trusting the data that comes out the other end.

As a Data Evaluation Associate, you will run evaluation workflows, build evaluation datasets, develop rubrics and perform quality review across work produced by the contributor network. Your judgment decides what ships to a model team.

You will pass a capability-specific assessment and project calibration before working on production data. This role is a fit for people who think rigorously about ambiguous problems, hold a consistent standard without supervision, and are genuinely curious about how modern AI systems work.

Responsibilities

  • Run structured evaluation workflows and apply rubrics consistently
  • Build evaluation datasets with careful attention to quality
  • Perform quality review, adjudication and rework loops on contributor output
  • Identify model failure patterns and document them systematically
  • Verify expert-generated datasets against domain-specific standards
  • Write clear structured analysis and technical summaries
  • Support research teams with structured, reliable data
  • Escalate edge cases and drive process improvements

Requirements

  • Strong analytical ability and rigorous thinking
  • Strong written and verbal communication
  • Ability to reason carefully about ambiguous problems and edge cases
  • Strong curiosity about AI and how modern models work
  • Comfort working with technical material and data
  • Programming experience is preferred (any language)

Nice to have

  • Python
  • PyTorch
  • Machine learning coursework
  • NLP or LLM experience
  • Statistics
  • Research experience (lab, thesis, or independent)
  • Competitive programming
  • Open-source contributions

What you'll learn

  • How evaluation methodology works in frontier AI development
  • How to design tasks and rubrics that reliably measure capability
  • How to run quality control and adjudication at production scale
  • The full loop from expert data to model post-training
  • Hands-on experience with LLMs, datasets and evaluation tooling

Ready to work on real AI problems?

Applications take about 15 minutes. If you're selected, we'll send a confirmation email and get back to you shortly.

Questions? hello@synthiumlabs.com