Skip to content
Synthium Labs
Synthium Labs

Open role

AI Data & Evaluation Specialist

Remote — global (any timezone)Part-time / Full-time (flexible)Selective — by capability assessmentCompetitive, discussed in the first call

Frontier LLMs are shaped as much by human feedback as by compute. Every preference comparison, evaluation and correction becomes a signal that steers how models reason, follow instructions and behave under pressure.

As an AI Data & Evaluation Specialist, you will be part of the human intelligence layer that produces those signals. You will work with detailed rubrics, compare model outputs, identify factual and reasoning errors, and write feedback that makes models measurably better.

This is selective work: you will pass a capability-specific assessment and project calibration before touching production data, and your output is continuously scored against project quality bars. The work is demanding, and it's paid accordingly.

Responsibilities

  • Evaluate LLM responses against detailed, task-specific rubrics
  • Compare model outputs and rank them by quality and correctness
  • Write high-quality, well-structured feedback on model behavior
  • Identify factual, logical and reasoning errors in model outputs
  • Follow complex annotation guidelines with high consistency
  • Create examples and edge cases used for model improvement
  • Participate in model evaluation projects across domains
  • Maintain strong attention to detail across long work sessions

Requirements

  • Strong written English with precise, unambiguous communication
  • Strong analytical and reasoning ability
  • Good attention to detail and consistency over sustained work
  • Ability to learn and apply detailed guidelines quickly
  • Comfortable working independently with asynchronous supervision
  • Basic familiarity with AI / LLMs (using ChatGPT, Claude or similar)

Nice to have

  • Python
  • Machine learning coursework or projects
  • Working knowledge of LLMs and prompt engineering
  • Mathematics, computer science or domain-specialist background
  • Experience writing technical documentation or structured reviews

What you'll learn

  • How RLHF and post-training pipelines actually work under the hood
  • How evaluation rubrics and annotation workflows are designed
  • How model behavior is measured, compared and improved
  • What high-quality training data looks like for frontier models
  • Working with a distributed, async team of evaluators

Ready to work on real AI problems?

Applications take about 15 minutes. If you're selected, we'll send a confirmation email and get back to you shortly.

Questions? hello@synthiumlabs.com