Avatar for Eros Innovation
Eros Innovation
Actively Hiring
Sovereign cultural AI infrastructure for creators, content and global markets

Video Evaluation, Verification & Data Engineer

  • Remote (
    Everywhere
    )
  • |5 years of exp
  • |Full Time
Posted: 1 day ago• Recruiter recently active
Hires remotely in
Everywhere
Remote Work Policy

Remote only

Company Location
Isle of Man
Visa Sponsorship

Not Available

RelocationNot Allowed

About the job

Eros is building sovereign cultural AI: systems designed to make advanced intelligence culturally relevant, rights-aware, governable, and commercially useful. We are bringing together a major film and media catalog, a new AI platform, and a founding technical team, with every training and evaluation asset subject to rights, consent, provenance, territory, and permitted-use controls.

We are hiring a senior evaluation and data engineer as a founding member of a new generative-video team. You will build the data controls, evaluation harness, blinded review process, verifier calibration, and evidence package used to determine whether model or conditioning changes produce real improvement.

This is not a general data-labeling role, a compliance-writing role, or a request for someone to score a few polished demos. We need a hands-on engineer who can make video-model results reproducible, comparable, auditable, and difficult to game.

What you will own

  • Create versioned dataset manifests with stable IDs, checksums, source lineage, split assignments, permitted-use fields, and exception tracking.
  • Define clean development, validation, and sealed holdout sets, then detect duplicates, contamination, identity leakage, and post-result test changes.
  • Build a frozen evaluation harness for identity consistency, appearance continuity, scene and action adherence, temporal stability, critical failures, latency, and cost.
  • Combine deterministic hard gates, automated scoring, and blinded human review without hiding disagreements between them.
  • Design reviewer instructions, adjudication rules, sampling plans, calibration cases, and inter-rater agreement checks.
  • Preserve immutable generation manifests, evaluator versions, scores, reviewer decisions, exceptions, and evidence links for every evaluated asset.
  • Calibrate verifier thresholds against human judgments and document false positives, false negatives, uncertainty, and known blind spots.
  • Produce decision-ready before-and-after evidence and complete technical handover materials, including code, schemas, runbooks, dashboards, and evaluation documentation.

Operating boundaries

  • Proprietary data and generated artifacts remain inside a controlled private environment.
  • Rights, consent, provenance, permitted use, retention, and revocation state must be represented in the data and evidence workflow.
  • The implementation team cannot change test cases, thresholds, routes, or evaluator versions after seeing results without creating a new declared experiment.
  • Missing evidence, rights uncertainty, holdout leakage, or incomplete manifests must cause a stop, quarantine, or explicit exception, not a silent pass.
  • You will build and operate the evidence machinery, but final acceptance remains with a designated internal authority.

Required experience

  • Built production or research-grade evaluation pipelines for generative video, computer vision, multimodal models, image generation, or a closely related domain.
  • Strong Python and data-engineering ability, including dataset versioning, manifests, validation, experiment tracking, and reproducible analysis.
  • Experience designing held-out evaluation, blinded human review, annotation quality controls, adjudication, and regression testing.
  • Able to distinguish model quality, data quality, reviewer disagreement, system failure, and policy failure in the final evidence.
  • Comfortable challenging impressive demos when the underlying comparison, data split, or evidence is weak.

Strong pluses

  • Identity or face consistency evaluation, temporal video metrics, scene or action adherence, video-quality assessment, or production-usability scoring.
  • Experience with tools such as FiftyOne, CVAT, Label Studio, W&B, MLflow, DVC, or equivalent internal systems.
  • Data provenance, consent, media rights, model governance, or high-integrity evaluation work.
  • Designed evaluation gates that stopped a bad model or system release.

To apply, please address these questions

  1. Describe the most relevant generative-model or computer-vision evaluation system you personally built. What data did it cover, what did you own, and what release or research decision did it support?
  2. How would you test character identity and temporal consistency without allowing the implementation team to tune against the holdout set?
  3. Describe one case where leakage, reviewer bias, a misleading metric, or an evaluation bug changed your initial conclusion.
  4. How would you combine automated scores and blinded human review when they disagree?
  5. Share one sanitized artifact you can walk through live, such as evaluation code, a dataset manifest, reviewer rubric, dashboard, calibration report, or experiment record.
  6. Are you able to work full time and provide reliable overlap with a distributed team?

About the company

Eros Innovation company logo

Eros Innovation

Actively Hiring
Sovereign cultural AI infrastructure for creators, content and global markets51-200 Employees
Learn more about Eros Innovation image