Avatar for Gengis AI
Gengis AI
Actively Hiring
Gengis is building AI-native software for film, content, and virtual production teams

Senior Applied AI Engineer (Agents & Evaluation)

  • $60k – $72k SGD
  • |Remote (
    Everywhere
    ) • 
  • |5 years of exp
  • |Full Time
Posted: 3 days ago• Recruiter recently active
Job Location
Remote Work Policy

Onsite or remote

Hires remotely in
Everywhere
Visa Sponsorship

Available

RelocationAllowed
Skills
Python
PostgreSQL
TypeScript
Model Evaluation
LLMs
AWS Bedrock
Openrouter
Agentic AI
Vercel AI SDK

About the job

Senior Applied AI Engineer (Agents & Evaluation)

Gengis AI
Location: Singapore (onsite preferred; hybrid or remote considered)
Employment: Full-time
Salary: S$60,000–S$72,000 / year

About Gengis

Gengis builds AI for film and TV production. Our platform, YAM, connects work from script to screen, carrying production context, creative decisions and approvals through the process.

Our founders also built Refinery Media and X3D Studio. YAM grew out of work on real productions, and you’ll work directly with the filmmakers and engineers building it.

The role

You’ll own how our product agents work and how we know they’re good.

That means designing their behaviour, choosing their models and building the evaluations behind every decision. Did a prompt change improve script scoring? Is a model worth three times the cost for this task? Does an agent need the whole script or just three scenes? You’ll answer with experiments the team can reproduce.

This is a hands-on engineering role. You’ll work in our repository, ship pull requests and turn findings into changes to prompts, skills, tools and model configurations. You’ll also write clear recommendations when the right next step needs a team decision.

Your first priorities will be script scoring and story breakdown quality, followed by the next generation of our Writer’s Room agents.

What you’ll own

  • Rubrics and quality standards. Work with film practitioners to define what good output looks like. Turn their judgment into versioned rubrics, labelled examples and acceptance thresholds. Practitioners own creative judgment; you make it measurable and understand where the measures fall short.
  • Agent design. Define each agent’s goal, prompts, skills, tools, accessible context and output contract. Set its triggers, limits on steps, tokens and cost, and behaviour when a run fails.
  • Model selection. Compare models across providers, including OpenRouter and AWS Bedrock, on quality, latency and cost. Maintain evidence for our model catalog and establish when a more efficient model is sufficient or additional reasoning effort is justified.
  • Evaluation infrastructure. Build golden fixtures, held-out datasets, experiment runners and regression suites. Prevent leakage between tuning and evaluation, and measure changes before they ship.
  • Production quality. Investigate failed or partial runs, quality regressions and cost drift. Turn findings into fixes and regression tests, working with engineering on tool and contract changes.

What you’ll work on first

Start by calibrating script scoring against practitioner judgments across short- and long-form web series. Establish a breakdown benchmark with labelled fixtures, a practitioner-approved rubric and comparisons of prompt, skill and model approaches.

Then apply that foundation to Writer’s Room agents: scene drafting from outlines, side-by-side writing proposals from multiple models, audience reactions, character interviews that respect what each character knows, and shot lists generated from reviewed scripts.

For each agent, establish a quality baseline and measure the cost and latency of an accepted output.

What you’ll bring

  • Substantial experience shipping LLM agents in production, including tool use, agent loops, structured outputs, context and retrieval design, retries and partial results.
  • A strong record in LLM evaluation: rubric design, golden sets, LLM-as-judge calibrated against humans, inter-rater agreement and recognising misleading metrics.
  • Practical judgment on model selection, supported by benchmarks on your own tasks and an understanding of token costs, reasoning effort, latency and provider differences.
  • Strong software engineering in TypeScript or Python. You can navigate an unfamiliar codebase, write useful tests and land clean pull requests.
  • Clear written communication. You can explain to a non-engineer why a result is trustworthy, where it is uncertain and what should happen next.
  • Experience working with domain experts and translating their judgment into measurable criteria.

Useful additional experience

  • Agent skills, tool contracts, MCP, progressive tool discovery or bounded graph retrieval.
  • Knowledge graphs or information extraction, particularly entity resolution, temporal reasoning and evidence-backed claims.
  • Human labelling workflows and annotation quality.
  • Screenwriting, film, TV, short-form drama or evaluation of creative writing.

Our stack includes a TypeScript monorepo, TanStack Start, oRPC, PostgreSQL with Drizzle, Inngest, Electric sync and Vercel’s AI SDK. Familiarity helps; experience with every component isn’t required.

How you’ll work

You’ll run short experiments with at least weekly checkpoints. Each experiment should preserve the versions, settings, inputs, outputs, cost and latency needed to reproduce it. Expected answers and held-out material stay out of tuning.

A well-documented negative result counts as delivery. You’ll share failures alongside successes, flag risks early and have autonomy over methods, tools and what to measure, including the freedom to challenge the framing of a problem.
The initial focus is evaluating and configuring agent systems. Foundation-model training is outside this role’s scope.

What success looks like after six months

  • Every shipped agent has a reproducible answer to “Did this change make it better?”
  • Agent designs are documented, with evidence behind model, prompt, tool and context choices.
  • Rubrics are practitioner-approved and calibrated, with known agreement levels.
  • Cost per accepted output is measured and improving without quality regressions.

Apply

Send your CV or profile and a short example of an agent system you shipped or an evaluation that changed a product decision. Tell us what you owned, how you measured the result and what you learned. A written summary is welcome if the work is confidential.

Apply through Wellfound.

About the company

Gengis AI company logo

Gengis AI

Actively Hiring
Gengis is building AI-native software for film, content, and virtual production teams1-10 Employees
Company Size
1-10
Company Industries
Artificial Intelligence / Machine Learning
Learn more about Gengis AI image

Founders

Karen Seah
Founder
Singapore
image
View the team image