Avatar for Dreamers
Dreamers
Actively Hiring
Specializing in building technological solutions for complex problems
  • Responds within two weeks
    Based on past data, Dreamers usually responds to incoming applications within two weeks
  • B2B
  • Early Stage
    Startup in initial stages
  • +1

ML Evaluation, Seeker of Truth

  • $159k – $239k • No equity
  • |Remote ()
  • |7 years of exp
  • |Contract
Posted: 3 weeks ago• Recruiter recently active
Hires remotely in
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

RelocationNot Allowed
Skills
Python
Machine Learning

About the job

My fellow humans of WellFound,

Time for some lighthearted world domination. This time we need someone who can tell whether the model is actually working, or merely looking confident in a lab coat.

Dreamers is building a contingency team for an Air Force SBIR effort involving AI/ML for JAG data, decision-making, and military-justice outcomes. We are not looking for a generic AI engineer, an impressive collection of model names, or a governance document that has never had to survive contact with data. We need a scientist or deeply quantitative ML practitioner who can own the evaluation evidence behind AI-generated assertions.

The proposed Phase I role owns the evaluation system behind a fact-checking workflow: a model reads a source record, produces assertions, and must show whether each assertion is supported, contradicted, or not established by the record, with exact citations. This person will define and maintain the evaluator-grade source of truth, keep humans and machines consistent as prompts and models change, and turn the results into honest go/no-go evidence for a Phase II decision.

The work may include:

  • Defining the assertion schema and labels for supported, contradicted, and insufficient-evidence findings.
  • Building and maintaining an evaluator-grade gold set with exact source citations and row-level provenance.
  • Designing clear labeling interfaces and instructions for nontechnical human reviewers.
  • Establishing adjudication, disagreement, ambiguity, and ground-truth revision procedures.
  • Comparing prompt, model, retrieval, and code versions through a reproducible regression harness.
  • Choosing metrics that reflect the real task, including per-class precision, recall, F1, citation correctness, groundedness, calibration or abstention where useful, and error severity.
  • Performing failure analysis and translating results into evaluator-visible limitations and technical go/no-go recommendations.

Strong evidence may come from AI or LLM evaluation, fact-checking, groundedness, information retrieval, NLP evaluation, annotation operations, human-in-the-loop systems, legal technology, healthcare, defense research, or another evidence-sensitive domain. An M.S. or Ph.D. in statistics, machine learning, data science, computer science, information science, or a directly equivalent body of work is preferred. DoD, Air Force, SBIR/STTR, DARPA, NSF, or other funded R&D experience is useful but not mandatory.

Generic full-stack work, dashboard validation, policy-only responsible AI, or ordinary prompt engineering without rigorous evaluation ownership will not be enough for this role. We need direct evidence involving assertion verification, source-grounded evaluation, benchmark or gold-set design, human adjudication, and reproducible comparison of model behavior.

Phase I is planned around public, synthetic, simulated, or Government-approved nonsensitive surrogate data. Please do not share classified, privileged, controlled, customer-confidential, or employer-protected information during intake.

The current plan is approximately 400 hours over six months, averaging about 15 hours per week. The role is remote, but all work must be performed in the United States. Because the current solicitation posture is ITAR-restricted and the proposal states no foreign-national participation, this intake is limited to U.S. citizens. The preferred engagement is award-contingent, part-time W-2 employment for the funded work. Submission, award, start date, and hours are not guaranteed.

The live pursuit is DAF26BZ03-NV018, with a government deadline of July 22, 2026. Near-term availability is useful, but Dreamers will request separate written permission before naming anyone or submitting a resume in this or another SBIR/STTR proposal.

Dreamers Inc. is a science and technology company built around a culture of high performance with low stress. We like to work on interesting things, learn from, and teach each other. Sometimes that means natural-language systems, quantum computing, physical robotics, or government systems that need to work correctly for real people. Other times it means web and mobile apps, and those can be pretty cool, too.

Please share your current resume or profile, links to relevant technical work or publications, two examples involving evaluation, fact-checking, groundedness, or human labeling, U.S. work location, availability, preferred hourly rate, and any current employer, client, IP, or outside-work conflict that could affect the work.

Questions? Curiosity is good. Inquire within.

About the company

Dreamers company logo

Dreamers

Actively Hiring
Specializing in building technological solutions for complex problems11-50 Employees
  • Responds within two weeks
    Based on past data, Dreamers usually responds to incoming applications within two weeks
  • B2B
  • Early Stage
    Startup in initial stages
  • Growing fast
    Showed strong hiring growth in the past month
Learn more about Dreamers image

Similar Jobs

Kero Sports company logo
Kero Sports
Real-time sports betting technology that turns live game data into in-play markets
Archesys company logo
Archesys
Improving the government services that impact everyday lives
Nextdoor company logo
Nextdoor
Nextdoor is the private social network for your neighborhood
Scale AI company logo
Scale AI
Accelerate the development of AI applications
Wonderschool company logo
Wonderschool
Quality in-home child care and preschools near you
Character AI company logo
Character AI
Character’s mission is to give everyone on earth access to their own deeply personalized superintell