Avatar for Alpaca
Alpaca
Actively Hiring
API-first Stock & Crypto Platform
  • Growing fast
    Showed strong hiring growth in the past month

Senior Data Scientist AI Evaluation

Posted: today• Recruiter recently active
Job Location
Hires remotely in
Everywhere
Visa Sponsorship

Not Available

RelocationNot Allowed
Hiring contact
Mike Flynn
Employee
image

About the job

Your Role: We're looking for a Senior Data Scientist, AI Evaluation to design how Alpaca measures whether our models and agents are actually right. You'll be a senior individual contributor who turns ambiguous quality questions into ground truth, scoring methods, and eval loops that the company can trust—and uses those results to make the systems better. You'll build on an established data foundation, so the focus is raising quality and speeding up safe rollout. You'll own the quality bar, independent of the teams that build and optimize those systems.

This role is for someone who cares as much about whether an answer is correct as about whether a model can generate one. You'll partner with Product, Engineering, Analytics Engineering, and business stakeholders to define what "good" looks like, build the evaluations that test it, and close the loop so evals drive iteration. If you have a strong quantitative background, have shipped rigorous, measurable work (evaluation, experimentation, or model validation), and want ownership over a greenfield eval practice at a fast-growing brokerage-infrastructure company, this is the role.

What You'll Do

  • Design AI evaluations: Define ground truth, metrics, and scoring methods for models and agents.
  • Build repeatable eval loops: Track quality over time and catch regressions before release.
  • Partner on infrastructure: Work with engineering and analytics engineering to operationalize eval harnesses.
  • Drive iteration: Translate eval results into actionable recommendations for system improvements.
  • Establish quality standards: Set evaluation guidelines, documentation, and review practices.
  • Mentor and align: Foster evaluation best practices and build a culture of measurable AI quality across the team.

What We're Looking For

  • Track record of quantitative measurement rigor (e.g., LLM/model evaluation, metric validation, or experimentation).
  • Strong statistical and ML foundation—you treat evaluations as experiments (sample sizing, confidence intervals, significance, handling non-determinism) and validate automated graders against human ground truth.
  • Proficiency in Python and SQL, with experience evaluating models in production environments.
  • Strong judgment in defining quality metrics and ground truth for ambiguous outputs.
  • Excellent communication and cross-functional collaboration skills to align technical teams and leadership.
  • Strong problem-solving ability in fast-paced, greenfield environments.
  • 6–10 years in quantitative data science or ML, with focused experience in measurement or evaluation. A quantitative degree is a plus; equivalent industry experience is equally welcome.

Nice to Have

  • Hands-on LLM/agent evaluation in production, including eval harnesses, LLM-as-judge calibration, and CI regression gates.
  • Experience evaluating text-to-SQL, analytics agents, or other systems where correctness is verifiable against data.
  • Background in fintech, brokerage, or other domains where a wrong answer has real business or risk consequences.
  • Fluency with AI tools in research and engineering workflows.

About the company

Alpaca company logo

Alpaca

Actively Hiring
More jobs
API-first Stock & Crypto Platform201-500 Employees
  • Growing fast
    Showed strong hiring growth in the past month
Learn more about Alpaca image

Funding

AMOUNT RAISED
$71.8M
FUNDED OVER
6 rounds
Rounds
B
$50000000
Series B - Jun 2021+5

Perks

Healthcare benefits
Benefits: Health benefits start on day 1. In the US this includes Medical, Dental, Vision.  In Canada, this includes supplemental healthcare.  Internationally, this includes a stipend value to offset medical costs.  
Equity benefits
Competitive Salary & Stock Options
Remote friendly
Remote-first
Miscellaneous
New Hire Home-Office Setup: One-time USD $500 Monthly Stipend: USD $150 per month via a Brex Card

Founders

Hitoshi Harada
CTO • 12 years
San Mateo
image
Yoshi Yokokawa
CEO • 12 years
Silicon Valley
image
View the team image