Avatar for Deffai
Deffai
Actively Hiring
Simulate FDA Approvals for Biotech!
  • Top 1% of responders
    Deffai is in the top 1% of companies in terms of response time to applications
  • Responds within a day
    Based on past data, Deffai usually responds to incoming applications within a day

Senior Founding AI Software Engineer

Posted: today• Recruiter recently active
Hires remotely in
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

Preferred Timezones
Pacific Time, Eastern Time
RelocationAllowed
Skills
Python
SQL
Databases
SQLAlchemy
Knowledge Management
API
Benchmarking
Data Annotation
SQL/MySQL/PostgreSQL
Model Evaluation
FastAPI
LLMs
Retrieval-Augmented Generation (RAG)
Agentic Workflow
Multiagent Systems
Agentic RAG

About the job

About the role

This is a rare and exciting opportunity for a hands-on builder who wants to take ownership and build 0 to 1 at a VC-backed, fast-growing startup. Build with former FDA decision-makers to define the future of biotech industry standards.

You don't have to know FDA or biotech, but you believe in vertical AI solutions and want to push the boundary of what is possible. Work across our entire stack from document ingestion to agents and evals, build 0 to 1 and grow with a young team. We're looking for a Senior Founding AI Software Engineer who is passionate about solving challenging problems in a highly regulated industry.

Headquartered in San Francisco, Deffai is an AI-native VC-backed seed-stage startup building an AI platform for simulating FDA regulatory intelligence, built together with former FDA reviewers. We help therapeutics companies and their investors anticipate regulatory outcomes.

Backed by top-tier tech investors in the SF Bay Area, with advisors from GitHub, OpenAI, and Mercor.

You get market standard comp, equity, viable path to Head of Engineering or beyond.

You will work closely with 2 engineers on the team, build, ship quickly and shape our product vision.

If you have a strong interest, qualify for some but not all of what's mentioned, please reach out with a short note! We manually review every application and we might make an exception.

Key Responsibilities

  1. Evals & Benchmarks: Build the eval pipeline for model benchmarking and regression suites. Use them to gate prompt, model and agent changes, measure precision and recall against expert judgment, and benchmark our performance with frontier models.
  2. Document Ingestion & Retrieval: Build ingestion and retrieval over diverse, long, and complex regulatory documents (scanned PDFs, Word files, tables, slide decks): parsing, structure-aware chunking, hybrid search, reranking, context assembly under token budgets, and knowledge graphs. Keep every answer traceable to a cited source, with provenance and versioning.
  3. Agents: Build and operate our multi-agent review workflows as they move to a shared LangGraph engine: tool use, memory, tracing, and handling cost, latency, and failures in long-running runs.
  4. Expert Collaboration: Work with former FDA experts to accelerate their regulatory review workflows. Turn their feedback into product changes and new evaluation cases.
  5. Build & Ship: Own architectural decisions that balance innovation, scalability, cost, and long-term maintainability in a high-stakes domain. Review AI-generated code for what a diff does not show: migration locks, environment boundaries, and tenant scope.
  6. Data Engineering: Work with the team on the data pipelines behind our product: document ingestion, background jobs, and the PostgreSQL schemas. Make every pipeline idempotent and observable.
  7. Infrastructure: Support the team on our AWS infrastructure: Terraform-managed ECS services, CI/CD, and separate development and production environments.
  8. Security & Tenant Isolation: Support the team with tenant isolation, role-based access control, encryption, audit logging, and data retention controls in every service that handles customer documents.

Who you are

  1. Cracked: You love to build.
  2. Strong opinions: You have strong opinions and strong product taste, and you're not afraid of pushing back.
  3. Relentless: You are tough, can handle the speed and execute quickly.
  4. Execution-Oriented: You move fast and focus on solving real problems.
  5. Clear Communicator: You collaborate well across functions and push decisions forward.

Requirements

  1. 5+ years building and operating production software, including owning the architecture of a production system
  2. Experience building evaluations against expert judgment, such as golden datasets and precision and recall, and using them to decide what ships
  3. Experience shipping LLM systems to production, such as retrieval over complex documents or agent workflows
  4. Daily use of AI coding tools, with the judgment to catch what they get wrong in migrations, infrastructure changes, and data-access code
  5. Strong Python and data engineering experience: schema design, zero-downtime migrations, and reliable pipelines and background jobs on a relational database
  6. Experience running production systems on a major cloud with infrastructure as code, CI/CD, and separate development and production environments

Nice to have

  1. Worked directly with domain experts (clinicians, lawyers, scientists) to turn their judgment into requirements and test cases
  2. Retrieval techniques: hybrid search, reranking, parsing scanned PDFs and tables, or knowledge graphs (GraphRAG or similar)
  3. Agent frameworks and runtimes: LangGraph, MCP, Claude Agent SDK, or sandboxed agent execution
  4. Post-training or fine-tuning (SFT, preference tuning, LoRA), and judgment on when it beats prompt and retrieval work
  5. Experience building secure multi-tenant systems: tenant isolation, access control, and audit logging
  6. AWS (ECS, RDS, S3) and Terraform

Our stack

You don't need prior experience with every tool in our stack. We value strong engineering fundamentals and the ability to take ownership of production systems.

  • Frontend: React, TypeScript, Vite, Tailwind CSS, Vercel AI SDK
  • Backend and persistence: Python, FastAPI, SQLAlchemy, PostgreSQL, Alembic
  • Cloud: AWS ECS/Fargate, ECR, RDS, S3, Application Load Balancer, Secrets Manager; Cloudflare for DNS
  • Delivery and infrastructure: Docker, GitHub Actions, Terraform and OpenTofu
  • AI and orchestration: Claude via the Anthropic API; application-managed skill packs, tool use, retrieval, streaming, and evaluation workflows
  • Data and documents: FDA corpus ingestion and SQL analytics, PDF parsing with PyMuPDF, DOCX processing and report generation with python-docx

Interview process

  1. 30min intro with CEO & Cofounder
  2. 40min technical interview with one of our engineers, including a short hands-on exercise
  3. Short paid take-home (5-8 hr), followed by a walkthrough where you explain your thought process
  4. Final round: meet one of our former FDA decision makers

We aim to decide within 2 weeks of your final round. Don't miss this rare opportunity to define the next generation of medicine approvals.

We currently do not provide visa sponsorship.

About the company

Deffai company logo

Deffai

Actively Hiring
Simulate FDA Approvals for Biotech!1-10 Employees
  • Top 1% of responders
    Deffai is in the top 1% of companies in terms of response time to applications
  • Responds within a day
    Based on past data, Deffai usually responds to incoming applications within a day
Learn more about Deffai image

Funding

AMOUNT RAISED
Undisclosed amount
FUNDED OVER
1 round
Round
PRE
Undisclosed amount
Pre-Seed - Jan 2026

Founders

Liz Wang
CEO • 1 year
San Francisco Bay Area
image
View the team image

Similar Jobs

Nextdoor company logo
Nextdoor
Nextdoor is the private social network for your neighborhood
Wonderschool company logo
Wonderschool
Quality in-home child care and preschools near you
Pulley company logo
Pulley
Helping project teams break ground faster
Character AI company logo
Character AI
Character’s mission is to give everyone on earth access to their own deeply personalized superintell
Rocket Money company logo
Rocket Money
The Money App That Works for You
Unlearn.AI company logo
Unlearn.AI
Transforming clinical development by making every trial smarter with AI
Avoma company logo
Avoma
AI Meeting Assistant, Collaboration, and Intelligence platform
tribe.ai company logo
tribe.ai
We embed elite AI engineers to ship real, production-grade AI for enterprises