AI / RAG / Document Intelligence Engineer
- $80k – $100k • 2.0% – 5.0%
- |
- |6 years of exp
- |Full Time
In office - WFH flexibility
Not Available
About the job
Stealth AI Workflow Platform
About the venture
We are building a stealth AI platform for expert technical and professional workflows. The product will help users capture, search, classify, structure, and generate complex work product from technical and professional information.
The first market is highly regulated and confidentiality-sensitive, so accuracy, retrieval quality, explainability, and source-grounding are core product requirements.
We are looking for an AI / RAG / Document Intelligence Engineer to build the core intelligence layer of the platform.
The role
You will design and build the AI systems that allow the platform to understand, retrieve, classify, and generate structured outputs from complex documents and technical inputs.
This is a hands-on engineering role. You will work closely with the CTO/founding engineer and domain founders to create a reliable AI workflow, not a generic chatbot.
What you will own
You will own:
Document ingestion pipeline
Chunking strategy
Embedding strategy
Retrieval architecture
RAG pipelines
Prompt and system design
Evaluation framework
Hallucination testing
Source-grounded generation
Structured extraction
Classification workflows
Model selection and benchmarking
Human-in-the-loop feedback loops
What you will build
You will build systems for:
Parsing long technical/professional documents
Extracting structured information
Searching across documents
Combining semantic and keyword retrieval
Ranking and re-ranking results
Generating draft outputs with references
Classifying technical content into defined workflow categories
Producing confidence indicators
Testing output accuracy and consistency
Reducing hallucination risk
Required skills
You should have experience with:
Python
LLM APIs such as OpenAI, Anthropic, or similar
RAG pipelines
Embeddings
Vector databases such as pgvector, Pinecone, Qdrant, Weaviate, or similar
Semantic search
Hybrid search
Prompt engineering
Long-context document workflows
Structured extraction
JSON/schema-constrained outputs
Model evaluation
Retrieval evaluation
Basic backend integration
Highly valuable skills
Highly valuable experience includes:
Legaltech, regtech, patent, compliance, or professional-services AI
Technical document automation
PDF parsing and OCR
Document comparison
Knowledge graphs
Fine-tuning or LoRA
Enterprise AI deployment
Local/open-source LLMs
Secure or private AI deployments
Air-gapped or on-premise AI environments
Human-in-the-loop review systems
What good looks like
A strong candidate will be able to answer questions such as:
How do we know the AI output is grounded in the source?
How do we test retrieval quality?
How do we reduce hallucination risk?
When should we use semantic search versus keyword search?
How should we chunk technical documents?
How do we compare generated outputs against expert review?
What should the system refuse to answer?
How do we design AI workflows that expert users can trust?
What we are not looking for
We are not looking for:
A pure prompt engineer with no software engineering depth
A researcher who cannot integrate models into product
Someone who relies blindly on frameworks
Someone who cannot explain evaluation methods
Someone who treats LLM output as automatically correct
Compensation
This is an early-stage role.
Compensation may include:
Salary, if funding is in place
Contract-to-full-time structure
Equity or options
Milestone-based compensation
Final structure will depend on candidate profile, availability, and funding stage.
Application
Please send:
A short note about relevant AI/document work
GitHub, demos, papers, products, or shipped examples
Experience with RAG, search, embeddings, or document AI
Your availability and preferred working arrangement
About the company
Similar Jobs








