
- Top 1% of respondersVectorStackAI is in the top 1% of companies in terms of response time to applications
- Responds within a dayBased on past data, VectorStackAI usually responds to incoming applications within a day
AI/ML Research Engineer
- $60k – $120k • No equity
- |Remote (Everywhere) •
- |4 years of exp
- |Full Time
Onsite or remote
Not Available
About the job
VectorStackAI · Applied research that ships as product.
Full-time · Amsterdam or remote (Netherlands preferred) · Competitive salary, set by experience
Who we are
VectorStackAI is a research-led team building vertically integrated GenAI products for legal and finance. We tune the whole stack, embeddings, retrieval, rerankers, and agent runtimes, against the metric that actually ships, and we beat public leaderboards in our domains along the way. We were founded by an ex-Apple ML and ex-Cerebras Principal Research Scientist.
The role
We are hiring our first Research Engineers, and this is one posting on purpose: we care more about the shape of the person than the title. You can do real applied research (form a hypothesis, read the literature, design the experiment, invent the objective) and you can ship the result into production against a hard metric. If you only do one of those two well, this is not the role. If you do both, the seniority sorts itself out in conversation.
You own problems the whole way: from a target metric to a feature running in production. And because we build our reputation in public, you write about what you find.
What you'll work on
The work spans the full stack, and the mix shifts with whatever moves the metric. Recent and near-term problems include:
- Fine-tuning LLMs with LoRA, QLoRA, and full fine-tunes for domain accuracy in legal and finance.
- Training best-in-class domain embedding models and rerankers, and beating public leaderboards in our domains.
- Model compression: distillation, pruning, and quantization to hold accuracy while cutting latency and cost.
- Tuning agentic harnesses end to end (our AgentGrad line): turning an agent runtime into something you can genuinely optimize against a KPI.
- Building retrieval stacks (PreciseSearch): hybrid search, index design, and the full pipeline tuned together rather than one component at a time.
- Frontier bets, for the person who wants to push there. One example: distilling diffusion models toward autoregressive-level capability.
You will not do all of these. You go deep on the ones that move the number in front of you.
What we're looking for
- Genuine applied-research ability. You take a vague target and a stack of papers and produce a method that works, not just a re-implementation.
- Strong ML engineering. You write clean, fast PyTorch, you profile before you optimize, and you train and serve models without hand-holding.
- Fluency with modern fine-tuning and efficiency techniques (LoRA, distillation, quantization, pruning), or the ability to get there fast.
- You measure everything, you are suspicious of leaderboard numbers, and you know why a single-component benchmark can lie.
- You can write. A result you cannot explain to a reader is half a result.
- Ownership. Small team, high trust, you drive problems to done.
- Communication that holds up across a remote team. Tight collaboration matters to us more than being in one room.
Nice to have
- Publications, strong technical writing, open-source models or libraries, or competitive leaderboard entries.
- Depth in legal, finance, or another regulated domain.
- Production experience with agent frameworks or RAG systems.
- Large-scale or hardware-aware training and inference optimization.
How we work
We are small, senior, and research-led. Your research lands as features in our own vertically integrated products, in production. The wins you produce become public writing that builds the company. If you want a pure lab with no shipping, or pure engineering with no research, we are not the fit. If you want the loop where research, product, and reputation are the same motion, that is exactly what we are.
Who you'll work with
You will work directly with the founder: a PhD from INRIA, then four years at Apple as a Senior ML Research Scientist shipping machine learning end to end across products, then Principal Research Scientist at Cerebras working on hardware-aware optimization of large-scale ML. It adds up to a decade learning the same lesson, the one this role is built on: vertical integration beats tuning and competing one layer at a time. This is a small, senior team, so you ship alongside that experience directly, not through three layers of management.
To apply
Email [email protected] with something real: a paper, a repo, a benchmark you beat, a model you compressed, a writeup you are proud of. Tell us the metric you moved and how you moved it. Skip the generic cover letter.
About the company
- Top 1% of respondersVectorStackAI is in the top 1% of companies in terms of response time to applications
- Responds within a dayBased on past data, VectorStackAI usually responds to incoming applications within a day
Similar Jobs
