
- Top 10% of respondersProduct.ai is in the top 10% of companies in terms of response time to applications
- Responds within a few daysBased on past data, Product.ai usually responds to incoming applications within a few days
- B2C
- +3
AI Engineer
- $375k – $450k • 1.0% – 3.0%
- |
- |Full Time
In office - WFH flexibility
Not Available
About the job
Agents write the software here. You design the harness they run inside and own the verdict on their work — an AI engineer's seat in the Office of the CEO.
Product.ai is the verified truth layer for shopping — the intelligence that tells you what's actually true about a product, including when not to buy. Our first proof at scale is SimplyCodes, the code verification service: roughly $22M in revenue at ~60% margins. Profitable. Bootstrapped since 2009. No outside investors. No board. Fewer than twenty operators, outbuilding companies 10× our size.
Strong people find us and keep finding us. They apply over months and years, because the field moves fast and the exact profile we need moves with it.
Why This Role Exists
AI agents now write most of our software. That moves the job up a level: you decide what gets built, you design the systems the agents run inside, and you prove the output is correct. You'll build like a product engineer, not a heads-down coder. Your leverage is judgment and taste, not typing speed.
You start with the agent harness behind the Office of the CEO: the live automation that lets fewer than twenty operators move like a company many times our size. A recruiting-evaluation pipeline scores candidates daily across every open role. A merchant-discovery pipeline lands new merchants daily at scale. Content, data, and ops automation run underneath it all. This year the fleet crossed a threshold: agent runs that go hours unattended became a normal unit of work for us, and that fleet now deserves a dedicated owner.
You won't stay boxed into one surface. A lot of the work is finishing what's nearly done: the founder starts a system, it runs but isn't yet hardened or owned, and you take it the last mile and own it from then on. You work inside his build loop, not beside it. The harness is where you start, not the ceiling of what you'll touch. This is a generalist builder's seat, working directly with the founder.
The System You'll Need to Model
- A fleet of production automations whose failure mode is silent death. Pipelines here rarely fail loudly. They stop, and the cost accrues invisibly until someone notices days later. The real engineering problem is liveness: dead-man switches, deterministic health checks, and alarms designed so that no automation in the fleet can die unnoticed.
- Long-horizon agent runs — hours unattended as a normal unit of work. Runs that long hold together not because someone watches them, but because the harness is right: the rules a run is bound by, the tool surface it may call, the token budget it can spend, and the verification wired into the loop. You design all four.
- Verification the agent cannot author. A generative model cannot reliably grade its own output, and a verifier that shares the generator's context will launder its own mistakes. What works is oracle separation: external truth anchors, regression corpora, and judges that never see what the builder saw. That separation is the whole game — and it's the company's thesis: verified truth a model can't fake.
- Token economics judged by what a run moved, not what it cost. Every run is instrumented for the outcome it produced, and budget shifts toward what's working while the run is still going. We're quality-maximalist: the expensive thing is a redo cycle, never tokens.
- Cortex — the shared AI brain you work inside. It isn't a tool you reach for now and then; it's the operating system the whole company runs on, and the same intelligence family we sell to the world. Every operator here works through governed AI sessions on it, and it answers its own questions from a base of 8,600+ documents. The automations you build stand on it, feed it, and answer to the same rules your own work does. Nobody else runs their company on the product they sell.
- Architecture that moves weekly. We built this harness before unattended runs were even possible at scale, and the ground keeps shifting: each model generation re-opens part of the design. You'll model where it's going and act without waiting for a brief.
If reading that energizes you, keep going. If it feels overwhelming or underspecified, this isn't the right fit.
What You Will Own
- Agentic orchestration at production scale. Ours is the agent harness behind the Office of the CEO: recruiting evaluations, merchant discovery, ops automation. You own the run designs, the runtime they execute in, and the rules that govern them. When a new automation is needed, you decide how it runs, what constrains it, and what proves it correct.
- LLM evaluation and verification design. Ad-hoc human review breaks down once volume passes a daily ceiling, and we're heading straight through that ceiling. You build what replaces it: regression corpora, oracle-separated checkers, sampling protocols, and escalation paths that put a human in the loop only where real judgment is needed.
- Liveness and observability for the whole fleet. Every automation ships with a deterministic companion that proves it's alive and correct — you own the standard and the coverage. Nothing runs in production without one; nothing dies unnoticed.
- Token economics and model routing. You decide which model runs which job and why, point real compute at business outcomes, and rebalance spend mid-run. You'll own falsifiable outcomes — each with an evidence test a stranger could run.
- Your seat, on a co-signed charter. This is the model we run here: within your first quarter, you and the founder co-sign a seat charter — one machine-checkable number that proves the seat is working, and a written split of what you decide on your own versus what you bring to him first. You own that number, and the authority that comes with it.
The craft you must already own: production automation and the verification that proves it. Comparable experience we accept: browser automation at fleet scale, LLM evaluation design, long-running pipeline orchestration, retrieval over a governed corpus — built anywhere. What you'll grow into here: long-horizon multi-agent orchestration, oracle-separated verification at company scale, and token economics as a first-class engineering discipline.
Who You Are
You form working models of running systems on your own. You can read a pipeline you've never seen and sketch its failure modes the same day, notice where your model is wrong, and update fast. When a build comes out wrong, you fix the spec that produced it, not just the symptom in front of you. You write clearly, because clear writing is evidence of clear thought — and here, what you write becomes the operating spec your agents execute.
You move between architecture and implementation without getting stuck at either altitude: designing a verification gate in the morning, shipping it alarmed and instrumented by the afternoon. You have strong, earned opinions about guardrails, evals, and agent behavior, and you make good calls in the gray instead of queuing questions. High agency is your resting state.
You can do this job by hand, and prove it — that depth of craft is exactly what lets you direct agents and still trust the result. You've built automations that ran in production for months, and you can say precisely how you knew when one was failing. Evaluation harnesses, CI gates that actually block, data pipelines with verification companions, agent systems with regression suites. You treat agents as leverage you verify, not autocomplete you trust. We care about the artifact and the reasoning behind it far more than where you built it or what's on your diploma.
Who this isn't for. This isn't the seat for everyone. It's wrong if your code is whatever the model handed you and you couldn't say why it's right — directing agents without depth of your own breaks down fast here. It's wrong if you're comfortable letting an agent grade its own work. It's wrong if you work best as a watched assistant or want a ticket queue to execute and report done. It's wrong if you think mainly in projects, timelines, and programs — the steering here happens in real time, mid-run. And it's wrong if you mostly optimize for logos on a resume. You'll be happiest if you ship the verification companion with the feature, fix the spec instead of the symptom, and want your visibility to come from architecture decisions on the record and outcomes moved — not hours, meetings, or activity.
How We Evaluate
We don't run traditional systems-engineering interviews.
- Video screen. Brief and async: about 15 minutes total, done whenever works for you.
- Calls with company stakeholders. Short conversations with key members of the team.
- Conversation with the founder. Chemistry and comprehension — can you model the system you just read about?
- Paid work trial. A paid one-to-two-week engagement — real work in our real environment. We watch how you ground yourself, whether you write the spec before the build, how you verify what your agents produce, and whether your self-assessment is honest. It is paid because it is real work, and because that respects your time.
If the work above reads like yours but your resume is unconventional, apply anyway. We hire on the work and the reasoning, not the pedigree.
Compensation & Ownership
Total first-year comp: $375,000–$450,000 (base + performance-based ownership and profit-share programs). Base: $250,000–$300,000. We set the number from what you've actually built, and every hire ships from day one.
Beyond base: eligibility for the company's ownership and profit-share programs — grants are performance-based, with terms discussed at the offer stage; 100% family premium coverage; and an effectively unlimited token budget, steered by ROI, never capped.
This is a partnership, not a salary line. The model is built to mint partners — when the company wins, you win, in real, liquid dollars, every year.
Based in Santa Monica, Los Angeles — in person, five days a week. The rooms are real rooms.
About the company
- Top 10% of respondersProduct.ai is in the top 10% of companies in terms of response time to applications
- Responds within a few daysBased on past data, Product.ai usually responds to incoming applications within a few days
- B2C
- B2B
- Early StageStartup in initial stages
- Growing fastShowed strong hiring growth in the past month
Perks
Similar Jobs








