Avatar for Lemonade Stand
Lemonade Stand
Actively Hiring
We're product focused software agency, committed to building successful applications

Senior Agentic AI Engineer

Reposted: 7 days ago• Recruiter recently active
Hires remotely in
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

Preferred Timezones
Eastern Time
RelocationNot Allowed
Skills
Python
Javascript
Machine Learning
AI
TypeScript
JavaScript (ES5/ES6) / TypeScript
Generative AI
LLMs
LLM Frameworks (Langchain, Claude, LLamaIndex) RAG Technologies Embedding Models Vect
Vector Databases (Pinecone Serverless, Qdrant, FAISS, ChromaDB)
Hiring contact
Chris Campbell
Employee
New York City
image

About the job

Summary
We're hiring a senior Agentic AI Engineer on a project-based engineer to audit, architect, and harden our production personalization engine. You'll work directly with our Head of Product and engineering team to take working prototypes to production-ready quality before November 2026.

Estimated 30–60 hours per month, ~3 months, with the possibility of extension.

We're building a career-intelligence and upskilling platform serving learners across MENA and Africa. We deliver outcomes — completed cohorts, secured placements, career progression — for government training contracts, university partnerships, and large-employer partnerships.

What you'll do:
We've prototyped a personalization engine on top of our new Learn app. The basic framework exists to validate the concept; we want a senior engineer to make it production-grade. Specifically:

  1. Architecture audit
    Review the personalization engine end-to-end: - Zone 1 — Surfaces: homepage canvas, in-course chat, events / jobs / comms cards - Zone 2 — Agents: LangGraph supervisor + vertical agents (Courses, Events, Jobs, Comms) - Zone 3 — Backends: MongoDB Atlas vector store, course content + transcript ingestion, employer pipeline, PostHog telemetry - Zone 4 — Self-improvement loop: scoring agent → user.md → tuned routing

  2. RAG / retrieval design review

  3. Chunking strategy for video transcripts + Markdown lessons

  4. Hybrid retrieval (dense + sparse) recommendations

  5. Reranking strategy

  6. Per-user scope enforcement (no cross-tenant leakage)

  7. Multilingual retrieval — Arabic + English minimum; Arabic word-error-rate is real

  8. Vector store choice review — MongoDB Atlas today; pgvector under evaluation

  9. Prompt + eval system

  10. Supervisor routing prompts

  11. Vertical-agent prompts (Courses, Jobs, Comms)

  12. Structured-output validation

  13. Regression eval set design + CI integration

  14. Failure-mode catalog

  15. Cost discipline

  16. Per-feature + per-organization token budgets with enforcement (we bill at org level)

  17. Cache strategy (we already cache canvas cards by content version)
    Multi-tier model routing — frontier (Sonnet / GPT-4o) for paid cohorts, mid-tier for general learners, cheap-tier or self-hosted for unverified
    Anti-abuse limits — topical-relevance classification, per-user daily caps
    Cost reporting to PostHog dashboard

Our current stack

  • LLMs: OpenAI + Anthropic (multi-provider posture)
  • Orchestration: LangChain.js + LangGraph (supervisor + sub-agent pattern)
  • Vector store: MongoDB Atlas (pgvector swap under evaluation)
  • Backend: Node.js, Express, BullMQ workers, MySQL (Aurora)
  • Frontend: Next.js 15 App Router, React, Tailwind
  • Eval / observability: PostHog (in-flight); LangSmith or Helicone under evaluation

What success looks like
First 3 months we should have:

  • Architecture assessment
  • Working RAG/retrieval pass with documented quality metrics on a fixture eval set
  • Production-ready prompt + eval pipeline in CI
  • Adaptive AI framework that will improve based on learners' interactions
  • Scaffolding for evaluations / quality control
  • Cost projection for ~10K learners with cap + cache + tier strategy locked

Who you are

  • Required: - Built production agentic systems before — not just chat wrappers around an LLM API
  • Strong production RAG experience — chunking, retrieval quality, eval discipline
  • Comfortable in * * * * JavaScript / TypeScript (Node + Next.js) - LangChain.js / LangGraph experience, or strong opinions on alternatives you can defend
  • Cost-aware — you've watched LLM bills explode and have systems-level opinions about budgets, caches, multi-tier routing
  • Strongly preferred: - Multilingual retrieval (especially Arabic)
  • Eval framework experience (LangSmith, Helicone, custom)
  • Vector store experience beyond Mongo (pgvector, Qdrant, Pinecone)
  • Worked on platforms (not just internal tools) — you've shipped to real users

About the company

Lemonade Stand company logo

Lemonade Stand

Actively Hiring
We're product focused software agency, committed to building successful applications11-50 Employees
Learn more about Lemonade Stand image

Similar Jobs

Enigma Technologies company logo
Enigma Technologies
Enigma provides trusted business data built on unparalleled entity resolution
impakter.com company logo
impakter.com
Empower your sustainable lifestyle - Take action everyday
Shef company logo
Shef
Authentic dishes. Homemade. Delivered. Explore who's cooking in your neighborhood
Kick Health company logo
Kick Health
The Online Performance Medicine Clinic for Energizing Sleep and Confident Presentations
Scale AI company logo
Scale AI
Accelerate the development of AI applications
Maybern company logo
Maybern
The operating system for modern fund finance
Enigma Technologies company logo
Enigma Technologies
Enigma provides trusted business data built on unparalleled entity resolution