
Artificial Intelligence Engineer / Architect
- |5 years of exp
- |Full Time
In office - WFH flexibility
Not Available
About the job
AI Engineer / Architect
Senior | Agentic AI & Applied ML
US-based — CA preferred, open to West Coast + remote. Hybrid-friendly.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
ABOUT PREDII
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Predii builds the intelligence layer that runs the automotive service and parts industry. Our platform, Predii 360, turns messy repair-order, DMS, and parts data into real-time intelligence — powering parts lookup, diagnostics, and repair search for dealership and aftermarket customers at scale, processing billions of repair orders and serving live search at sub-second latency.
We're small, fast, and allergic to red tape. No 12-layer approval chains, no work that disappears into a backlog forever. If you build something here, it ships — and real customers use it. Learn more at www.predii.com.
PREDII RESEARCH
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
We do real research, not just integration. We continue to submit state-of-the-art research on topics including: engineering-diagram and technical-document understanding, domain-calibrated evaluation frameworks for technical content, multi-agent architectures that optimize for correctness and honesty, detecting "confident-but-wrong" failures that standard monitoring misses, moving from reactive detection to causal, explainable prognosis, and multilingual evaluation of technical and repair content. We've found that multi-agent systems that just concatenate outputs get less trustworthy as they get more capable, so we design ours to contest and qualify each other's findings instead. And we run open-weight models in production at enterprise scale, because repair-grade accuracy shouldn't cost frontier-model money. All of it is deliberately vertical: deep automotive domain expertise applied to automotive problems, not a general-purpose model with an automotive skin.
THE VIBE
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
We need an AI Engineer/Architect who wants more than a model to fine-tune — someone ready to own agentic AI products end-to-end and shape the platform they run on. This is real ownership, not busywork. You'll touch:
- Agentic products across dealership operations — technician diagnostics, service advisor recommendations, pricing/inventory intelligence
- A shared architecture that supports multiple agentic products, not one-off builds per use case
- Inference performance and cost — latency and cost-per-token as design inputs, not afterthought metrics
- Evaluation frameworks tuned to what "correct" means for each application
- Safety guardrails and escalation paths for decisions with real diagnostic and financial impact
- Integration reality across varying DMS/shop management systems
Expect to shape architecture and set technical direction, not just execute someone else's roadmap.
WHAT YOU'LL ACTUALLY DO
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
- Own Agentic Products End-to-End — from concept to production, across technician diagnostics, service advisor recommendations, and pricing/inventory intelligence.
- Build the AI Layer on Shop/Dealer Systems — agentic applications that reason over past repair orders, monitor DMS activity in real time, and act on unstructured pricing/sourcing data.
- Design a Shared Platform — one architecture that supports multiple agentic products, not a pile of bespoke builds per use case.
- Build Evaluation Frameworks — tailored to each application (diagnostic accuracy, recommendation relevance, pricing correctness), both offline and in production.
- Architect Inference for Scale — high-throughput, low-latency serving on open-weight models; own cost-per-token and latency-per-request as design inputs, not just dashboards to watch.
- Define Safety & Escalation — guardrails and human-in-the-loop paths for high-stakes recommendations with diagnostic or financial impact.
- Work the Real Integration Surface — varying DMS/shop management system integrations as a core design constraint, not an afterthought.
YOUR TOOLKIT
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
- Models & Serving: Open-weight LLMs (Llama, Mistral, Qwen, or similar), vLLM/TGI/TensorRT-LLM, GPU inference optimization
- Agentic Frameworks: LangGraph, LangChain, custom agent/orchestration stacks, tool-use & function calling
- Evaluation & Observability: Offline eval harnesses, production monitoring (accuracy, relevance, drift), LLM-as-judge, tracing (LangSmith or similar)
- Data & Retrieval: RAG pipelines, vector databases, structured + unstructured data reasoning over repair orders and DMS records
- Languages & Systems: Python, distributed systems fundamentals, API design for real-time DMS integrations
- Cloud & Infra: Azure/GCP/AWS, containers/Kubernetes, cost and latency instrumentation
- Safety: Guardrails, escalation/human-in-the-loop design, risk classification for high-stakes outputs
WHAT YOU BRING
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
- 5+ years building and shipping production ML/AI systems, with real users depending on the output — not just research or labs.
- Hands-on experience building agentic or LLM-powered applications end-to-end, from prototype to production.
- Track record architecting shared platforms/services that support multiple product use cases, not single-purpose builds.
- Practical experience with evaluation frameworks for AI systems — designing metrics that map to business correctness, not just model benchmarks.
- Experience optimizing inference for throughput/latency/cost on open-weight models in production.
- Comfort designing safety guardrails and escalation logic for high-stakes, high-consequence recommendations.
- Experience integrating with messy, heterogeneous third-party systems (APIs, DMS, ERPs, or similar) as a given constraint.
- A self-starter mindset — comfortable owning ambiguity and working independently across a distributed US–India team.
BONUS POINTS
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
- Background in automotive, dealership, or other data-heavy vertical platforms.
- Experience with diagnostic, recommendation, or pricing systems where wrong answers have real financial/safety stakes.
- Prior work fine-tuning or serving open-weight models at scale (not just calling a hosted API).
- Startup or small-team energy — you've owned a product, not just a component.
- Technical leadership or mentoring experience.
HOW WE ROLL
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
- Own it — flag issues early, make the call, follow through.
- Be proactive — don't wait to be told; spot the risk, bring the fix.
- Share the load — quality and reliability are everyone's job, not just yours.
- Make it count — your work ships to production and touches real customers, real fast.
How To Apply
- Email: [email protected]
About the company

Predii
Similar Jobs








