
Sr AI Engineer
- $200k – $230k
- |Remote (Everywhere) •
- |5 years of exp
- |Full Time
Onsite or remote
Not Available
About the job
Senior AI Engineer
Location: US
Type: Full-time
Department: AI & Automation — IT Operations
────────────────────────────────────────────────────────────
About the Role
We're building an AI-powered autonomous operations platform that uses multi-agent AI to detect, diagnose, and resolve IT infrastructure incidents — with minimal human intervention. Our platform coordinates specialized AI agents that collaborate across the incident lifecycle, from initial alert through resolution.
We have a working MVP processing real incidents end-to-end. We need someone to take the AI layer from functional to intelligent — improving agent reasoning, building smarter retrieval, and making the system genuinely learn from every incident it handles.
You'll own the AI layer: agent reasoning, multi-step planning, tool use, knowledge retrieval, and continuous learning.
────────────────────────────────────────────────────────────
What You'll Do
• Advance agent reasoning — Design and refine multi-step reasoning pipelines where AI agents dynamically plan, execute, and re-plan based on results. Push toward 80%+ autonomous incident resolution
• Build intelligent knowledge retrieval — Implement advanced RAG patterns including hybrid search, re-ranking, chunking strategies, and adaptive retrieval across enterprise knowledge bases
• Develop continuous learning capabilities — Build pattern mining that discovers recurring incident patterns, constructs causal models, and auto-generates operational runbooks from historical resolutions
• Design multi-agent collaboration — Architect how agents share context, negotiate priorities, and dynamically route work based on incident complexity and confidence levels
• Implement safety and guardrails — Extend the safety layer with risk scoring, blast radius estimation, and progressive autonomy models (shadow → supervised → autonomous)
• Optimize cost and latency — Implement intelligent model routing (powerful models for complex reasoning, lightweight models for classification), response caching, and token optimization strategies
────────────────────────────────────────────────────────────
What You Bring
Must Have
• 5+ Python — async/await, FastAPI or similar frameworks, strong software engineering fundamentals
• Hands-on LLM application development — You've built production systems where LLMs reason, plan, and use tools. Prompt engineering is second nature
• Experience with agentic AI frameworks — LangChain, LangGraph, CrewAI, AutoGen, or similar. You understand state management, checkpointing, and conditional routing in agent workflows
• RAG implementation experience — Vector databases, embedding models, hybrid search, retrieval evaluation. You know why naive RAG fails and how to fix it
• Comfortable with ambiguity — You'll design, build, test, and iterate in a fast-moving environment. You figure out the right approach, not wait for detailed specs
• Observability & production readiness — Tracing (Langfuse/OpenTelemetry), structured logging, prompt/version tracking, monitoring token usage, model drift detection
• Cost-performance optimization — Model selection tradeoffs (latency vs reasoning depth), caching strategies, batching, tool-call efficiency, and context window optimization
• Security & governance awareness — Prompt injection mitigation, data isolation, secrets handling, PII redaction, and secure tool access patterns
Nice to Have
• Experience with IT operations, NOC, or ITSM platforms (ServiceNow, PagerDuty, etc.)
• Knowledge of causal inference or pattern mining from operational data
• Experience with MCP (Model Context Protocol) or similar tool-use frameworks
• Familiarity with vector databases (pgvector, Pinecone, Weaviate, Qdrant)
• Prior work on AI safety and guardrails for autonomous systems
• Experience deploying LLM systems in cloud-native environments (AWS/GCP/Azure, Kubernetes)
────────────────────────────────────────────────────────────
Why Join
• Greenfield AI architecture — You're building autonomous agents that reason, learn, and collaborate — not maintaining legacy ML pipelines
• Real production impact — Every improvement directly reduces mean time to resolution and operational toil
• Cutting-edge applied AI — Multi-agent coordination, agentic RAG, continuous learning from operations data
• Ownership — Small team, high autonomy. You'll shape the AI architecture, not just implement tickets
About the company

Mapgenesys
Similar Jobs








