In office - WFH flexibility
Not Available
About the job
About the Role
The GEOS platform is LLM-powered at its core. Every application uses AI as its reasoning engine — from generating data-driven research and analysis reports, to co-piloting complex workflow sessions, to supporting operational teams in the field. The quality, reliability, and safety of these AI features is not a secondary concern: in several applications, the AI output is acted on directly by operational teams, and in others it informs sensitive care and support decisions.
As AI / LLM Engineer, you will be responsible for everything from the LLM integration layer downward: prompt design, output validation, guardrails, structured data extraction, cost management, and the evaluation pipelines that tell us whether the AI is behaving correctly in production. You will work across all 8 applications, partnering with the Senior Engineers who own each application's full-stack delivery.
This is a hands-on, production-facing role. You will be building LLM-powered features that go live across six countries on a defined schedule.
What You Will Own
• Design and maintain the platform-wide LLM integration layer: API client, retry logic, token budgeting, cost tracking, and model fallback handling across Anthropic Claude, Together AI, and OpenAI.
• Build the prompt engineering system for all applications — structured prompt templates, dynamic context injection, multi-step chain orchestration, and system prompt management per application domain.
• Implement output validation and guardrail pipelines for safety-critical applications — bounded output sets, human-in-the-loop confirmation flows, and audit trails for every AI-generated recommendation.
• Build structured data extraction pipelines: converting unstructured LLM outputs (research reports, entity analyses, operational assessments) into validated, schema-conformant structured records.
• Design and run LLM evaluation frameworks — offline evals, regression testing when prompts change, and production monitoring for output quality drift.
• Manage LLM cost across applications — instrument token usage, surface per-application cost metrics, and optimise prompts for cost-efficiency without quality degradation.
• Integrate third-party AI APIs: image analysis APIs for visual data processing workflows, and text-to-speech services for content delivery applications.
• Build AI co-pilot features for a complex workflow application — real-time suggestion and guidance, session support, and reflective prompts — with appropriate safety guardrails designed in close collaboration with the functional lead.
• Support the multi-section report generation pipeline: source citation propagation, verification logic, and section re-run and diff detection.
• Use Claude Code as your primary development environment — including for prompt iteration, eval scripting, and integration code.
What We Are Looking For
Essential
• 3–6 years of software engineering experience with at least 1–2 years focused on LLM / AI integration in production systems.
• Hands-on experience with Anthropic Claude, Together AI, or OpenAI APIs in production: prompt design, structured outputs (JSON mode / tool use), streaming, and token management.
• Strong Python — all AI integration and evaluation code lives here.
• Demonstrated ability to design prompts that produce reliable, structured outputs at production scale — not just in demos.
• Experience building guardrail and validation layers for AI outputs — schema validation, confidence thresholds, fallback handling, human-in-the-loop flows.
• Understanding of LLM cost economics: how to instrument usage, where costs accumulate, and practical optimisation techniques.
• Ability to build and run offline LLM evaluations with rigorous test sets and regression tracking.
Strongly Preferred
• Experience building agentic or multi-step LLM pipelines (research agents, document generation pipelines, multi-section report generation).
• Prior work in safety-critical or high-stakes AI applications — healthcare, legal, social services, financial compliance, or similar domains where AI output errors have real consequences.
• Familiarity with RAG (Retrieval-Augmented Generation) patterns and vector databases for knowledge-grounded AI features.
• Experience integrating non-LLM AI services: image analysis APIs (AWS Rekognition, Azure Face API) or TTS APIs.
• Knowledge of data privacy considerations in AI systems — handling PII in prompts, preventing training data leakage, and audit trail design for AI-generated content.
• Familiarity with data residency and cross-border compliance concepts as they relate to AI data flows.
• React / frontend experience sufficient to contribute to AI co-pilot UX components.
The AI Stack You Will Work With
Primary LLM Anthropic Claude
Additional LLMs Together AI; OpenAI
Image analysis Azure Face API / AWS Rekognition (visual data processing workflows)
TTS TTS API (content delivery applications)
Backend Python — all AI integration code lives here
Vector DB To be confirmed — for RAG and knowledge-grounded features
Eval tooling Custom eval harness (you build this); open to lightweight frameworks
Dev tooling Claude Code (primary co-authoring environment)
AI Feature Map Across the Platform
To give you a sense of scope — these are the categories of AI features you will build or own across the 8 applications:
• Data aggregation and reporting — multi-section report generation from open-source and structured data inputs; entity extraction; source citation verification; diff detection on re-runs.
• Workflow co-pilot — real-time guidance and suggestions during live sessions; AI-assisted plan generation; alert and escalation logic; safety guardrails on all AI recommendations.
• Field data and operations — predictive analytics pipeline; image analysis API integration; network and pattern inference from structured inputs.
• Relationship and engagement management — profile generation from aggregated data; engagement strategy suggestions.
• Risk and assessment tools — structured risk scoring from multiple data inputs; AI-assisted intervention and action plan drafting.
• Governance and operational dashboards — AI-assisted document drafting, status summarisation, and metric narrative generation.
How Success Is Measured
• AI features in the two highest-risk applications pass functional lead review with no safety-critical failures at UAT.
• LLM cost per application within agreed targets.
• Eval regression suite in place before each application goes to UAT; no prompt changes deployed without an eval run.
• Zero incidents attributable to unguarded AI output in any production application.
• All third-party AI API integrations functional and tested before the relevant application's UAT window opens.
About the company
Similar Jobs









