
LLM Cost Reduction Engineer (Contract)
- No equity
- |Remote (Canada •+2)
- |Contract
Remote only
Not Available
About the job
About the role
MAX is an AI astrology assistant serving thousands of members through a FastAPI backend. It combines Swiss Ephemeris, retrieval and LLMs to answer questions about natal charts, transits and forecasts. We handle tens of thousands of LLM calls each week across free, basic and paid tiers.
We're looking for an experienced engineer to reduce LLM serving costs by at least 40% without reducing answer quality or accuracy.
This is not a prompt-tuning or model-selection role. Changing models is not an option. The focus is on finding savings within the existing model setup through better use of tokens, caching, request flow and deterministic alternatives.
What you'll do
Audit LLM spend by model, stage, tier and endpoint using raw logs and SQL.
Identify and cost potential savings before implementing them.
Optimize caching, prompts, payloads, request flow and unnecessary LLM calls.
Identify opportunities to replace deterministic work with code where appropriate.
Ship changes incrementally through staging and validate cost, quality and accuracy before production.
Document what worked, what didn't and why.
What we're looking for
Proven experience working with LLMs at scale and reducing LLM serving costs.
Strong understanding of token economics, caching and API pricing.
Strong SQL and production debugging skills.
Experience with Python/FastAPI and Anthropic or similar LLM APIs.
A measurement-first approach. We care about evidence, not assumptions.
Please don't apply if you haven't worked with LLMs at scale or directly reduced LLM serving costs.
Success
40%+ reduction in LLM serving cost
No regression in answer quality or factual accuracy
Savings verified against deduplicated production logs
Format
Contract, scoped engagement. Remote.
To apply, send your background and relevant experience to [email protected].
About the company

TheFutureSociety
Similar Jobs








