Avatar for Eros Innovation
Eros Innovation
Actively Hiring
Sovereign cultural AI infrastructure for creators, content and global markets

Founding AI Platform Engineering Lead

  • Remote (
    Everywhere
    )
  • |8 years of exp
  • |Full Time
Posted: 2 days ago• Recruiter recently active
Hires remotely in
Everywhere
Remote Work Policy

Remote only

Company Location
Isle of Man
Visa Sponsorship

Not Available

RelocationNot Allowed
Hiring contact
Sam Knight
Employee
image

About the job

Eros is building sovereign cultural AI: systems designed to make advanced intelligence culturally relevant, rights-aware, governable, and commercially useful. We are bringing together a major film and media catalog, a new AI platform, and a founding team, with every data and model decision governed by rights, provenance, evidence, and policy controls.

A first gateway, router, and foundation-model release are already in flight. We need a founding engineering lead who can own the architecture, ship production code every week, and raise the level of a small cross-time-zone team.

This is a full-time founding role for someone who wants long-term ownership and remains deeply hands-on.

What you will own

  • Define and evolve the common platform architecture across model access, inference, model lifecycle, evaluation, and agent-runtime interfaces.
  • Turn ambiguous product and technical goals into executable boundaries, contracts, milestones, and acceptance tests.
  • Write and review production code, not only diagrams or strategy documents.
  • Establish secure multi-tenancy, identity propagation, observability, cost controls, release gates, and failure handling.
  • Make practical build-versus-buy and open-weight-versus-hosted decisions using quality, latency, cost, security, and portability evidence.
  • Lead design reviews, incident reviews, and technical hiring panels.
  • Create a clean handoff path across internal engineers, contractors, and product owners.

You are likely a fit if you have

  • Eight or more years in backend, platform, distributed-systems, or ML infrastructure engineering.
  • Staff, principal, founding, or early-engineer ownership of a production distributed system.
  • Personally built or operated LLM-era infrastructure under real traffic, such as an AI gateway, inference service, evaluation platform, model control plane, or agent platform.
  • Owned incidents, rollbacks, capacity decisions, observability, and measurable quality, latency, and cost tradeoffs.
  • Led a small team while remaining deeply hands-on in code and design.
  • Strong written communication and several hours of reliable overlap with US Eastern time.

Experience with vLLM, TGI, TensorRT-LLM, KServe, Ray Serve, Kubernetes, OpenTelemetry, policy engines, or multi-provider model orchestration is useful, but named production ownership matters more than keyword coverage.

To apply, please address these questions

  1. Name the most relevant production platform you personally owned. What traffic or scale did it serve, what code or architecture did you own, and what measurable outcome changed because of your work?
  2. Describe one serious production incident or failed rollout. What did you diagnose, what did you change, and what permanent control did you add?
  3. How would you separate policy eligibility from model or route selection in a multi-tenant AI platform?
  4. Share one sanitized artifact you can walk through live, such as a repository, code sample, design document, incident review, or system diagram.
  5. How many hours of reliable overlap can you provide with US Eastern time?

About the company

Eros Innovation company logo

Eros Innovation

Actively Hiring
Sovereign cultural AI infrastructure for creators, content and global markets51-200 Employees
Learn more about Eros Innovation image