Avatar for Altimetrik
Altimetrik
Actively Hiring
IT Services

Senior SRE (AI)

Posted: 3 weeks ago
Job Location
Remote Work Policy

In office

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
AWS
Kubernetes
Prometheus
Grafana
Go
OpenTelemetry
Langfuse
Arize
WhyLabs

About the job

AI-First SRE (Senior)

Key Responsibilities

  • Build and extend O11Y pipelines: metrics, distributed tracing, logging, SLO/SLI definitions and dashboards consumed platform-wide.
  • Instrument AI systems in production: token usage and cost, latency-per-inference, tool-call success rates, response-quality signals, agent loop detection, runaway-cost alerting.
  • Extend tracing across agentic flows — planner → executor → retrieval → tool calls — spanning GenOS, AI Gateway, and MCP Gateway.
  • Define SLOs where "available" includes response quality and tool-call success, not just HTTP 200s.
  • On-call and incident command for AI platform surfaces: model/provider degradation, Bedrock/Gemini failover, semantic-cache issues, prompt-injection events — blameless postmortems with tracked remediations.
  • Own the SRE side of the progressive-delivery seam: canary analysis, automated rollback decisioning, chaos/resilience testing, blast-radius controls — DevOps builds the pipeline; you define the gates.
  • Build AIOps agents in the Traffic Agent pattern: intelligent triage, anomaly detection, auto-remediation with hard guardrails.
  • Drive UX Availability (R30) and Foundations Capabilities (R30); land network, O11Y, and cloud cost savings (tracked YTD KPIs).

Must-Have Qualifications

  • 7+ years SRE/production engineering on Tier-0/Tier-1, high-traffic systems.
  • Strong Go or Python — this is a build role: tooling, automation, instrumentation.
  • Has built observability stacks, not just consumed them: Prometheus/Grafana, OpenTelemetry, or equivalent at scale, including cardinality and cost control.
  • Production LLM/ML monitoring: Langfuse, Arize, WhyLabs, or homegrown — token/cost tracking, drift and quality metrics.
  • Working fluency in AI-system failure modes: nondeterminism, provider limits and outages, context-window overflow, agent loops, cache poisoning.
  • Kubernetes + AWS operational depth — debugs across cluster, mesh, and gateway layers.
  • Structured incident-command and postmortem experience.

Nice-to-Have

  • AIOps / LLM-applied-to-ops: auto-triage, incident summarization, remediation agents.
  • Chaos engineering (Litmus, Gremlin, or homegrown); eBPF or deep network debugging.
  • FinOps / cost engineering; fintech or regulated-industry reliability experience.

About the company

Funding

AMOUNT RAISED
$1.5B
FUNDED OVER
1 round
Round
S
$1500000000
Seed - Jun 2024

Similar Jobs

EarnIn company logo
EarnIn
You worked today. Get paid today. Why wait for money you’ve already earned?