
Altimetrik
Actively Hiring
IT Services
Job Location
Remote Work Policy
In office
Visa Sponsorship
Not Available
RelocationAllowed
Skills
Python
AWS
Kubernetes
Prometheus
Grafana
Go
OpenTelemetry
Langfuse
Arize
WhyLabs
About the job
AI-First SRE (Senior)
Key Responsibilities
- Build and extend O11Y pipelines: metrics, distributed tracing, logging, SLO/SLI definitions and dashboards consumed platform-wide.
- Instrument AI systems in production: token usage and cost, latency-per-inference, tool-call success rates, response-quality signals, agent loop detection, runaway-cost alerting.
- Extend tracing across agentic flows — planner → executor → retrieval → tool calls — spanning GenOS, AI Gateway, and MCP Gateway.
- Define SLOs where "available" includes response quality and tool-call success, not just HTTP 200s.
- On-call and incident command for AI platform surfaces: model/provider degradation, Bedrock/Gemini failover, semantic-cache issues, prompt-injection events — blameless postmortems with tracked remediations.
- Own the SRE side of the progressive-delivery seam: canary analysis, automated rollback decisioning, chaos/resilience testing, blast-radius controls — DevOps builds the pipeline; you define the gates.
- Build AIOps agents in the Traffic Agent pattern: intelligent triage, anomaly detection, auto-remediation with hard guardrails.
- Drive UX Availability (R30) and Foundations Capabilities (R30); land network, O11Y, and cloud cost savings (tracked YTD KPIs).
Must-Have Qualifications
- 7+ years SRE/production engineering on Tier-0/Tier-1, high-traffic systems.
- Strong Go or Python — this is a build role: tooling, automation, instrumentation.
- Has built observability stacks, not just consumed them: Prometheus/Grafana, OpenTelemetry, or equivalent at scale, including cardinality and cost control.
- Production LLM/ML monitoring: Langfuse, Arize, WhyLabs, or homegrown — token/cost tracking, drift and quality metrics.
- Working fluency in AI-system failure modes: nondeterminism, provider limits and outages, context-window overflow, agent loops, cache poisoning.
- Kubernetes + AWS operational depth — debugs across cluster, mesh, and gateway layers.
- Structured incident-command and postmortem experience.
Nice-to-Have
- AIOps / LLM-applied-to-ops: auto-triage, incident summarization, remediation agents.
- Chaos engineering (Litmus, Gremlin, or homegrown); eBPF or deep network debugging.
- FinOps / cost engineering; fintech or regulated-industry reliability experience.
About the company
Similar Jobs

Kodiak Robotics
Building the world's safest driver

Kodiak Robotics
Building the world's safest driver

EarnIn
You worked today. Get paid today. Why wait for money you’ve already earned?

Kodiak Robotics
Building the world's safest driver