
- Top 10% of respondersSeekr is in the top 10% of companies in terms of response time to applications
- Responds within a weekBased on past data, Seekr usually responds to incoming applications within a week
About the job
We're looking for an AI Site Reliability Engineer to ensure the reliability, scalability, and safe operation of Seekr's AI-powered services and supporting infrastructure. You will combine software engineering and site reliability practices with AI/ML operational expertise to improve how models, APIs, data pipelines, and platform services are deployed, monitored, and operated in production. You will partner closely with AI/ML, Platform, Security, and Product teams to build dependable systems that are performant, resilient, and ready to scale.
The Impact
You will help define how Seekr operates mission-critical AI systems in production. Your work will make releases safer, incidents less frequent and easier to resolve, and service health more visible and measurable. By establishing meaningful SLOs, strengthening observability, improving resilience, and reducing operational toil, you will directly improve the experience of our customers and engineering teams.
Duties and Responsibilities
- Build safe release processes using CI/CD automation, progressive deployments, automated validation, rollback mechanisms, and deployment health metrics.
- Define and operate SLIs/SLOs for AI APIs and critical services, covering availability, latency, errors, throughput, model quality, output safety, user-facing correctness, and cost.
- Develop actionable observability through metrics, logs, traces, dashboards, and SLO-based alerts.
- Participate in a sustainable on-call rotation; lead incident response, improve runbooks, and facilitate blameless postmortems.
- Reduce operational toil and improve resilience through infrastructure as code, automation, and disaster-recovery planning.
- Design automated load, stress, spike, soak, and scalability tests that model realistic AI production traffic.
- Establish performance baselines, release thresholds, and capacity forecasts for latency, throughput, concurrency, resource utilization, and cost per inference.
- Validate autoscaling, rate limiting, backpressure, graceful degradation, and recovery from infrastructure and dependency failures.
- Design and scale GPU-backed model inference on Kubernetes using replicas, autoscaling, batching, caching, and load balancing to meet SLAs/SLOs.
- Optimize and troubleshoot inference engines and infrastructure across models, Kubernetes, compute, networking, storage, GPU memory, and cost.
- Support reliable training workloads, including distributed training, scheduling, checkpointing, observability, and failure recovery.
- Partner with AI/ML, Platform, Security, and Product teams to establish production-readiness standards.
Qualifications and Skills Required
- Bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience.
- 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, production infrastructure engineering, or a similar role.
- Strong knowledge of distributed systems, cloud infrastructure, networking, containers, and Kubernetes.
- Experience building and operating CI/CD pipelines, automated validation, and progressive deployment strategies.
- Proficiency with infrastructure-as-code and configuration-management tools, such as Terraform.
- Experience with monitoring, logging, distributed tracing, alerting, and incident-management practices, including tools such as Prometheus, Grafana, and OpenTelemetry.
- Strong scripting and software-development skills in Python, Go, or a similar language.
- Practical knowledge of SLIs, SLOs, error budgets, capacity planning, production-readiness reviews, and blameless postmortems.
- Ability to troubleshoot complex systems across application, model-serving, and infrastructure layers.
- Strong communication and collaboration skills, with a commitment to sustainable and blameless operations.
- Experience operating machine-learning or generative-AI systems in production is preferred.
- Familiarity with model serving, inference optimization, LLM gateways, vector databases, GPU infrastructure, ML observability, and model evaluation is preferred.
- Experience effectively utilizing AI technologies and tools, including large language models, agents, or AI coding assistants, to enhance workflows and operational effectiveness.
- Hands-on experience with performance-testing tools such as k6, Locust, JMeter, or Gatling.
- Experience testing distributed systems and interpreting latency percentiles, saturation, throughput, concurrency, and resource-consumption metrics.
- Familiarity with AI inference benchmarking, GPU profiling, autoscaling, and performance-versus-cost optimization.
- Experience with inference engines such as vLLM, NVIDIA Triton, TensorRT-LLM, or similar platforms, and scaling GPU workloads on Kubernetes.
- Working knowledge of GPU infrastructure, memory management, quantization, distributed training, and frameworks such as PyTorch, DeepSpeed, or FSDP.
#LI-CT1
About the company

Seekr
- Top 10% of respondersSeekr is in the top 10% of companies in terms of response time to applications
- Responds within a weekBased on past data, Seekr usually responds to incoming applications within a week
Similar Jobs









