
Senior Site Reliability Engineer
- $160k – $200k
- |Remote () • Rancho Cordova
- |5 years of exp
- |Full Time
Onsite or remote
Not Available
About the job
About FarmGPU
FarmGPU is redefining the future of GPU-powered cloud computing, delivering cost-effective, scalable, high-performance GPU infrastructure tailored for AI developers, startups, and enterprises globally. Our vertically integrated platform transforms data centers into AI-optimized facilities, accelerates storage-intensive training and inference workflows, and delivers on-demand compute via strategic partnerships such as with RunPod Secure Cloud. With sustainability, performance, and innovation at our core, we challenge the status quo of traditional cloud providers.
As we scale our infrastructure to support high-bandwidth, low-latency AI workloads, we're seeking a Senior Site Reliability Engineer to own the reliability of our production GPU clusters, storage systems, and datacenter network, not just operate them. You'll build the automation, tooling, and operational discipline that let a small team run a growing bare-metal fleet without heroics, and set the technical bar for how FarmGPU does production ops.
What You'll Own
- Fleet health end to end: design the metrics pipelines, alerting, and unified health views that give a true, real-time picture of every GPU server, storage system, and network path in production
- Turn toil into pipelines: take recurring manual work (node repair, firmware updates, drive swaps, configuration drift) and replace it with automation; own the Ansible playbooks and scripting frameworks other engineers build
- Incident leadership: serve as a senior on-call responder and incident lead for production issues: drive the incident, write the postmortem, fix the systemic root cause, not just the symptom
- SLI/SLO ownership: define, track, and continuously improve reliability targets for GPU compute, storage, and networking in partnership with engineering leadership
- Hardware qualification & lifecycle: build and improve the process for bringing new hardware generations (H100/H200/B200 and beyond) into production, including burn-in, performance baselining, and BMC/IPMI-level tooling
- Cross-functional reliability: partner with software engineering and network architecture to ensure infrastructure changes are safe, observable, and reversible
What You Bring
- 5+ years in a production SRE, DevOps, or infrastructure operations role, with demonstrated ownership of reliability for critical systems
- Deep Linux systems expertise: comfortable diagnosing issues at the kernel, service, and hardware level in a bare-metal environment
- Strong track record with monitoring and observability tooling (Grafana, Prometheus), you build dashboards and alerts, not just read them
- Fluency in Python and/or Go for building automation and tooling that other teams depend on; extensive experience running and improving Ansible (or equivalent) at fleet scale
- Solid grasp of distributed systems and datacenter networking (VLANs, switching, OSI layers 3/4) sufficient to debug cross-layer issues independently
- Hands-on experience with bare-metal server environments, including BMC/Redfish/IPMI tooling and hardware diagnostics
- Experience defining and tracking SLIs/SLOs for production services
- Comfortable being the senior voice in an incident, calm under pressure, clear in communication, decisive
- Working knowledge of containerization (Docker/Kubernetes) at an operational level
- Willingness to participate in on-call rotation, including evenings, nights, and weekends
Preferred Qualifications
- Experience with GPU server environments (NVIDIA H100/H200/B200) or HPC infrastructure at scale
- Experience with storage platforms such as NVMe, NAS, or VAST Data in production
- Exposure to security and compliance practices: secret management, access control, Linux hardening, SOC 2 familiarity
- Experience with cloud platforms (AWS, GCP, Azure) or hybrid datacenter/cloud environments
- Fluency with AI coding tools (Claude Code, Cursor, or similar) to accelerate tooling and automation work
- Relevant certifications such as RHCSA, CKA, or AWS Certified DevOps Engineer
Why FarmGPU?
- Hands-on ownership of some of the most advanced AI compute infrastructure available, with real authority to fix what's broken at the root
- Small, senior team: your technical judgment shapes how the fleet is run, not just how tickets are closed
- Direct impact: your reliability work is the thing standing between a GPU failure and a customer's training run
- Located in Rancho Cordova, CA, in the heart of a growing AI and robotics ecosystem