Avatar for FarmGPU
FarmGPU
Actively Hiring
Neocloud GPU compute with focus on storage
  • Growing fast
    Showed strong hiring growth in the past month

Network Automation Engineer

Posted: yesterday• Recruiter recently active
Job Location
Remote Work Policy

In office - WFH flexibility

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Automation
Linux
Engineering
RDMA
SRE / DevOps
Hiring contact
Jessica Cordell
Operations Manager • 10 months
Portland
image

About the job

Network Automation Engineer

Rancho Cordova, CA (On-site)

Reports to: Lab Lead, FarmGPU AI Lab

Employment status: Full-time

About FarmGPU

FarmGPU operates bare-metal NVIDIA GPU clusters (H100, H200, and B200) for AI developers, enterprises, and research organizations through RunPod Secure Cloud and direct enterprise deployments. From our facility in Rancho Cordova, we own the full stack: datacenter networking, RDMA fabrics, GPU server operations, and high-performance storage built for AI training and inference.

The FarmGPU AI Lab is our independent validation lab. Storage vendors, silicon makers, and ISVs bring us their hardware and software, and we put it through its paces on production-grade infrastructure. Most engagements end in published work: whitepapers, reference architectures, and technical and solutions briefs, plus industry benchmark submissions such as MLPerf Storage v3.0.

We're building a small, senior team dedicated to a flagship validation program with a strategic storage partner. The team builds the systems, runs the tests, and stands behind the numbers when they go public.


The Role

This role blends network engineering and DevOps. You'll design and operate the high-speed RDMA fabrics that connect the lab's GPU, CPU, and storage systems, and write the software that provisions hardware, runs test campaigns, and collects results without anyone babysitting them.

The goal is a lab where dropping in a new device and running a full benchmark suite is a scheduled job, not a week of manual work. If you'd rather write the playbook than run the same procedure by hand twice, this is your role.


What You'll Do

Network Fabric

  • Design, build, and operate the lab network: 100G to 800G Ethernet with RoCEv2 for GPU and storage traffic, plus management and out-of-band networks.
  • Tune for lossless, low-latency performance: PFC, ECN, congestion control, QoS, MTU, and NIC settings, validated with perftest, NCCL tests, and real storage traffic.
  • Reconfigure topologies between test campaigns: re-cable, re-VLAN, and re-address as each engagement demands, with NetBox kept as the source of truth.
  • Troubleshoot across layers, from optics and cabling through switch configuration, NIC firmware, and the RDMA stack.

Automation & Test Harness

  • Build the lab's automation framework: bare-metal provisioning, OS imaging, firmware and BIOS configuration, and test orchestration, all driven from code.
  • Turn benchmark procedures into pipelines: scheduled, repeatable runs that capture full system configuration and push results to a central store automatically.
  • Build the results data path: parsing, storage, dashboards, and report generation, so analysis starts from clean, trustworthy data.
  • Integrate with FarmGPU's internal provisioning and observability platforms, and improve them as the lab's needs grow.

Infrastructure & GPU Fleet Operations

  • Operate lab infrastructure and the partner-dedicated GPU fleet: NVIDIA drivers, CUDA, container toolkit, DCGM, firmware updates, and health monitoring.
  • Keep systems secure and access controlled in line with our SOC 2 practices, including partner access to lab resources.
  • Keep the performance engineers moving with environments that are ready when the test plan is.

What You Bring

  • 5+ years in network engineering, DevOps, SRE, or infrastructure automation, with hands-on ownership of data center networks and the automation that runs them.
  • Strong data center networking: L2/L3, VLANs, BGP, and switch configuration on platforms such as SONiC, NVIDIA Cumulus, or Arista EOS.
  • RDMA experience: RoCEv2 (InfiniBand a plus), lossless Ethernet design, and NVIDIA ConnectX NICs or BlueField DPUs.
  • Software automation fluency: Python and/or Go, Bash, Ansible, and Git-based workflows with CI/CD. Infrastructure as code is your default.
  • Bare-metal Linux expertise: PXE/iPXE provisioning, Redfish/IPMI, kernel and driver management, and debugging at the hardware boundary.
  • Monitoring and observability: Prometheus and Grafana. You build the dashboards and alerts, not just read them.
  • Hands-on hardware comfort: cabling, optics, racking, and component swaps.
  • Clear communication: documentation other engineers can follow, and status updates that tell the truth.

Preferred Qualifications

  • GPU cluster experience (NVIDIA H100, H200, or B200), including GPUDirect RDMA and GPUDirect Storage.
  • Storage networking: NVMe-oF over RDMA or TCP, and parallel file systems such as WEKA, VAST Data, Lustre, or Ceph.
  • Kubernetes, Slurm, or other schedulers for test and workload orchestration.
  • Experience building benchmark or test automation frameworks in a lab or validation environment.
  • Network automation tooling: NetBox, Nornir, Netmiko, or vendor APIs.
  • Fluency with AI coding tools (Claude Code, Cursor, or similar) to accelerate automation work.
  • Certifications such as CCNP, NVIDIA networking, CKA, or RHCSA are welcome; demonstrated ability matters more.

What Success Looks Like

  • A lab fabric that's documented in NetBox, monitored, and reconfigurable in hours, not days.
  • New hardware goes from racked to test-ready through automation, not a checklist.
  • Full benchmark suites run unattended, with results and configurations captured automatically.
  • No surprises for partners: the GPU fleet stays healthy, access stays controlled, and monitoring catches issues before anyone else does.

Why FarmGPU?

  • Build it right from the start. You'll define how the lab automates, not inherit someone else's scripts.
  • Serious networking. 400G and 800G RDMA fabrics and GPU clusters that stress every link.
  • Real infrastructure. Production B200 clusters and petabyte-scale AI storage, with next-generation hardware landing regularly.
  • Small, senior team. Your code runs the lab, and your judgment shapes the architecture.
  • Located in Rancho Cordova, CA, at the heart of the Sacramento region's storage and semiconductor community.

Compensation

  • $120,000 to $200,000 base salary, depending on experience.
  • Full-time, on-site position in Rancho Cordova, CA. Remote work is not available for this role.

Culture & Fit

FarmGPU is a small team running serious infrastructure. We value people who are close to the work, communicate directly, and close the loop, not people who create process for its own sake.

AI-first by default. We use AI to move faster across every function. For this role that means using AI to write and review automation code, generate configs, and speed up troubleshooting.

High agency. You see a gap, you own it end to end: identify → decide → execute → document.

Direct communication. Say the thing early. A status update that says "on track" when there's a risk is worse than no update.

Systems thinking. The lab is a system. A switch change can shift a benchmark result, and a firmware update can break a pipeline. You see those dependencies before they become problems.

Radically transparent. Configs, runbooks, and metrics are visible internally, and we expect you to use them.

Similar Jobs

CalHR company logo
CalHR
CalHR supports departments within the state of California
FarmGPU company logo
FarmGPU
Neocloud GPU compute with focus on storage
FarmGPU company logo
FarmGPU
Neocloud GPU compute with focus on storage