
Senior Performance Engineer, Systems & Storage
- $120k – $200k
- |
- |7 years of exp
- |Full Time
In office - WFH flexibility
Not Available
About the job
Senior Performance Engineer, Systems & Storage
Rancho Cordova, CA (On-site)
Reports to: Lab Lead, Farm GPU AI Lab
Employment status: Full-time
About FarmGPU
FarmGPU operates bare-metal NVIDIA GPU clusters (H100, H200, and B200) for AI developers, enterprises, and research organizations through RunPod Secure Cloud and direct enterprise deployments. From our facility in Rancho Cordova, we own the full stack: datacenter networking, RDMA fabrics, GPU server operations, and high-performance storage built for AI training and inference.
The FarmGPU AI Lab is our independent validation lab. Storage vendors, silicon makers, and ISVs bring us their hardware and software, and we put it through its paces on production-grade infrastructure. Most engagements end in published work: whitepapers, reference architectures, and technical and solutions briefs, plus industry benchmark submissions such as MLPerf Storage v3.0.
We're building a small, senior team dedicated to a flagship validation program with a strategic storage partner. The team builds the systems, runs the tests, and stands behind the numbers when they go public.
The Role
You'll own test methodology and benchmark execution for the lab: building test platforms, running workloads, and working out why a system performs the way it does, down to the silicon. Your results feed competitive analyses, published benchmarks, and reference architectures that partners take to market, so they have to be right, reproducible, and defensible to people who know the hardware as well as you do.
You read the block diagram before the spec sheet. When a result shifts by 4%, you can trace it to a PCIe link, a NUMA boundary, or an SSD's garbage collection. When a result looks too good, you're the first one to go find out why.
What You'll Do
Benchmarking & Methodology
- Design test plans and methodology for device, platform, and full-system benchmarks: workload selection, preconditioning to steady state, run rules, and statistical treatment that hold up to partner and public scrutiny.
- Execute benchmark campaigns across AI and storage workloads, including MLPerf Storage, DLIO, GDSIO, fio, and inference stacks such as vLLM with LMCache, with power and efficiency measured alongside performance.
- Run competitive and comparative analysis across SSDs, CPUs, GPUs, NICs, and complete platforms, with apples-to-apples configurations and every variable documented.
Analysis & Root Cause
- Explain the results, not just report them. Profile systems end to end (CPU, memory, PCIe, NVMe, network, GPU) to find the true bottleneck and quantify the headroom.
- Characterize devices under test: latency distributions, tail QoS, thermal and power behavior, and how performance moves with queue depth, block size, fill state, and firmware revision.
- Turn data into findings with clear charts, tables, and written analysis that technical marketing can build from and partners can trust.
Hardware Buildout & Configuration
- Spec and build test platforms, including custom whitebox systems designed to remove bottlenecks so the component under test is what you actually measure.
- Configure systems end to end: BIOS and firmware tuning, PCIe topology and lane mapping, NUMA placement, drive and NIC firmware, and kernel and driver settings.
- Make every result reproducible by capturing full configuration state (hardware, firmware, OS, drivers, tool versions) with every run and keeping lab records current in NetBox.
Partner Collaboration
- Work directly with partner engineering teams on test plans, early hardware and firmware samples, and result reviews.
- Pair with the lab's Network Automation Engineer to turn manual procedures into automated, repeatable test harnesses.
What You Bring
- 7+ years in performance engineering, benchmarking, or systems validation for storage, servers, silicon, or HPC/AI infrastructure.
- Working knowledge of chip and system architecture: CPU microarchitecture and memory hierarchy, NUMA, PCIe topology and bandwidth, NVMe and NAND flash behavior, and how GPUs, NICs, and storage interact across a system.
- Deep hands-on benchmarking experience with tools such as fio, MLPerf Storage/DLIO, GDSIO, elbencho, IOR, or vdbench, and the judgment to know when a benchmark is lying to you.
- Strong Linux performance analysis skills: perf, eBPF/bpftrace, iostat, blktrace, numactl, and vendor profilers such as Intel VTune, AMD uProf, or NVIDIA Nsight.
- Python and Bash fluency for test automation, data parsing, and analysis (pandas, Jupyter, or similar).
- Rigor with data: you think in variance, outliers, and confidence intervals, and you can defend every number you publish.
- Clear technical writing: test reports an engineer can reproduce and an executive can act on.
- Hands-on hardware comfort: racking, cabling, swapping drives and NICs, flashing firmware, and working from BMC consoles.
Preferred Qualifications
- Validation, performance, or architecture experience at an SSD, silicon, server OEM, or storage vendor.
- SSD depth: SNIA PTS, steady-state preconditioning, write amplification, over-provisioning, TLC vs. QLC behavior, and endurance.
- AI workloads on GPU systems: training data pipelines, checkpointing, KV cache offload, and GPUDirect Storage.
- Parallel and distributed file systems (WEKA, VAST Data, Lustre, Ceph) and NVMe-oF over RDMA or TCP.
- Published or audited benchmark results, such as MLPerf or SPEC submissions.
- Power and thermal measurement: PDU and BMC telemetry, inline power analyzers, and performance per watt.
- Fluency with AI coding tools (Claude Code, Cursor, or similar) to accelerate test tooling and analysis.
What Success Looks Like
- A documented, versioned test methodology that partners trust and anyone on the team can execute.
- A new device goes from drive-in-the-slot to a complete, reviewed performance profile in weeks, not months.
- Every published number traces back to a specific run, configuration, and dataset.
- Bottlenecks found and explained before a partner asks, with the data to back it up.
- Your results show up in published whitepapers, reference architectures, and industry benchmark submissions.
Why FarmGPU?
- Hardware first. Hands-on time with next-generation SSDs, CPUs, GPUs, and NICs, often as engineering samples.
- Real infrastructure. Production B200 clusters, 400G and 800G RDMA fabrics, and petabyte-scale AI storage. Not a simulation.
- Work that goes public. Your results get published, submitted to industry benchmarks, and presented at major events.
- Small, senior team. Your judgment sets the methodology, not a committee.
- Located in Rancho Cordova, CA, at the heart of the Sacramento region's storage and semiconductor community.
Compensation
- $120,000 to $200,000 base salary, depending on experience.
- Full-time, on-site position in Rancho Cordova, CA. Remote work is not available for this role.
Culture & Fit
FarmGPU is a small team running serious infrastructure. We value people who are close to the work, communicate directly, and close the loop, not people who create process for its own sake.
AI-first by default. We use AI to move faster across every function. For this role that means using AI to build test tooling, parse results, and draft reports, so more of your time goes to the hard part: understanding the system.
High agency. You see a gap, you own it end to end: identify → decide → execute → document.
Direct communication. If a result looks wrong, say so early. A benchmark we can't defend is worse than no benchmark.
Systems thinking. Performance is a property of the whole system. A BIOS setting, a firmware revision, or a cable in the wrong port can change the story, and you catch it before it reaches the data.
Radically transparent. Methodology, configurations, and raw data are visible internally, and we expect you to use them.

