Design, build, and operate the lab network: 100G to 800G Ethernet with RoCEv2 for GPU and storage traffic, plus management and out-of-band networks.
Tune for lossless, low-latency performance: PFC, ECN, congestion control, QoS, MTU, and NIC settings, validated with perftest, NCCL tests, and real storage traffic.
Design test plans and methodology for device, platform, and full-system benchmarks: workload selection, preconditioning to steady state, run rules, and statistical treatment that hold up to partner and public scrutiny.
Execute benchmark campaigns across AI and storage workloads, including MLPerf Storage, DLIO, G...
About FarmGPU
FarmGPU is a GPU cloud infrastructure company operating bare-metal NVIDIA H100, H200, and B200 clusters from our facility in Rancho Cordova, CA. We operate under a revenue-sharing model with hardware partners, deliver compute capacity through the RunPod Secure Cloud marketplace and ...