Avatar for Tracevision
Tracevision
Actively Hiring
AI video analysis for people, vehicles, and everything in between
  • B2B
  • Growth Stage
    Expanding market presence

AI Infrastructure Engineer

  • $120k – $180k • 0.001% – 0.01%
  • |
    Garden Grove
  • |6 years of exp
  • |Full Time
Posted: 4 days ago• Recruiter recently active
Job Location
Garden Grove
Remote Work Policy

In office

Visa Sponsorship

Available

RelocationAllowed
Skills
Enterprise Storage (SAN, NAS, DFS)
Ansible
Slurm
PyTorch
Hiring contact
Claire Roberts-Thomson
Employee
image

About the job

tl;dr
Own Trace’s on-prem AI training cluster.

Build it. Run it. Fix it. Make it boringly reliable.

Why this role exists
Trace builds computer-vision products for the real world. Our researchers and engineers train models constantly, and that work runs on our own local GPU cluster — not a rented cloud training setup.

Most small companies don’t do this. We do, because we care about owning the full AI stack for CV: data, models, training infrastructure, and the operational reality underneath it.

We’re hiring an engineer to be the solo owner of that cluster. The cluster (named tinyDans) is currently 12 pflops of bfp16 and we’re looking to 10x it over the next year.

*What you’ll own
*
You will be accountable for the health and evolution of Trace’s on-prem training environment, including:

  1. Multi-GPU training servers and related hosts
  2. Slurm scheduling and cluster operations
  3. Storage and dataset movement paths
  4. Lab networking tied to training performance and recovery
  5. Rack/hardware work, power needs, spares, cabling, and physical bring-up
  6. Incident response, root cause, and prevention
  7. Docs, runbooks, automation, and standards that make the system supportable

If training is blocked, that’s your problem to clear. If the cluster is fragile, that’s your system to harden.

As the lab grows, you’ll help design and build the next version (including the physical space it runs in):

  1. Design the room and cooling strategy
  2. Coordinate with finance, the exec team, the GC, electrician, and trades
  3. Migrate current machines into the new space
  4. Make purchase decisions and determine expansion strategy as we scale

What the work looks like
This is a hands-on staff role. You will:

  • Keep the training cluster up and highly usable for ML/CV teams
  • Administer Linux GPU servers and the NVIDIA software/runtime stack
  • Operate and improve Slurm: node state, limits, packing, maintenance windows, job failure debugging
  • Own storage workflows between NAS/shared storage and local high-speed node storage
  • Diagnose and resolve hardware and systems failures under time pressure
  • GPUs, NICs, boot devices, mounts, thermal issues, network paths, flaky nodes
  • Do physical infrastructure work
  • racking, cabling, GPU installs/swaps, parts organization, airflow/cooling sanity, bench-to-rack bring-up
  • Build automation and standardization
  • reproducible OS installs, config management, monitoring/alerting, guardrails that stop bad jobs from taking down machines
  • Write the docs and runbooks that make incidents shorter the second time
  • Partner tightly with engineering users, while remaining the clear owner of the platform
  • Bring a high-ownership Trace personality: curious, intense, practical, and proud to make hard systems work

Nice-to-have
Slurm administration in production
Ansible or equivalent config management
NAS/NFS performance and failure-mode experience
BMC/IPMI/out-of-band management
Containerized training environments and base image maintenance
Observability for machine health and job health
Prior on-prem lab or small datacenter buildout
Enough ML/PyTorch context to debug user training failures without becoming an ML Engineer

What good looks like

  • In the first 90 days, you should be able to:
  • Independently recover failed nodes and restore cluster usability
  • Make cluster health visible: monitoring, top failure modes, and runbooks
  • Reduce researcher-facing breakage from drains, bad staging, mounts, and resource contention
  • Standardize machine setup so every box is less special and less fragile
  • Earn trust as the obvious owner of training infrastructure

Who thrives here
Someone who likes ownership more than ceremony.
Someone who is energized by the fact that a small company is doing something unusually ambitious: running serious on-prem AI training infrastructure for computer vision, instead of pretending cloud credits are a strategy.

Someone who can rack a machine in the afternoon, debug a nasty NFS/NIC failure in the evening, and write the runbook so it never wastes the team’s time again.

About the company

Tracevision company logo

Tracevision

Actively Hiring
AI video analysis for people, vehicles, and everything in between11-50 Employees
Company Size
11-50
Company Type
Artificial Intelligence
Company Type
Software Development
Company Industries
B2B · SaaS · Mobile · Artificial Intelligence / Machine Learning
  • B2B
  • Growth Stage
    Expanding market presence

Employees joined from

Learn more about Tracevision image

Funding

AMOUNT RAISED
$60M
FUNDED OVER
1 round
Round
A
$60000000
Series A - Feb 2014