Avatar for Predii
Predii
Actively Hiring

AI/ML Infrastructure Engineer

Posted: 2 weeks ago• Recruiter recently active
Job Location
Remote Work Policy

In office - WFH flexibility

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
Automation
Linux
Bash
DNS
Network Security
Capacity Planning
TCP/IP
MacOS
Azure
Load Balancing
Segmentation
Dr
Incident Response
Monitoring
Docker
Ansible
Disaster Recovery
Windows Server
AWS
Elk
Kubernetes
Firewalls
Terraform
GitLab CI
Business Continuity
Prometheus
Grafana
Okta
RBAC
GCP
logging
helm
Alerting
CI/CD
Azure DevOps
Cloud Cost Optimization
GitHub Actions
AKS
SOC 2 Type 2
Tracing
VPNs
Access Controls
Change Controls
NSGs
Cloud IAM
Vuln Management
Multi-Tenant Auth

About the job

AI/ML Infrastructure Engineer

Mid-Level to Senior | Engineering & Platform Ops

US-based — CA preferred, open to West Coast + remote. Hybrid-friendly.

Address: 2211 Park Blvd, Palo Alto, CA 94306

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

ABOUT PREDII

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Predii builds the intelligence layer that runs the automotive service and parts industry. Our platform, Predii 360, turns messy repair-order, DMS, and parts data into real-time intelligence — powering parts lookup, diagnostics, and repair search for dealership and aftermarket customers at scale, processing billions of repair orders and serving live search at sub-second latency.

We're small, fast, and allergic to red tape. No 12-layer approval chains, no work that disappears into a backlog forever. If you build something here, it ships — and real customers use it. Learn more at www.predii.com.

PREDII RESEARCH

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

We do real research, not just integration. We continue to submit state-of-the-art research on topics including: engineering-diagram and technical-document understanding, domain-calibrated evaluation frameworks for technical content, multi-agent architectures that optimize for correctness and honesty, detecting "confident-but-wrong" failures that standard monitoring misses, moving from reactive detection to causal, explainable prognosis, and multilingual evaluation of technical and repair content. We've found that multi-agent systems that just concatenate outputs get less trustworthy as they get more capable, so we design ours to contest and qualify each other's findings instead. And we run open-weight models in production at enterprise scale, because repair-grade accuracy shouldn't cost frontier-model money. All of it is deliberately vertical: deep automotive domain expertise applied to automotive problems, not a general-purpose model with an automotive skin.

THE VIBE

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

We need a AI/ML Infrastructure Engineer who wants more than tickets — someone ready to actually own infrastructure across multiple clouds and help shape how we build. This is real ownership, not busywork. You'll touch:

  • Multi-cloud infra (Azure, GCP, AWS)
  • Kubernetes, CI/CD, automation-everything
  • Security, compliance, access — keeping the house locked
  • Monitoring & reliability — catching problems before customers do
  • Incident response & disaster recovery
  • Cloud cost optimization (yes, we care about the bill too)

Senior folks: expect to shape architecture and mentor the team, not just execute someone else's roadmap.

WHAT YOU'LL ACTUALLY DO

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  • Cloud & Platform — build + run deployments on Azure, GCP, AWS with Kubernetes/AKS, Terraform, Ansible, Helm.
  • DevOps & CI/CD — ship pipelines that are reliable and repeatable, not held together with duct tape.
  • Reliability & Observability — build monitoring, logging, alerting, tracing; hunt down root causes, not just symptoms.
  • Security & Compliance — RBAC, auth, network security, vuln management, SOC 2 Type 2 controls.
  • Resilience & Ops — own backups, DR, capacity planning, cloud costs, and production support.
  • Keep Leveling Up — evaluate new tools across DevOps, infra, and DevSecOps; you're not stuck with 2019's stack.
  • IT Support — jump in on Windows Server / macOS support when needed.

YOUR TOOLKIT

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  • Cloud & Containers: Azure, GCP, AWS, Docker, Kubernetes, AKS
  • Automation: Terraform, Ansible, Helm, Bash, Python
  • CI/CD: GitHub Actions, GitLab CI, Azure DevOps (or similar)
  • Observability: Grafana, Prometheus, ELK (or equivalent)
  • Networking: TCP/IP, DNS, load balancing, VPNs, firewalls/NSGs, segmentation
  • Identity: RBAC, cloud IAM, Okta, multi-tenant auth
  • Systems: Linux, Windows Server, macOS
  • Security & Compliance: SOC 2 Type 2, access/change controls, business continuity + DR

WHAT YOU BRING

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  • 3–5+ years hands-on AI/ML Infrastructure / DevOps / SRE / Platform Engineering, in production — not just labs.
  • Strong cloud + Kubernetes chops — Azure/AKS preferred; GCP/AWS/Docker is a big plus.
  • Solid networking and security fundamentals.
  • Comfortable with Terraform, Ansible, Helm, Bash, Python (or similar).
  • Real Git-based CI/CD experience — automated deploys, security scanning included.
  • Battle-tested on observability & prod ops — monitoring, incident response, RCA, runbooks, backup, DR.
  • Working knowledge of security/compliance across Linux, Windows Server, macOS.
  • A self-starter mindset — comfortable working independently across a distributed US–India team.

BONUS POINTS

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  • Auth/IAM experience, especially multi-tenant setups.
  • Background in automotive or data-heavy platforms.
  • Been through a SOC 2 Type 2 audit before.
  • Startup or small-team energy — you've worn more than one hat.
  • Cloud, Kubernetes, or Terraform certs.
  • Senior folks: mentoring or technical leadership experience.

HOW WE ROLL

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  • Own it — flag issues early, make the call, follow through.
  • Be proactive — don't wait to be told; spot the risk, bring the fix.
  • Share the load — security and reliability are everyone's job, not just yours.
  • Make it count — your work ships to production and touches real customers, real fast.

How To Apply

About the company

Predii company logo

Predii

Actively Hiring
11-50 Employees
Company Size
11-50
Company Type
Small And Medium Business
Learn more about Predii image

Similar Jobs

Archesys company logo
Archesys
Improving the government services that impact everyday lives
Astranis company logo
Astranis
Building next-generation internet satellites to get the world online
Scale AI company logo
Scale AI
Accelerate the development of AI applications
Orchard Robotics company logo
Orchard Robotics
Securing America's food supply by building the AI farmer that automates our nation's farms
EliseAI company logo
EliseAI
Building AI agents that transform complex healthcare and housing systems
Postman company logo
Postman
Postman is the world’s leading collaboration platform for API development
Mercor company logo
Mercor
Mercor is at the intersection of labor markets and AI research
Mercor company logo
Mercor
Mercor is at the intersection of labor markets and AI research