
IR Labs
Actively Hiring
Evidence-first verification for systems code
- Growing fastShowed strong hiring growth in the past month
Principal Platform Engineer
- $240k – $320k
- |Remote ()
- |5 years of exp
- |Full Time
Posted: yesterday• Recruiter recently active
Hires remotely in
Remote Work Policy
Remote only
Company Location
Visa Sponsorship
Not Available
Preferred Timezones
Pacific Time, Mountain Time, Central Time, Eastern Time
RelocationNot Allowed
Hiring contact
Geoff Sieron
Founder & General Manager • 2 years

About the job
Who We’re Looking For:
The founding Principal Platform Engineer will join a lean team and work closely with senior technical leaders to accelerate execution and deliver high-impact infrastructure to allow AI teams to focus on innovation. If you want to own the core platform that turns AI systems into a reliable, secure, and scalable product and build the backend services that power onboarding and operations, create streamlined “paved roads” for engineering, and productionize AI workloads with strong MLOps foundations, this role was written for you.
What You’ll Do:
- Own the product platform, not just the infrastructure: build the core platform layer that turns our AI systems into a reliable, supportable product customers can onboard to and use daily.
- Ship critical backend systems: design and implement core services/APIs that power signup/sign-in, tenancy, permissions, billing/entitlements, provisioning, and internal workflows so customer onboarding is seamless and repeatable.
- Build “paved roads” that prevent engineering drag: create a standardized, self-serve path from feature → deploy → operate so teams can ship without repeated bespoke work, manual steps, or fragile runbooks.
- Keep the AI critical path clean: absorb platform/ops/MLOps work that would otherwise pull our AI Scientist/MLE into toil and customer support, protecting our highest-leverage innovation hours.
- Productionize AI workloads: build pragmatic MLOps and GPU operations foundations (serving/training workflows, model artifact distribution, utilization-aware scheduling, caching/cold-start strategies) so AI cost and latency remain controlled as we scale.
- Create reliability as a system property: define SLIs/SLOs and error budgets, then enforce them through release standards, testing discipline, rollout patterns, and incident practices.
- Embed security into the platform primitives: implement safe-by-default patterns for identity, access boundaries, secrets, and dependency hygiene so security is “how the platform works,” not a separate checklist.
- Design for resilience with cost discipline: architect multi-region strategies (backup/restore, DR, failover) that are as simple as possible while meeting availability goals.
- Make observability actionable: build high-signal metrics/logs/traces and operational dashboards that reduce noise and accelerate diagnosis; run blameless incident reviews that permanently reduce repeat failures.
- Be a force multiplier across a senior team: partner closely with data infra, compiler, and AI leadership to review designs, unblock execution, and ship high-impact work directly when needed.
- What You Bring to the Table:
- Principal-level “builder/operator” experience: you’ve built and run production systems end-to-end, including design, implementation, deployment, on-call reality, and iteration under customer pressure.
- Backend engineering strength: you can ship significant backend systems (APIs, workflows, auth boundaries, tenancy) and are comfortable owning core product services, not just infrastructure automation.
- Strong systems coding: deep proficiency in Go and/or Rust (Python acceptable as supporting), plus solid shell skills; you can deliver large, correct changes quickly with tests and quality gates.
- Kubernetes and cloud fluency: strong Kubernetes/EKS experience; you understand service networking, zero-downtime rollouts, autoscaling, and safe multi-tenant patterns. Comfortable with AWS primitives (VPC, IAM, EC2, S3, RDS) and multi-region architecture.
- AI Ops / GPU awareness: experience operating GPU-backed workloads or adjacent high-performance compute. You understand utilization, scheduling tradeoffs, and practical cost/performance management in production.
- Reliability engineering mindset: you’ve used SLIs/SLOs/error budgets and know how to make reliability measurable and enforceable through engineering practice (not meetings).
- Security-by-design fundamentals: strong instincts and experience with identity/access patterns, secrets, dependency/supply-chain hygiene, and audit-friendly operational practices.
- Automation-first execution: track record of eliminating operational toil through automation across CI/CD, provisioning, testing, and operations.
- Customer-facing technical judgment: you can take customer feedback and translate it into platform capabilities that reduce friction, reduce support burden, and accelerate adoption.
- Clear communication in ambiguity: you write crisp design docs, make tradeoffs explicit, and collaborate effectively across ML/data/backend in a fast-moving environment.
About the company
1-10
Public
Statistic Analysis
- Growing fastShowed strong hiring growth in the past month
Perks
Competitive Health, Dental, and Vision
Multiple plan tiers to choose from with highly competitive company HSA contributions.
401k Contribution
3% Company Contribution
Stock Grants
IR is a ASX-listed public company - employee stock grants that vest over 3 years (at board discretion).
100% Remote
Flexible environment so you can balance work and life.
Founders
Nick Brown
Cofounder & CTO • 2 years

Geoff Sieron
Founder & General Manager • 2 years

Similar Jobs

Klaviyo
Klaviyo is the AI-first CRM built for B2C brands

LogicMonitor
We expand what’s possible for businesses by advancing the technology behind them