
Sr Platform Engineer
- |6 years of exp
- |Full Time
In office
Not Available
About the job
Senior Platform Engineer — Wrench.ai
Location: Remote (US-overlapping hours) · Type: Full-time or long-term contract
Reports to: CEO · Start: Immediate
About the role
Wrench.ai is an AI-driven sales and marketing intelligence platform — predictive lead scoring, audience segmentation, competitive creative intelligence, and CRM-connected outreach. We're a small team, and our platform runs in production for enterprise and Fortune 100 clients and universities. That's the job: a small number of engineers carrying serious production weight.
We build through a lens of orchestrated automation and governance — the platform itself runs on agentic systems that automate a large share of delivery: CI/CD, monitoring, data pipelines, even parts of code review. You would own the backend and infrastructure this all runs on, extend and maintain the automation layer, and make real architectural calls with full visibility to the CEO. If you've designed or
operated sophisticated agentic systems yourself — not just used one — this role is built for you.
What you'll work on
Backend platform (~45%)
Python services behind the Wrench.ai API — REST endpoints, job orchestration, idempotent write paths, webhook handlers
PostgreSQL schema design, migrations, and backward-compatible rollout of DDL changes
Multi-tenant workspace isolation, entitlements, usage metering and coverage billing
MCP server surface and OAuth/authorization endpoints; WorkOS-based identity
LLM integration for enrichment, entity resolution, and creative analysis
Infrastructure and delivery (~25%)
AWS: ECS Fargate, Lambda, S3, SQS, SNS, RDS, Step Functions
Terraform for infrastructure-as-code
GitHub Actions CI/CD, including OIDC-based deploys and workflow_dispatch release flows
A develop → qa → prod promotion model with hotfix branching
Datadog monitoring and incident response, including refining alert thresholds so signal stays trustworthy as usage scales
Data and ML pipelines (~20%)
ELT ingestion (Fivetran, custom Lambda extractors, external API requesters)
Lead-scoring model input assembly and serving; Shapley-value driver attribution
Competitive intelligence scrapers — ad transparency sources, advertiser resolution, relevance and dedup guardrails, rate/volume caps
Creative processing: video handling via ffmpeg, perceptual-hash grouping, feature extraction
Quality and tooling (~10%)
pytest suites — the backend suite currently runs ~6,300 tests
ruff for linting, ty for type checking, Taskfile for task running
Keeping the test suite meaningful rather than merely green
What we need you to be good at
Required
6+ years building and operating production backend systems, at least 2 of them with meaningful production-ownership responsibility (deploys, on-call, incident response)
Strong Python. You should be comfortable in a large existing codebase you did not write.
PostgreSQL beyond CRUD — schema evolution, migration safety, query performance
AWS in production, and infrastructure-as-code (Terraform or equivalent)
CI/CD ownership: you have built and debugged pipelines, not just used them
A real testing practice, and the judgement to know which tests are worth writing
Strongly preferred
Experience designing or operating agentic/automated delivery systems — CI/CD, autonomous review, orchestration frameworks
Data pipeline or ELT experience
Observability practice — you have tuned alerting systems and know why that matters
LLM application work in production (integration and evaluation, not model training)
Multi-tenant SaaS, ideally serving enterprise or regulated customers
How you work — this matters as much as the stack
- You let the process write itself and evolve. Runbooks, decision records, and PR descriptions that explain
the why — you set the standard as the team grows.
- You are reachable and you take calls. Small-team engineering runs on direct
conversation, not asynchronous position papers.
You can be the only engineer in a room with a client-facing problem and handle it.
You are comfortable being reviewed and reviewing others.
What you get
Direct ownership of a platform serving enterprise, Fortune 100, and university clients — at a company where your work is visible to the CEO weekly, not filtered through four layers
A governance-and-automation-first engineering culture: you'll extend systems that already do real delivery work, not just talk about AI tooling
Genuine architectural latitude — the constraints are real but the decisions are yours
Compensation: competitive, commensurate with experience
Practical notes
Redundancy is part of the role: documentation, cross-training, and a second pair of eyes on every system, built in as the team scales.
Your first four weeks are spent mapping the system as it exists and setting up a structured onboarding path for whoever joins next.
On-call: production alerting is live via Datadog. Expect real incidents, and real support in handling them.