Avatar for BAM Ventures
BAM Ventures
Actively Hiring

Site Reliability Engineer

Posted: 4 days ago• Recruiter recently active
Job Location
Remote Work Policy

In office

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Node.js
DataDog
IAM
Cloud Functions
GCP
Pub/Sub
Vercel
AI Tools
AI/agent Workflows

About the job

What is Aisle?

Physical retail has always been a black box - brands ship product, run static promotions, and wait weeks (or months) to understand what actually happened.

Aisle builds a real-time control layer on top of physical retail. We connect shopper actions, incentives, and outcomes into a continuous feedback loop, so brands can trigger, measure, and optimize retail performance while it’s happening - not after the fact.

This is the foundation for autonomous retail execution: systems that don’t just report on what happened, but actively decide what should happen next - which offers to show, when to show them, and how to maximize outcomes in real time.

Today, 1600+ brands and 5M+ consumers use Aisle to drive measurable results in brick-and-mortar. Under the hood, that means high-volume event ingestion, real-time decisioning, and infrastructure that has to be both reliable and low-latency to influence behavior in the moment.

We’ve gotten here with a small, fast-moving team, and now we’re scaling the system to handle significantly more volume, complexity, and automation - including AI systems that are part of the production path, not just offline analysis.

You'll win here if…

  • You enjoy communicating across technical and non-technical teams looking to build shared understanding that can turn into collaborative actio
  • Can solve problems in stages with appropriate level of urgency when needed: immediate triage, temporary patch, and long term fix that increases durability of the system
  • You’re comfortable moving quickly, shipping improvements, and iterating in production
  • AI is a core part of how you work - you’ve gone well beyond basic copilots and actively use AI tools to design, debug, and ship systems faster
  • You treat AI agents as chaotic, untrusted components - your instinct is isolation, rate limiting, and permissions scoping, etc. (gist of harness engineering)
  • You have experience building or orchestrating AI/agent workflows - in production or through serious side projects

About Your Role

Reliability & Infrastructure

  • Own the reliability, scalability, and observability of our infrastructure across GCP and Vercel
  • Design and implement monitoring and alerting in Datadog, including monitors-as-code and product-level dashboards
  • Manage IAM, service accounts, and security best practices across our cloud environment
  • Participate in on-call rotation, incident response, and post-mortems — and turn learnings into systemic improvements

Core Infrastructure & State

  • Stabilize core infrastructure and stateful systems under heavy concurrent load from both consumer traffic and agent-driven workflows
  • Build and maintain our event-driven GCP infrastructure, including Cloud Functions, Pub/Sub, Workflows, and GKE/Kubernetes deployments
  • Investigate performance inconsistencies (e.g., identical replicas behaving differently) and drive durable root-cause resolution
  • Stabilize and simplify parts of the stack that create friction today (e.g., Prisma, connection management, cron workflows)
  • Automate infrastructure provisioning, deployments, and operational workflows

AI & Next-Gen Tooling: Agent Ops

  • Build agent operations infrastructure that enables AI agents to run safely and reliably in production
  • Develop secure isolation layers, LLM gateways, workflow monitoring, anomaly detection, and production telemetry
  • Help define CI/CD for an agent-driven engineering environment - including how code gets shipped, reviewed, validated, and deployed when agents are in the loop
  • Own visibility into AI usage, reliability, and spend as our agent footprint scales

Cross-functional Impact

  • Partner closely with engineering and product teams to maintain reliability without slowing development velocity
  • Act as a force multiplier across the team — helping engineers ship faster and more safely

About your skillsMust haves

  • 4+ years in SRE, DevOps, or infrastructure/platform engineering
  • Strong, hands-on experience with a major cloud platform (preferable GCP)
  • Experience with event-driven or serverless systems (Pub/Sub, Cloud Functions, etc.)
  • Solid understanding of IAM, security, and cloud best practices
  • Experience with observability tools like Datadog
  • Familiarity with Node.js environments
  • AI is part of your daily engineering workflow you use AI tools thoughtfully to improve engineering velocity, reliability, and operational excellence beyond basic code generation.
  • Experience building or orchestrating AI/agent workflows (work or serious personal projects)
  • High ownership, strong curiosity, and a bias toward action

Nice to haves

  • Experience with harness engineering patterns for untrusted agents, including isolation, rate limiting, permission scoping, auditability, and safety controls
  • Familiarity with OpenClaw, agent orchestration frameworks, or LLM gateway patterns is a plus
  • Hands-on experience with GCP Workflows for orchestration
  • Familiarity with Prisma, pgbouncer, and PostgreSQL connection pooling
  • Experience with Vercel deployment and edge computing
  • Familiarity with the k8s ecosystem
  • Familiarity with Redis and BullMQ
  • Understanding of SOC 2 compliance requirements and implementation
  • Previous experience in a high-growth startup environment
  • Previous backend engineering experience to bridge the gap between infrastructure and code

About the stack

  • Cloud: Google Cloud Platform (GCP)
  • Infrastructure: Cloud Functions, Pub/Sub, Workflows, Kubernetes
  • Database: PostgreSQL with pgbouncer, Prisma
  • Observability: Datadog
  • Runtime: Node.js, TypeScript
  • Deployment: Vercel, GCP
  • Frontend: React, Next.js, TypeScript
  • Backend: TypeScript, PostgreSQL, Next.js

About the company

Similar Jobs

Scale AI company logo
Scale AI
Accelerate the development of AI applications
Pulse company logo
Pulse
Transforming healthcare by creating remarkable experiences for doctors and patients
Kalepa company logo
Kalepa
We're on a mission to transform the $7 trillion insurance industry
Yuzu Health company logo
Yuzu Health
Create your own health plan
Merciv company logo
Merciv
The Future of Enterprise Intelligence. Read your data’s past. Write your company’s future