Avatar for YellowPad
YellowPad
Actively Hiring
The Document Data Layer for AI
  • B2B
  • Early Stage
    Startup in initial stages

Senior DevOps Engineer

Reposted: 1 month ago
Hires remotely in
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

Preferred Timezones
Coordinated Universal Time
RelocationNot Allowed
Skills
Python
PostgreSQL
Linux
Bash
Docker
Kubernetes
Terraform
CICD
GCP
Hiring contact
Ananda Sen
Founder, CEO
image

About the job

We're looking for a Senior DevOps Engineer to lead the effort of transforming our cloud-native GCP infrastructure into portable, containerized solutions that can be deployed fully on-premises in client environments. This is a high-impact role where you'll architect the bridge between our current Cloud Run-based platform and self-contained on-prem deployments — including scenarios with and without locally-hosted LLMs.

You'll work directly with our engineering team and collaborate with clients to deliver reliable, secure, production-grade deployments in their infrastructure.

This is a fully remote position.

WHAT YOU'LL DO

  • Architect on-prem deployment strategy: Design and implement a containerized deployment model that replaces our GCP-managed services (Cloud Run, AlloyDB, Pub/Sub, Secret Manager, Cloud Storage) with portable, self-hosted equivalents
  • Build and maintain Kubernetes infrastructure: Stand up and manage K8s clusters (or similar orchestration) for on-prem environments, translating our 20+ Cloud Run microservices into Helm charts or Kustomize manifests
  • Enable on-prem LLM serving: Design GPU-accelerated infrastructure for local LLM inference (e.g., vLLM, Ollama, TGI) as an alternative to Vertex AI, supporting models like Llama and others in air-gapped or restricted environments
  • Harden Docker images: Optimize and secure our 50+ existing Dockerfiles for production on-prem use — smaller images, non-root users, vulnerability scanning, reproducible builds
  • Replace GCP-managed services with self-hosted equivalents: Identify and implement replacements for AlloyDB (PostgreSQL + pgvector), Memorystore (Redis), Pub/Sub (NATS/RabbitMQ/Redis Streams), Secret Manager (Vault), Cloud Storage (MinIO/S3-compatible), and Artifact Registry
  • Develop infrastructure-as-code: Extend our existing Terraform codebase or introduce complementary tooling (Ansible, Pulumi) for repeatable on-prem provisioning
  • Build CI/CD for on-prem: Adapt our Cloud Build pipelines into portable CI/CD workflows (GitHub Actions, GitLab CI, or similar) that support both cloud and on-prem deployment targets
  • Implement observability: Deploy self-hosted monitoring and logging (we already use Loki + Grafana) with alerting appropriate for client environments
  • Support client implementations: Assist with deploying, configuring, and troubleshooting YellowPad in client infrastructure — including networking, security, and integration requirements
  • Document everything: Create runbooks, deployment guides, and operational documentation for both internal teams and client IT staff

WHAT WE'RE LOOKING FOR

Required:

  • 5+ years in DevOps, SRE, or Platform Engineering with production infrastructure responsibility
  • Deep Kubernetes expertise: Cluster provisioning, Helm/Kustomize, networking (Ingress, service mesh), storage classes, RBAC, and operations in on-prem or bare-metal environments
  • Strong Docker skills: Multi-stage builds, image optimization, security hardening, private registries
  • GCP experience: Familiarity with Cloud Run, Cloud SQL/AlloyDB, Pub/Sub, Secret Manager, VPC networking, IAM — you need to understand what you're migrating away from
  • Terraform proficiency: Comfortable with modular Terraform (we have multi-environment setups with shared modules for database, networking, secrets, KMS, and Redis)
  • PostgreSQL administration: Deployment, configuration, backups, replication, and performance tuning — including familiarity with pgvector for vector search workloads
  • Networking fundamentals: VPNs, firewalls, DNS, TLS/certificate management, private networking in enterprise environments
  • Linux systems administration: Comfortable operating and troubleshooting in production Linux environments
  • CI/CD pipeline design: Experience building deployment pipelines that support multiple target environments
  • Scripting: Proficiency in Bash and at least one of Python or Go for automation

Highly Valued:

  • On-prem LLM infrastructure: Experience deploying and operating GPU-accelerated inference servers (vLLM, Triton, TGI, Ollama, MLC-LLM) — including NVIDIA driver/CUDA management, model quantization, and performance tuning
  • Air-gapped / restricted environment deployments: Experience with environments that have limited or no internet access
  • GPU infrastructure management: NVIDIA GPU provisioning, monitoring, driver management, and container runtime configuration (nvidia-container-toolkit)
  • Secrets management: HashiCorp Vault or similar for on-prem secrets
  • Object storage: MinIO or S3-compatible storage deployment and operations
  • Message queue systems: NATS, RabbitMQ, or Redis Streams as Pub/Sub alternatives
  • Managed database alternatives: Experience running highly-available PostgreSQL (Patroni, Crunchy Data, CloudNativePG operator)
  • Security and compliance: SOC 2, FedRAMP, or similar frameworks; container vulnerability scanning; network segmentation

Nice to Have:

  • Client-facing experience: Comfortable communicating with non-technical stakeholders and client IT teams
  • Monitoring stack deployment: Prometheus, Grafana, Loki, Alertmanager in production
  • Firebase Auth alternatives: Keycloak, Auth0, or similar identity provider deployment
  • Cloud Run to Kubernetes migration: Specific experience translating Cloud Run services into K8s workloads
  • Document processing pipelines: Familiarity with OCR (Tesseract, PaddleOCR), PDF processing, or ML pipeline orchestration

OUR CURRENT STACK (WHAT YOU'LL BE WORKING WITH)

Compute — GCP Cloud Run (20+ services), Compute Engine (GPU workloads)
Backend — Python 3.11+ / FastAPI, Node.js / TypeScript / Fastify
Frontend — Next.js 15, React 19, TypeScript
Database — AlloyDB (PostgreSQL 15 + pgvector), Redis (Memorystore)
AI/LLM — Google Vertex AI (Gemini), MLC-LLM (local inference), LangChain
Messaging — Google Pub/Sub (async task queues for ingest, extraction, agent workers)
Storage — Google Cloud Storage
Auth — Firebase Auth + Okta SAML, JWT
IaC — Terraform (modular, multi-environment)
CI/CD — Google Cloud Build
Monitoring — Loki + Grafana, GCP Cloud Logging, Sentry
Containers — 50+ Dockerfiles, Artifact Registry
Feature Flags — Unleash (self-hosted)

WHY THIS ROLE MATTERS

Our enterprise clients increasingly require fully on-premises deployments — whether for data sovereignty, regulatory compliance, or security policy reasons. This role is the linchpin in making that possible. You'll be taking a mature, well-containerized cloud platform and making it truly portable, opening up an entirely new deployment model for the company.

About the company

YellowPad company logo

YellowPad

Actively Hiring
The Document Data Layer for AI1-10 Employees
  • B2B
  • Early Stage
    Startup in initial stages
Learn more about YellowPad image

Founders

Ananda Sen
Founder, CEO
image
View the team image

Similar Jobs

Archesys company logo
Archesys
Improving the government services that impact everyday lives
Astranis company logo
Astranis
Building next-generation internet satellites to get the world online
Scale AI company logo
Scale AI
Accelerate the development of AI applications
Orchard Robotics company logo
Orchard Robotics
Securing America's food supply by building the AI farmer that automates our nation's farms
EliseAI company logo
EliseAI
Building AI agents that transform complex healthcare and housing systems
NumeralHQ company logo
NumeralHQ
Sales tax on autopilot for Ecommerce & SaaS ✨ Spend 5 mins or less per month on compliance
Postman company logo
Postman
Postman is the world’s leading collaboration platform for API development
Mercor company logo
Mercor
Mercor is at the intersection of labor markets and AI research