
- B2B
- Early StageStartup in initial stages
Senior DevOps Engineer
- $36k – $48k • No equity
- |Remote (+4)
- |5 years of exp
- |Full Time
Remote only
Not Available
About the job
We're looking for a Senior DevOps Engineer to lead the effort of transforming our cloud-native GCP infrastructure into portable, containerized solutions that can be deployed fully on-premises in client environments. This is a high-impact role where you'll architect the bridge between our current Cloud Run-based platform and self-contained on-prem deployments — including scenarios with and without locally-hosted LLMs.
You'll work directly with our engineering team and collaborate with clients to deliver reliable, secure, production-grade deployments in their infrastructure.
This is a fully remote position.
WHAT YOU'LL DO
- Architect on-prem deployment strategy: Design and implement a containerized deployment model that replaces our GCP-managed services (Cloud Run, AlloyDB, Pub/Sub, Secret Manager, Cloud Storage) with portable, self-hosted equivalents
- Build and maintain Kubernetes infrastructure: Stand up and manage K8s clusters (or similar orchestration) for on-prem environments, translating our 20+ Cloud Run microservices into Helm charts or Kustomize manifests
- Enable on-prem LLM serving: Design GPU-accelerated infrastructure for local LLM inference (e.g., vLLM, Ollama, TGI) as an alternative to Vertex AI, supporting models like Llama and others in air-gapped or restricted environments
- Harden Docker images: Optimize and secure our 50+ existing Dockerfiles for production on-prem use — smaller images, non-root users, vulnerability scanning, reproducible builds
- Replace GCP-managed services with self-hosted equivalents: Identify and implement replacements for AlloyDB (PostgreSQL + pgvector), Memorystore (Redis), Pub/Sub (NATS/RabbitMQ/Redis Streams), Secret Manager (Vault), Cloud Storage (MinIO/S3-compatible), and Artifact Registry
- Develop infrastructure-as-code: Extend our existing Terraform codebase or introduce complementary tooling (Ansible, Pulumi) for repeatable on-prem provisioning
- Build CI/CD for on-prem: Adapt our Cloud Build pipelines into portable CI/CD workflows (GitHub Actions, GitLab CI, or similar) that support both cloud and on-prem deployment targets
- Implement observability: Deploy self-hosted monitoring and logging (we already use Loki + Grafana) with alerting appropriate for client environments
- Support client implementations: Assist with deploying, configuring, and troubleshooting YellowPad in client infrastructure — including networking, security, and integration requirements
- Document everything: Create runbooks, deployment guides, and operational documentation for both internal teams and client IT staff
WHAT WE'RE LOOKING FOR
Required:
- 5+ years in DevOps, SRE, or Platform Engineering with production infrastructure responsibility
- Deep Kubernetes expertise: Cluster provisioning, Helm/Kustomize, networking (Ingress, service mesh), storage classes, RBAC, and operations in on-prem or bare-metal environments
- Strong Docker skills: Multi-stage builds, image optimization, security hardening, private registries
- GCP experience: Familiarity with Cloud Run, Cloud SQL/AlloyDB, Pub/Sub, Secret Manager, VPC networking, IAM — you need to understand what you're migrating away from
- Terraform proficiency: Comfortable with modular Terraform (we have multi-environment setups with shared modules for database, networking, secrets, KMS, and Redis)
- PostgreSQL administration: Deployment, configuration, backups, replication, and performance tuning — including familiarity with pgvector for vector search workloads
- Networking fundamentals: VPNs, firewalls, DNS, TLS/certificate management, private networking in enterprise environments
- Linux systems administration: Comfortable operating and troubleshooting in production Linux environments
- CI/CD pipeline design: Experience building deployment pipelines that support multiple target environments
- Scripting: Proficiency in Bash and at least one of Python or Go for automation
Highly Valued:
- On-prem LLM infrastructure: Experience deploying and operating GPU-accelerated inference servers (vLLM, Triton, TGI, Ollama, MLC-LLM) — including NVIDIA driver/CUDA management, model quantization, and performance tuning
- Air-gapped / restricted environment deployments: Experience with environments that have limited or no internet access
- GPU infrastructure management: NVIDIA GPU provisioning, monitoring, driver management, and container runtime configuration (nvidia-container-toolkit)
- Secrets management: HashiCorp Vault or similar for on-prem secrets
- Object storage: MinIO or S3-compatible storage deployment and operations
- Message queue systems: NATS, RabbitMQ, or Redis Streams as Pub/Sub alternatives
- Managed database alternatives: Experience running highly-available PostgreSQL (Patroni, Crunchy Data, CloudNativePG operator)
- Security and compliance: SOC 2, FedRAMP, or similar frameworks; container vulnerability scanning; network segmentation
Nice to Have:
- Client-facing experience: Comfortable communicating with non-technical stakeholders and client IT teams
- Monitoring stack deployment: Prometheus, Grafana, Loki, Alertmanager in production
- Firebase Auth alternatives: Keycloak, Auth0, or similar identity provider deployment
- Cloud Run to Kubernetes migration: Specific experience translating Cloud Run services into K8s workloads
- Document processing pipelines: Familiarity with OCR (Tesseract, PaddleOCR), PDF processing, or ML pipeline orchestration
OUR CURRENT STACK (WHAT YOU'LL BE WORKING WITH)
Compute — GCP Cloud Run (20+ services), Compute Engine (GPU workloads)
Backend — Python 3.11+ / FastAPI, Node.js / TypeScript / Fastify
Frontend — Next.js 15, React 19, TypeScript
Database — AlloyDB (PostgreSQL 15 + pgvector), Redis (Memorystore)
AI/LLM — Google Vertex AI (Gemini), MLC-LLM (local inference), LangChain
Messaging — Google Pub/Sub (async task queues for ingest, extraction, agent workers)
Storage — Google Cloud Storage
Auth — Firebase Auth + Okta SAML, JWT
IaC — Terraform (modular, multi-environment)
CI/CD — Google Cloud Build
Monitoring — Loki + Grafana, GCP Cloud Logging, Sentry
Containers — 50+ Dockerfiles, Artifact Registry
Feature Flags — Unleash (self-hosted)
WHY THIS ROLE MATTERS
Our enterprise clients increasingly require fully on-premises deployments — whether for data sovereignty, regulatory compliance, or security policy reasons. This role is the linchpin in making that possible. You'll be taking a mature, well-containerized cloud platform and making it truly portable, opening up an entirely new deployment model for the company.
About the company
Similar Jobs








