Avatar for Pica
Agentic tooling platform ⚡
  • B2B
  • Early Stage
    Startup in initial stages

Senior DevOps Engineer

  • $60k – $80k • No equity
  • |Remote (
    Everywhere
    )
  • |7 years of exp
  • |Full Time
Posted: 6 months ago
Hires remotely in
Everywhere
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

Preferred Timezones
Eastern Time
Collaboration Hours
9:00 AM - 6:00 PM Eastern Time
RelocationNot Allowed
Skills
PostgreSQL
AWS Cloud Services
Terraform
CloudFlare
Teleport
Docker / Docker Compose / Kubernetes
Growth and Scale-up
GitOps
K6 LoadTesting
Google Cloud Platform (GCP)
ArgoCD
Helm Chart
Prometheus & Grafana
SRE, CI/CD, DevOps, GitOps
Grafana Stack (Loki, Tempo, Mimir, Grafana)
Large-Scale Infrastructure
Hiring contact
Krish Parekh
Founding Engineer
Mumbai
image

About the job

Senior DevOps Engineer

Department: DevOps Head
Location: Remote
Type: Full-Time (EST Hours)


About the Role

We're looking for a Senior DevOps Engineer who will own and drive the entire DevOps function of our platform. This isn't a "maintain the pipeline" role — this is a hands-on leadership position where you take full ownership of infrastructure, reliability, and delivery velocity across the organization. You'll be the person the engineering org leans on when it comes to uptime, scalability, and shipping fast without breaking things.

We operate a high-traffic, production-grade platform serving enterprise customers. Our infrastructure is complex, distributed, and built to scale. We need someone who has been through the fire, someone who has operated large-scale systems under real pressure and knows what it takes to keep them resilient, observable, and fast.

If you thrive in high-velocity environments, have strong opinions on how infrastructure should be built, and move with urgency, we want to talk to you.


What You'll Own

  • Full ownership of our cloud infrastructure across GCP (primary) and AWS, including compute, networking, storage, databases, and security
  • Kubernetes at scale — managing and evolving our GKE clusters (regional and zonal) running production workloads on high-performance node pools (c3d-highcpu series)
  • Infrastructure as Code — maintaining and extending our Terraform codebase across multiple GCP projects, AWS accounts, and Cloudflare zones
  • GitOps delivery pipeline — owning our ArgoCD-driven deployment workflow with automated image updates, self-healing sync, and Slack-integrated deployment notifications
  • CI/CD systems — GitHub Actions pipelines for build, test, lint, image publishing, and Helm chart distribution to Google Artifact Registry
  • Observability stack — full Grafana ecosystem including Grafana, Loki (log aggregation), Tempo (distributed tracing), Prometheus (metrics), and Alloy (telemetry pipelines), all backed by GCS and Minio
  • Database infrastructure — Cloud SQL PostgreSQL 17 (Enterprise HA with pgvector, pgcron, pgstat_statements), Memorystore Redis 7.2 (Standard HA), CloudNativePG operator, and MongoDB Atlas
  • Networking and edge — Cloudflare DNS/CDN/WAF across multiple domains, GCP VPCs with private subnets, Cloud NAT, private service access, and global load balancing
  • Security posture — cert-manager with Let's Encrypt, External Secrets Operator syncing from GCP Secret Manager, Cloud KMS key management, Workload Identity Federation, and Teleport for zero-trust access to clusters, databases, and internal applications
  • Uptime and incident response — Better Stack and UptimeRobot monitoring across all production and development endpoints with escalation policies and status pages
  • Performance engineering — k6-based load testing infrastructure with custom test scenarios for availability and throughput validation

What We're Looking For

Must-Haves

  • 7+ years of hands-on DevOps/SRE/Infrastructure experience, with at least 3 years operating at senior level in high-scale environments
  • Deep expertise with Google Cloud Platform — GKE, Cloud SQL, Memorystore, VPC networking, IAM, Workload Identity, Cloud KMS, Artifact Registry, Cloud NAT, and Cloud DNS
  • Production Kubernetes mastery — you've run, scaled, debugged, and recovered Kubernetes clusters handling real enterprise traffic. Helm chart authoring and management is second nature
  • Terraform proficiency — you write clean, modular Terraform with remote state, multiple providers (GCP, AWS, Cloudflare, Helm, Kubernetes), and can manage complex multi-environment configurations
  • Strong GitOps experience — ArgoCD (or equivalent) for continuous delivery with automated sync, image update strategies, and environment promotion
  • CI/CD pipeline ownership — GitHub Actions (or equivalent) for building, testing, publishing containers, and managing release automation
  • Observability expertise — designing and operating monitoring stacks (Grafana, Prometheus, Loki, Tempo or similar). You know how to instrument systems, build actionable dashboards, set meaningful alerts, and trace issues across distributed services
  • Database operations — managing PostgreSQL and Redis in production at scale, including HA configurations, backups, point-in-time recovery, and performance tuning
  • Security-first mindset — secrets management, zero-trust access patterns, certificate automation, IAM least-privilege, and encryption at rest/in transit
  • Enterprise resilience — you've designed and operated systems where downtime means real business impact. You understand HA architectures, disaster recovery, failover strategies, and capacity planning
  • Bias for speed — you ship fast, iterate constantly, and don't let perfect be the enemy of done. You know when to move quickly and when to be careful

Strong Preferences

  • Experience with Cloudflare (DNS, WAF, Workers, redirect rulesets)
  • Experience with Teleport for secure infrastructure access and audit
  • Experience with k6 or similar tools for performance and load testing
  • Experience with PostgreSQL in production
  • Familiarity with release automation tooling and scripting (shell, Python)
  • Experience operating in multi-cloud environments (GCP + AWS)

Who You Are

  • You take ownership. When something is your responsibility, you don't wait to be told what to do. You see the problem, you fix it, you improve the system so it doesn't happen again.
  • You move fast. You understand that velocity matters. You ship infrastructure changes with confidence because you've built the guardrails — not because you skip them.
  • You've seen scale. You've operated systems handling serious traffic. You know what breaks at scale and how to prevent it. You've been paged at 3 AM and you've built the systems that stop those pages from happening again.
  • You're a team player. You work closely with product engineering, you unblock developers, you make the platform better for everyone.
  • You communicate clearly. You can explain complex infrastructure decisions to technical and non-technical stakeholders. You document what matters and don't over-engineer what doesn't.

Our Stack at a Glance

| Layer | Technologies |
|---|---|
| Cloud | GCP (primary), AWS |
| Compute | GKE (Kubernetes), c3d-highcpu node pools |
| IaC | Terraform (multi-provider: GCP, AWS, Cloudflare, Helm, Kubernetes, GitHub) |
| GitOps / CD | ArgoCD, ArgoCD Image Updater |
| CI | GitHub Actions |
| Containers | Docker, Google Artifact Registry, Helm |
| Databases | Cloud SQL PostgreSQL 17, Memorystore Redis 7.2, CloudNativePG |
| Observability | Grafana, Loki, Tempo, Prometheus, Alloy, Minio |
| Networking | Cloudflare (DNS/CDN/WAF), GCP VPC, Cloud NAT, Global LB |
| Security | Teleport, cert-manager, External Secrets Operator, Cloud KMS, Workload Identity |
| Monitoring | Better Stack, UptimeRobot |
| Load Testing | k6, custom Go mock servers |
| Scripting | Python 3.11, Go 1.21, Bash |


Why Join Us

You'll have the autonomy and ownership to shape how infrastructure is built and operated at a company that's scaling fast. This isn't a role where you'll be maintaining someone else's decisions — you'll be making them. If you want to build world-class infrastructure and move at startup speed with enterprise standards, this is your role.

About the company

Pica company logo
Agentic tooling platform ⚡11-50 Employees
  • B2B
  • Early Stage
    Startup in initial stages
Learn more about Pica image

Similar Jobs

Astranis company logo
Astranis
Building next-generation internet satellites to get the world online
Scale AI company logo
Scale AI
Accelerate the development of AI applications
Assured company logo
Assured
Automated claims are now a reality
Orchard Robotics company logo
Orchard Robotics
Securing America's food supply by building the AI farmer that automates our nation's farms
EliseAI company logo
EliseAI
Building AI agents that transform complex healthcare and housing systems
NumeralHQ company logo
NumeralHQ
Sales tax on autopilot for Ecommerce & SaaS ✨ Spend 5 mins or less per month on compliance