Avatar for WALT
WALT
Actively Hiring
WALT — Your AI Data Engineer
  • Responds within three weeks
    Based on past data, WALT usually responds to incoming applications within three weeks

DevOps Engineer

Posted: 2 days ago• Recruiter recently active
Job Location
Remote Work Policy

In office

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Microsoft Azure
AWS
Terraform
GCP
Kubernetes Helm
ArgoCD
Hiring contact
Dixita Sharma
Employee
India
image

About the job

We're hiring a DevOps Engineer with 4–5+ years running production cloud infrastructure
who has built at least one environment from scratch. You will own how our environments
are provisioned, deployed and kept healthy , using Kubernetes, Helm, Terraform and Argo
CD .

Experience: 4–5+ years managing cloud infrastructure (A WS, GCP or Azure)

Core stack: Kubernetes, Helm, Terraform, Argo CD

Must have done: taken at least one environment from an empty cloud account to
production.

What you'll do

You will make every environment reproducible from code, from the first VPC to the running application.

  • Stand up new environments end to end (network, Kubernetes cluster, databases, DNS, TLS, secrets and deploy pipeline) with Terraform, not by hand.
  • Bring existing hand-built environments under Terraform and Helm so any of them can be rebuilt from the repository.
  • Run our Kubernetes clusters: node pools, autoscaling, resource requests and limits, ingress and version upgrades.
  • Write and maintain Helm charts for application services, background workers and their stateful dependencies, with clean per-environment values.
  • Own CI/CD: build container images, promote releases from dev to staging to production, and roll back safely.
  • Manage short-lived preview environments for pull requests, including cleanup when a PR closes.
  • Set up monitoring, logging and alerting; respond to infrastructure incidents and write the follow-ups.
  • Manage access and secrets with least-privilege IAM, rotated credentials and no secrets in code.
  • Keep cloud spend visible and right-sized.
  • Write runbooks so any engineer can deploy, debug or recover an environment.

Must-haves

All of these are required; the first two are the bar we screen on.

  • 4–5+ years managing production infrastructure on at least one major cloud: AWS, GCP or Azure.
  • At least one environment built from scratch and taken to production, covering networking, compute, a Kubernetes cluster, databases, DNS/TLS and CI/CD . You can explain the decisions you made and what you would change.
  • Kubernetes in production: deployments, services, ingress, autoscaling, resource tuning, cluster upgrades and debugging pods that are pending or crash-looping.
  • Helm: you have written charts, not only installed them, including templates, per- environment values, upgrades and rollbacks.
  • Argo CD (GitOps): you have run Argo CD in production, with every environment synced from Git and releases promoted and rolled back through commits, not manual kubectl changes.
  • Terraform: you have written and maintained modules, used remote state with locking, run plan and apply in CI, and imported existing resources. CI/CD with GitHub Actions, GitLab CI, Jenkins or similar, building and deploying container images.
  • Docker and Linux fundamentals, plus scripting in Bash or Python.
  • Networking: VPCs, subnets, load balancers, DNS, TLS certificates and private connectivity between services.

Nice to have

None of these is required, but each one shortens your ramp-up.

  • MLOps or AIOps: deploying and serving ML or LLM workloads (model serving, GPU node pools, model versioning), or using AI-driven tooling for alerting and incident response. GCP specifically: GKE, Cloud SQL, Artefact Registry and IAM.
  • Running stateful services on or alongside Kubernetes, such as PostgreSQL, Redis, Neo4j and queue-backed workers.
  • Observability with Prometheus, Grafana, Loki or ELK, and OpenTelemetry.
  • Secrets management with Vault, External Secrets or a cloud secret manager.
  • Running separate, single-tenant environments for individual customers.
  • Certifications such as CKA, CKAD, HashiCorp Terraform Associate or a cloud provider's associate or professional track.

About the company

WALT company logo

WALT

Actively Hiring
More jobs
WALT — Your AI Data Engineer11-50 Employees
Company Size
11-50
Company Type
Artificial Intelligence
Company Type
Internet
Company Type
Information Technology
  • Responds within three weeks
    Based on past data, WALT usually responds to incoming applications within three weeks
Learn more about WALT image

Similar Jobs

Sosuv Consulting company logo
Sosuv Consulting
We provide specialised Product, Software Development and Support services to the global fi
Facets.Cloud company logo
Facets.Cloud
Building a DevOps Platform for Cloud Infrastructure Management
Obrive Industries company logo
Obrive Industries
Obrive Industries Private Limited is pioneering the future of immersive technology through cutting-e
Interact AI company logo
Interact AI
Interact AI puts an agentic layer on your site that demos your product
Flex.ai company logo
Flex.ai
The AI Infrastructure company
Jumeau Capital company logo
Jumeau Capital
Building developer infrastructure in the API economy. Real product. Real revenue. Lean team
Pilot (pilotplans.com) company logo
Pilot (pilotplans.com)
Multiplayer consumer AI for travel where groups plan, decide, and book together