Avatar for Pratilipi
Pratilipi
Actively Hiring
The largest Indian language story telling platform
  • Scale Stage
    Rapidly increasing operations
  • Top Investors
    This company has received a significant amount of investment from top investors

Infrastructure Engineer (DevOps) — SDE2

Posted: 2 months ago• Recruiter recently active
Job Location
Remote Work Policy

In office - WFH flexibility

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
DevOps
Jenkins
AWS/EC2/ELB/S3/DynamoDB
Docker
AWS
IAM
Containerization
DevOps/Linux/Docker/Jenkins/Chef/Puppet/Git
CI/CD
Hiring contact
Neelima Sahu
Talent Acquisition Manager • 2 years
Bangalore Urban
image

About the job

About the role
Pratilipi is building a generational company in storytelling, and infrastructure is core to that bet. We're looking for an Infrastructure Engineer — a hands-on role embedded in a high-growth product team using AWS, CI/CD, our self-managed data layer, and the operational substrate for the AI workloads now landing in production. You'll help us self-host our own models and take the infrastructure multi-region as we expand globally.

Expect a lot of breadth, with depth in a few things that matter. You won't be a specialist hiding behind a narrow remit — you'll move across AWS, CI/CD, databases, CDN, and AI infra, and go deep where it counts.

What you'll do

  • Own our AWS infrastructure through Terraform — ECS on EC2, VPCs, IAM, autoscaling, rolling deploys. Clickops-free console as the target.
  • Build Jenkins pipelines that are fast, safe, and self-serve. Be in the deploy paths, not just the platform.
  • Drive cost optimisation.
  • Own media delivery at scale — Cloudflare CDN, image/video/audio transcoding, format/device delivery — tuned for latency, cache hit ratio, and egress cost.
  • Operate the self-managed data layer with service teams — RDS MySQL, MongoDB, Valkey, plus managed MSK and ScyllaDB Cloud. Lead migrations end-to-end (e.g. the ongoing Redis → Valkey).
  • Stand up the operational substrate for AI in production — GPU capacity, inference gateways, key rotation, rate limits, cost-per-request visibility. Build observability for LLM agents: traces, token/cost accounting, eval hooks, alerts on silent regressions.
  • Partner with the AI/ML team on self-hosting models — capacity planning, vLLM / TGI serving, canary rollouts, and cost/latency trade-offs vs managed providers.
  • Own reliability across services and infra — SLOs, alerting, incident response, blameless postmortems. Move teams from firefighting to proactive reliability.
  • Once you have the context, add the guardrails — IaC checks, deploy gates, paved paths — that make the right thing the easy thing and quietly remove whole classes of human error.

What we're looking for

  • 4–6 years in DevOps, SRE, or infrastructure — production systems at meaningful scale, with ownership beyond tickets.
  • A problem solver with high agency. You reason from first principles, don't wait to be told, and dig in rather than deflect — whether it's a developer stuck on Terraform or an ML engineer asking for GPUs.
  • Strong AWS hands-on — ECS on EC2, VPCs, IAM, multi-AZ design — and proficiency with Terraform (you write modules others reuse).
  • Jenkins in production plus a real sense of developer experience in CI/CD — you've looked at deploy-time metrics and changed them.
  • Python, Ansible, and shell — you automate work rather than repeat it.
  • Operated databases in production — at least some of RDS MySQL, MongoDB, Redis/Valkey. Done migrations, failovers, and perf tuning, not just provisioning. Working familiarity with Kafka and Cassandra/ScyllaDB.
  • Strong grasp of Linux internals and networking fundamentals (TCP/IP, DNS, TLS, load balancing) and common failure modes.
  • Some hands-on exposure to LLM-based systems — inference endpoints, agent tracing, token/cost, or evals in CI. Not a researcher; you should reason clearly about latency, cost, and failure modes of LLM workloads.
  • Comfortable with cloud security fundamentals: IAM least-privilege, secrets management, network segmentation. Prior exposure to ISO 27001 / DPDP-style controls is a plus.

Tech stack
AWS · ECS (on EC2) · Terraform · Jenkins · Ansible · Python · Cloudflare · Prometheus · Grafana · InfluxDB · RDS MySQL · MongoDB · Valkey · MSK · ScyllaDB Cloud · growing: GPU inference · vLLM / TGI · OpenTelemetry GenAI · Langfuse-style tracing · multi-region AWS

Security & Data Handling
All employees are expected to handle sensitive data responsibly in compliance with the DPDP Act, ISO-27001:2022, and Pratilipi's internal security policies — ensuring data privacy, confidentiality, and NDA obligations at all times.

About the company

Pratilipi company logo

Pratilipi

Actively Hiring
The largest Indian language story telling platform201-500 Employees
Company Size
201-500
Company Type
Consumer Technology
Company Type
Internet
Company Type
Information Technology
Company Type
Content Publishing
Company Industries
Reading Apps
  • Scale Stage
    Rapidly increasing operations
  • Top Investors
    This company has received a significant amount of investment from top investors
Learn more about Pratilipi image

Funding

AMOUNT RAISED
$78.8M
FUNDED OVER
5 rounds
Rounds
D
$48000000
Series D - Jul 2021+4

Perks

Medical insurance for Fulltime employees & Interns
Mental health consultation
Every employee can avail the benefit of having a confidential mental health consultation with an in-house clinical psychologist & we also provide on-call therapist support.
Maternity & Paternity leaves
Participate in our ESOP and become a stakeholder in the company's success.
Breakfast
Pet-friendly office
We actively invest in Learning & Development

Founders

sahradayi modi
Head of product • 11 years
Bengaluru
image
Sankaranarayanan Devarajan
Head - Language Operations • 12 years
Bengaluru
image
Ranjeet Pratap Singh
CEO • 12 years
Bengaluru
image
View the team image

Similar Jobs

Sosuv Consulting company logo
Sosuv Consulting
We provide specialised Product, Software Development and Support services to the global fi
TopGrep Tech company logo
TopGrep Tech
AI Enabled Platform for STEM Education in Quality Engineering
Flex.ai company logo
Flex.ai
The AI Infrastructure company
AuxoAI company logo
AuxoAI
We help companies—turn their strategies into practical digital and AI solutions
Jumeau Capital company logo
Jumeau Capital
Building developer infrastructure in the API economy. Real product. Real revenue. Lean team
Planso company logo
Planso
we are a early stage startup looking to build AI agents for enterprise grade applications
Astra Security company logo
Astra Security
Uncover & fix vulnerabilities in record time with Astra's Pentest Platform