
- Growth StageExpanding market presence
- Top InvestorsThis company has received a significant amount of investment from top investors
- YC FundedStartup funded by Y Combinator
Site Reliability Engineer
- ₹20L – ₹30L
- |
- |3 years of exp
- |Full Time
In office - WFH flexibility
Not Available
About the job
About Decentro
Decentro is a Y-Combinator backed banking & payments infrastructure company. Decentro provides building blocks that help companies stitch their fintech workflows in a few weeks. While starting our first fintech venture back in 2014, we spent years convincing banks to partner with us. Before we could launch our fintech product (more here), we had to convince different departments within banks - including technology, business, leadership, legal, and support to get the green signal. Fast forward to today, we realise that banks & regulatory institutions find it difficult to help the longer tail of companies build innovative & compliant fintech solutions.
What if there was a platform (think AWS for cloud or Twilio for messaging) that abstracts the complexities of banking, identity, payments, lending, and provides simple APIs so that companies do not have to spend years dealing with banks?
We’re solving this pain point at Decentro since 2020. We’ve scaled to process $4 billion + in payment volumes & have 1400+ customers across India & Singapore in multiple verticals such as marketplaces, banks, regulated institutions, fintechs, lenders, gaming, and more.
About the role
We are looking for a highly motivated and experienced Site Reliability Engineer to join our engineering team. This role will focus on improving the reliability, visibility and operational excellence of our production systems. You will work closely with the SRE and engineering teams to build comprehensive observability and SLO coverage across our products and infrastructure.
What is expected from you -
Operate, stabilise and continuously improve our production observability stack, with hands-on experience in Prometheus, Grafana and related monitoring systems.
Implement and maintain SLIs, SLOs, error budgets, recording rules, dashboards and burn-rate alerts for critical user journeys and services.
Drive OpenTelemetry adoption across product teams, including service instrumentation, telemetry pipelines and correlation across metrics, logs and traces.
Operate and troubleshoot observability workloads on Kubernetes, including stateful, multi-node and highly available production workloads.
Own PromQL and Prometheus-based monitoring, including query optimisation, recording rules, metric types, histograms and identifying high-cardinality issues.
Own alert quality and incident response, improving alerting, routing and escalation while reducing noisy and non-actionable alerts.
Manage and optimise the AWS telemetry surface across CloudWatch, ALB, WAF, VPC Flow Logs and ECS, including multiple accounts and regions.
Build automation and tooling using Python, including custom exporters, log-processing Lambdas, SLI computation jobs and other data-processing workflows.
Improve telemetry efficiency and cost through cardinality management, retention, sampling and data-volume optimisation, while ensuring appropriate security and compliance for sensitive data.
Work closely with engineering and SRE teams to audit and improve observability coverage across services and infrastructure, contribute to the migration towards a highly available LGTM stack (Loki, Grafana, Tempo and Mimir) with Prometheus on Kubernetes, and continuously improve production reliability based on operational feedback.
What we are looking for -
3+ years of hands-on experience in SRE, Observability, Production Engineering or Reliability Engineering, with 4-5 years of total engineering experience.
Strong production experience with Prometheus and Grafana, including PromQL, recording rules, dashboards, alerting and troubleshooting monitoring issues.
Hands-on experience operating observability workloads on Kubernetes, including stateful, multi-node and highly available production environments.
Practical experience with OpenTelemetry, including SDKs, collectors, service instrumentation and telemetry pipelines across metrics, logs and traces.
Proven experience implementing SLIs, SLOs, error budgets and burn-rate based alerting for production systems.
Strong understanding of observability best practices, including alert quality, incident management, cardinality management, retention, sampling and telemetry cost optimisation.
Hands-on experience with AWS observability and monitoring, particularly CloudWatch, ALB, WAF, VPC Flow Logs and ECS, preferably across multiple accounts and regions.
Strong Python scripting and automation skills, with experience building exporters, pipelines, Lambda functions or data-processing workflows.
Experience with Terraform and Helm, along with exposure to LGTM components, continuous profiling (Pyroscope), eBPF-based observability (Beyla/Alloy), SQL and analytical stores (Athena, Doris, Iceberg), frontend/RUM observability, or load/chaos testing is a strong plus.
Experience with live system/platform migrations, observability checks integrated into CI/CD (Jenkins), and Payments / Fintech / Banking is a plus. Strong problem-solving, ownership and communication skills, with the ability to collaborate effectively with engineering teams, are essential.
Why join Decentro
You will work on high-scale payment infrastructure, solve real production reliability challenges, and contribute to systems that power critical financial workflows. This is an opportunity to build strong platform and reliability depth while working with a fast-moving team in one of the most interesting segments of fintech.
What We Offer
The ability for you to make an impact and lay a foundation for the upcoming fin-tech innovations.
A multicultural and diverse team of colleagues from different states that speak in total of 6 Indian and global languages.
Progressive and flexible work hours that match your personality and lifestyle.
The best-in-class perks and benefits for the team. Check out our careers page for the same: https://decentro.tech/careers/https://decentro.tech/careers/
Backed by global investors such as Ycombinator & Rapyd, we're a contrarian and progressive culture of independent thinkers and systematic executors that are driven to build cool things that matter.
If this aligns with you, time to hop on!
We are an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, colour, religion, sex, disability status, or any other characteristic protected by law.
About the company
- Growth StageExpanding market presence
- Top InvestorsThis company has received a significant amount of investment from top investors
- YC FundedStartup funded by Y Combinator
Employees joined from
Similar Jobs




