Avatar for Decentro
Decentro
Actively Hiring
Fintech Infrastructure Platform
  • Growth Stage
    Expanding market presence
  • Top Investors
    This company has received a significant amount of investment from top investors
  • YC Funded
    Startup funded by Y Combinator

Site Reliability Engineer

  • ₹20L – ₹30L
  • |
  • |3 years of exp
  • |Full Time
Posted: 2 months ago• Recruiter recently active
Job Location
Remote Work Policy

In office - WFH flexibility

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
Site Reliability Engineering
DataDog
Grafana
CloudWatch

About the job

About Decentro

Decentro is a Y-Combinator backed banking & payments infrastructure company. Decentro provides building blocks that help companies stitch their fintech workflows in a few weeks. While starting our first fintech venture back in 2014, we spent years convincing banks to partner with us. Before we could launch our fintech product (more here), we had to convince different departments within banks - including technology, business, leadership, legal, and support to get the green signal. Fast forward to today, we realise that banks & regulatory institutions find it difficult to help the longer tail of companies build innovative & compliant fintech solutions.

What if there was a platform (think AWS for cloud or Twilio for messaging) that abstracts the complexities of banking, identity, payments, lending, and provides simple APIs so that companies do not have to spend years dealing with banks?

We’re solving this pain point at Decentro since 2020. We’ve scaled to process $4 billion + in payment volumes & have 1400+ customers across India & Singapore in multiple verticals such as marketplaces, banks, regulated institutions, fintechs, lenders, gaming, and more.

About the role

We are looking for a highly motivated and experienced Site Reliability Engineer to join our engineering team. This role will focus on improving the reliability, visibility and operational excellence of our production systems. You will work closely with the SRE and engineering teams to build comprehensive observability and SLO coverage across our products and infrastructure.

What is expected from you -

  1. Operate, stabilise and continuously improve our production observability stack, with hands-on experience in Prometheus, Grafana and related monitoring systems.

  2. Implement and maintain SLIs, SLOs, error budgets, recording rules, dashboards and burn-rate alerts for critical user journeys and services.

  3. Drive OpenTelemetry adoption across product teams, including service instrumentation, telemetry pipelines and correlation across metrics, logs and traces.

  4. Operate and troubleshoot observability workloads on Kubernetes, including stateful, multi-node and highly available production workloads.

  5. Own PromQL and Prometheus-based monitoring, including query optimisation, recording rules, metric types, histograms and identifying high-cardinality issues.

  6. Own alert quality and incident response, improving alerting, routing and escalation while reducing noisy and non-actionable alerts.

  7. Manage and optimise the AWS telemetry surface across CloudWatch, ALB, WAF, VPC Flow Logs and ECS, including multiple accounts and regions.

  8. Build automation and tooling using Python, including custom exporters, log-processing Lambdas, SLI computation jobs and other data-processing workflows.

  9. Improve telemetry efficiency and cost through cardinality management, retention, sampling and data-volume optimisation, while ensuring appropriate security and compliance for sensitive data.

  10. Work closely with engineering and SRE teams to audit and improve observability coverage across services and infrastructure, contribute to the migration towards a highly available LGTM stack (Loki, Grafana, Tempo and Mimir) with Prometheus on Kubernetes, and continuously improve production reliability based on operational feedback.

What we are looking for -

  1. 3+ years of hands-on experience in SRE, Observability, Production Engineering or Reliability Engineering, with 4-5 years of total engineering experience.

  2. Strong production experience with Prometheus and Grafana, including PromQL, recording rules, dashboards, alerting and troubleshooting monitoring issues.

  3. Hands-on experience operating observability workloads on Kubernetes, including stateful, multi-node and highly available production environments.

  4. Practical experience with OpenTelemetry, including SDKs, collectors, service instrumentation and telemetry pipelines across metrics, logs and traces.

  5. Proven experience implementing SLIs, SLOs, error budgets and burn-rate based alerting for production systems.

  6. Strong understanding of observability best practices, including alert quality, incident management, cardinality management, retention, sampling and telemetry cost optimisation.

  7. Hands-on experience with AWS observability and monitoring, particularly CloudWatch, ALB, WAF, VPC Flow Logs and ECS, preferably across multiple accounts and regions.

  8. Strong Python scripting and automation skills, with experience building exporters, pipelines, Lambda functions or data-processing workflows.

  9. Experience with Terraform and Helm, along with exposure to LGTM components, continuous profiling (Pyroscope), eBPF-based observability (Beyla/Alloy), SQL and analytical stores (Athena, Doris, Iceberg), frontend/RUM observability, or load/chaos testing is a strong plus.

  10. Experience with live system/platform migrations, observability checks integrated into CI/CD (Jenkins), and Payments / Fintech / Banking is a plus. Strong problem-solving, ownership and communication skills, with the ability to collaborate effectively with engineering teams, are essential.

Why join Decentro

You will work on high-scale payment infrastructure, solve real production reliability challenges, and contribute to systems that power critical financial workflows. This is an opportunity to build strong platform and reliability depth while working with a fast-moving team in one of the most interesting segments of fintech.

What We Offer

The ability for you to make an impact and lay a foundation for the upcoming fin-tech innovations.
A multicultural and diverse team of colleagues from different states that speak in total of 6 Indian and global languages.
Progressive and flexible work hours that match your personality and lifestyle.

The best-in-class perks and benefits for the team. Check out our careers page for the same: https://decentro.tech/careers/https://decentro.tech/careers/

Backed by global investors such as Ycombinator & Rapyd, we're a contrarian and progressive culture of independent thinkers and systematic executors that are driven to build cool things that matter.

If this aligns with you, time to hop on!

We are an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, colour, religion, sex, disability status, or any other characteristic protected by law.

About the company

Decentro company logo

Decentro

Actively Hiring
Fintech Infrastructure Platform51-200 Employees
  • Growth Stage
    Expanding market presence
  • Top Investors
    This company has received a significant amount of investment from top investors
  • YC Funded
    Startup funded by Y Combinator
Learn more about Decentro image

Funding

AMOUNT RAISED
$9.9M
FUNDED OVER
4 rounds
Rounds
B
$3600000
Series B - Jun 2025+3

Perks

Health Insurance with good coverage
ESOPs that are actually put in place.
Work from home friendly
Generous vacation
Bring along those cute little friends :)
Latest devices & gadgets

Similar Jobs

EquityList company logo
EquityList
Cap table and equity grant automation for global businesses
Ratch AI company logo
Ratch AI
Ratch: evidence-based hiring. Real work. Clear trade-offs. You only pay when you hire
Pilot (pilotplans.com) company logo
Pilot (pilotplans.com)
Multiplayer consumer AI for travel where groups plan, decide, and book together
Guile company logo
Guile
Autonomous AI agents that find, chain, and prove real exploit paths in your software