
Sr. Data Engineer
- ₹40L – ₹60L • No equity
- |
- |5 years of exp
- |Full Time
In office
Not Available
About the job
Roles and Responsibilities
• Own the real-time path — Steward the real-time path (Kafka into ClickHouse) that powers low-latency analytics, overage detection, and operational dashboards. Tune brokers, partitions, materialised views, and consumer offsets.
• Own the batch path — Steward the batch path (Spark into Snowflake) as a durable, replayable backup of the usage ledger and the source for warehouse loads.
• CDC pipeline — Own the AWS DMS flows from Aurora MySQL / Postgres into the warehouse. Make sure every operational change is captured, durable, and joinable with traffic data.
• Airflow and orchestration — Build and maintain the DAGs that produce semantic-layer tables, internal BI fact tables, customer-facing analytics aggregates, and search/ranking signals. Keep DAG failures rare, observable, and easy to retry.
• Quota and ledger correctness — Partner with the platform team on the Kinesis → Lambda → DynamoDB → Redis quota loop and the ClickHouse-backed Quota Service. The usage ledger has to reconcile across paths.
• Schema and contract discipline — Own event schemas, evolve them safely, and keep producer/consumer contracts honest as the platform changes.
• Pipeline scalability — Scale the pipeline ahead of traffic: Kafka partition strategy, ClickHouse cluster sizing, Airflow concurrency, and Spark job tuning. The pipeline should get cheaper per event as volume grows, not more fragile.
• Cost discipline — Own the data-platform spend as a first-class metric. Tune storage tiers, retention, partitioning, and compute — Snowflake, ClickHouse, storage, and streaming throughput are all on your scoreboard, reported monthly.
• Agent-native data foundations — Help define the metering and governance signals needed to bill and observe agent traffic distinctly from human traffic.
What You Bring
• 5–8 years building production data systems; multiple years owning a pipeline that revenue depends on.
• Deep hands-on experience with most of: Kafka / MSK, ClickHouse, Spark, Snowflake, Airflow, AWS DMS, S3, Parquet, DynamoDB.
• Strong SQL, comfort tuning queries on a columnar OLAP engine (ClickHouse) and on a cloud warehouse (Snowflake).
• Proficient in Python for data work (PySpark, Airflow operators, scripting); bonus for Scala, Go, or Node.js.
• Hands-on AWS at scale: EKS, Lambda, Kinesis, IAM, S3 lifecycle, Aurora, Elasticache.
• Terraform / IaC for data infrastructure.
• Have lived through the real-time vs batch trade-off — and have an opinion on when each is the right answer.
• Understand CDC pipelines (DMS, Debezium, or similar) and the schema-evolution headaches that come with them.
• Bonus: experience with semantic layers (dbt, Cube, LookML), event-schema registries, billing/metering correctness, or agentic / LLM-driven workloads.
• You write runbooks. You add observability before you need it. You believe the ledger is sacred.
About the company

Jumeau Capital
Similar Jobs









