Data Engineer

  • ₹30L – ₹50L
  • |Remote (
    Everywhere
    )
  • |8 years of exp
  • |Full Time
Posted: today• Recruiter recently active
Hires remotely in
Everywhere
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

RelocationNot Allowed
Skills
Python
SQL
Big Data
Gdk
Pyspark
Dataproc
GCP
databricks
CI/CD
ETL/ELT
Delta Lake
Dataflow
Generative AI
LLMs
RAG

About the job

6+ years in data engineering, ETL/ELT, analytics engineering, data warehouse development, or big data pipelines; minimum 4 years on GCP cloud data platforms preferred.
Strong SQL and Python.
Data modelling, warehousing concepts, partitions, clustering, indexing, optimisation.
ETL/ELT pipeline design and production support.
Batch and streaming pipeline fundamentals.
Orchestration using Airflow, Composer, Control-M, ADF, dbt, or equivalent.
Big data processing using Spark, PySpark, Beam, Dataflow, Dataproc, Databricks, EMR, or equivalent.
Cloud storage and data lake concepts.
Data quality, reconciliation, schema validation, lineage, monitoring, and alerting.
CI/CD for data pipelines and code versioning.
Ability to explain scale scenarios, not just tool definitions.
Generic Skills (Must Have)
Python and advanced SQL.
Data quality checks, reconciliation, alerting, retries, and backfills.
Production pipeline monitoring and troubleshooting.

GCP Skills (Must Have)
Mandatory: BigQuery: datasets, tables, views, partitioning, clustering, SQL optimisation, cost awareness.
Option 1: Dataflow / Apache Beam or equivalent streaming/batch framework.
Option 2: Dataproc / PySpark or strong Spark experience.
Pub/Sub or equivalent event streaming.
Cloud Composer / Airflow DAG development.

Nice to have (Trainable)
Dataform, dbt, CI/CD for data.
BigQuery ML.
Dataflow templates.
Looker / LookML basics.
Terraform for data infrastructure.