Intermediate Data Engineer
- ₹12L – ₹18L • 0.01% – 0.1%
- |Remote (Everywhere)
- |3 years of exp
- |Full Time
Posted: 4 months ago
Hires remotely in
Everywhere
Remote Work Policy
Remote only
Company Location
Visa Sponsorship
Not Available
Preferred Timezones
Eastern Time
Collaboration Hours
1:30 AM - 10:30 AM Eastern Time
RelocationNot Allowed
Skills
Python
SQL
PostgreSQL
Spark
AWS
DBT
About the job
Century Health is a clinical data intelligence company turning messy real-world clinical data into structured, analysis-ready insights for researchers and pharma teams.
The Role
We're looking for an Intermediate Data Engineer (3–5 years) to be a core builder on our data platform — designing ETL pipelines, profiling raw clinical datasets, and ensuring data flowing through our systems is clean and reliable.
What You'll Do
- Design, build, and maintain ETL/ELT pipelines ingesting data from CSVs, Parquet, XLSX, APIs, and databases
- Optimize pipeline performance — tune queries, manage compute, reduce latency and cost
- Profile raw datasets to identify quality issues: missing values, duplicates, schema drift, outliers
- Build and maintain dbt models for clean, documented, analysis-ready data layers
- Orchestrate workflows using Apache Airflow on AWS MWAA
- Process large-scale data using PySpark on EMR
- Collaborate with ML engineers and GTM teams on downstream use cases
- Document pipelines, models, assumptions, and known issues clearly
What We're Looking For
Must-Have
- 3–5 years of professional data engineering experience
- Strong SQL — window functions, CTEs, query optimization
- Solid Python — clean, modular, production-grade code
- Hands-on PySpark for large-scale processing
- Experience with dbt, Snowflake, and cloud-based ETL/ELT
- Familiarity with AWS: S3, MWAA, ECS Fargate, EMR, RDS, Bedrock
- Strong data intuition and ability to work independently in ambiguous environments
- Experience with test suites — pytest, Great Expectations, dbt tests
- Active use of AI coding tools (Cursor, Claude Code, etc.)
Nice to Have
- Healthcare data experience (HIPAA, FHIR, HL7, OMOP)
- Data science background (statistical modeling, ML pipelines, feature engineering)
- DevOps exposure (CI/CD, Docker, Terraform/CDK)
Tech Stack
- Cloud: AWS (MWAA, ECS Fargate, EMR, RDS, S3, Bedrock)
- Data Warehouse: Snowflake
- Transformation: dbt
- Processing: PySpark, Python
- Orchestration: Apache Airflow (MWAA)
- AI Tools: Cursor, Claude Code
Hiring Process
- Application Review
- Take-Home Assignment (~3 hours, real clinical data)
- Technical Interview
- Managerial Interview
Why Century Health
- Hard data problems with real clinical impact
- Small, high-ownership team — your work ships and matters
- Modern cloud-native stack with architectural influence
- AI-first engineering culture
Similar Jobs

Cypherock Wallet
Personal Fort Knox for Your Crypto

Infinite Analytics
Consumer Insights AI Platform

Sciative - We Price Right
Price Optimization SaaS platform

EveoAI
A novel approach to managing fashion, style, and personality through Generative AI

Zintellix
AI Product Studio

Carbon Trail
AI powered sustainability platform for fashion and retail industry

Khushi Baby
Last Mile Wearable Medical Passport

Sciative - We Price Right
Price Optimization SaaS platform

NoScrubs
Lightning fast laundry delivery service

