Avatar for Century Health
Accelerating access to breakthrough treatments, with the power of AI

Intermediate Data Engineer

  • ₹12L – ₹18L • 0.01% – 0.1%
  • |Remote (
    Everywhere
    )
  • |3 years of exp
  • |Full Time
Posted: 4 months ago
Hires remotely in
Everywhere
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

Preferred Timezones
Eastern Time
Collaboration Hours
1:30 AM - 10:30 AM Eastern Time
RelocationNot Allowed
Skills
Python
SQL
PostgreSQL
Spark
AWS
DBT

About the job

Century Health is a clinical data intelligence company turning messy real-world clinical data into structured, analysis-ready insights for researchers and pharma teams.


The Role

We're looking for an Intermediate Data Engineer (3–5 years) to be a core builder on our data platform — designing ETL pipelines, profiling raw clinical datasets, and ensuring data flowing through our systems is clean and reliable.


What You'll Do

  • Design, build, and maintain ETL/ELT pipelines ingesting data from CSVs, Parquet, XLSX, APIs, and databases
  • Optimize pipeline performance — tune queries, manage compute, reduce latency and cost
  • Profile raw datasets to identify quality issues: missing values, duplicates, schema drift, outliers
  • Build and maintain dbt models for clean, documented, analysis-ready data layers
  • Orchestrate workflows using Apache Airflow on AWS MWAA
  • Process large-scale data using PySpark on EMR
  • Collaborate with ML engineers and GTM teams on downstream use cases
  • Document pipelines, models, assumptions, and known issues clearly

What We're Looking For

Must-Have

  • 3–5 years of professional data engineering experience
  • Strong SQL — window functions, CTEs, query optimization
  • Solid Python — clean, modular, production-grade code
  • Hands-on PySpark for large-scale processing
  • Experience with dbt, Snowflake, and cloud-based ETL/ELT
  • Familiarity with AWS: S3, MWAA, ECS Fargate, EMR, RDS, Bedrock
  • Strong data intuition and ability to work independently in ambiguous environments
  • Experience with test suites — pytest, Great Expectations, dbt tests
  • Active use of AI coding tools (Cursor, Claude Code, etc.)

Nice to Have

  • Healthcare data experience (HIPAA, FHIR, HL7, OMOP)
  • Data science background (statistical modeling, ML pipelines, feature engineering)
  • DevOps exposure (CI/CD, Docker, Terraform/CDK)

Tech Stack

  • Cloud: AWS (MWAA, ECS Fargate, EMR, RDS, S3, Bedrock)
  • Data Warehouse: Snowflake
  • Transformation: dbt
  • Processing: PySpark, Python
  • Orchestration: Apache Airflow (MWAA)
  • AI Tools: Cursor, Claude Code

Hiring Process

  1. Application Review
  2. Take-Home Assignment (~3 hours, real clinical data)
  3. Technical Interview
  4. Managerial Interview

Why Century Health

  • Hard data problems with real clinical impact
  • Small, high-ownership team — your work ships and matters
  • Modern cloud-native stack with architectural influence
  • AI-first engineering culture

About the company

Century Health company logo
Accelerating access to breakthrough treatments, with the power of AI1-10 Employees
Learn more about Century Health image

Founders

Sanjay Hariharan
Founder
image
View the team image