Avatar for Tredence Analytics Solutions
Connect the dots

Sr. Data Engineer

Posted: 2 months ago
Job Location
Remote Work Policy

In office

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
SQL
WorkFlows
Pyspark
Airflow
Star Schema
Azure Data Factory
DLT
DBT
CI/CD
Delta Lake
Great Expectations
Unity Catalog
Medallion Architecture
Repos
Databricks DLT
Databricks Notebooks
Databricks Jobs
Git Integration
Slowly Changing Dimensions (SCD)
DBFS
Databricks Core Lakehouse

About the job

Role description

About Tredence¬: -

Tredence focuses on last mile delivery of insights into actions by uniting its strengths in business analytics, data science, and software engineering. The largest companies across industries are engaging with Tredence and deploying its prediction and optimization solutions at scale –empowering end users to improve decision making. Headquartered in the San Francisco Bay Area, the company serves clients in the US, Canada, Europe, and SE Asia. Learn more at www.tredence.com

Job Summary

We are seeking an experienced Databricks Architect to design, build, and optimize our modern Lakehouse platform. You will leverage the full Databricks ecosystem — including Delta Lake, Lakehouse architecture, and unified data governance — to deliver scalable, high-performance data solutions. You will be responsible for data modeling, quality enforcement, and ETL/ELT pipelines using PySpark, Python, and SQL.

Key Responsibilities

Databricks Platform & Architecture

Architect and implement end-to-end Lakehouse solutions on Databricks.

Design and optimize Delta Lake storage, including ACID transactions, time travel, and schema evolution.

Configure clustering, partitioning, Z-order, and vacuuming for performance tuning.

Data Modeling & Governance

Build data models (Kimball, Inmon, Data Vault, or medallion architecture: bronze/silver/gold).

Implement Data Catalog (Unity Catalog) for metadata management, lineage, and access control.

Define and enforce Data Quality rules using Great Expectations, DBT, or Databricks DLT (Delta Live Tables).

Development & Pipelines

Develop scalable ETL/ELT pipelines using PySpark, Python, and SQL.

Optimize Spark jobs for performance, cost, and reliability.

Automate workflows with Databricks Jobs, workflows, and orchestration tools (Airflow, Azure Data Factory, etc.).

Collaboration & Best Practices

Partner with data scientists, analysts, and business stakeholders to understand data requirements.

Establish CI/CD for Databricks notebooks and repositories (DBFS, Repos, Git integration).

Monitor and troubleshoot pipeline failures, data drift, and performance bottlenecks.

Required Qualifications

Area Skills

Databricks Core Lakehouse, Delta Lake, Unity Catalog, DLT, Workflows

Languages PySpark, Python, SQL

Data Modeling Star schema, slowly changing dimensions (SCD), medallion architecture

Data Quality Validation, anomaly detection, DQ rules implementation

Catalog & Governance Unity Catalog, Hive Metastore, lineage, access controls

Experience 5+ years in data engineering; 2+ years hands-on with Databricks

Preferred Qualifications

Databricks Certification (e.g., Data Engineer Professional or Associate).

Experience with cloud platforms (AWS S3/Glue, Azure Data Lake/ADF, GCP).

Knowledge of streaming (Kafka, Kinesis, Structured Streaming).

Familiarity with DBT, Terraform, or MLflow.

Soft Skills

Strong analytical and problem-solving abilities.

Excellent communication for technical and non-technical audiences.

Self-starter capable of leading architectural decisions.

About the company

Funding

AMOUNT RAISED
$175M
FUNDED OVER
1 round
Round
B
$175000000
Series B - Jan 2023