Avatar for GenebyGene.com
GenebyGene.com
Actively Hiring
Genetic testing for ancestry, health, wellness, and COVID-19 tests

Senior Data Pipeline Engineer/Developer

Posted: 3 days ago• Recruiter recently active
Job Location
Remote Work Policy

In office

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
SQL
PostgreSQL
Unit Tests
Batch
Genomics
Snowflake
Microsoft SQL Server
Kafka
EMR
Spark
LAMBDA
Docker
S3
Kubernetes
Terraform
Redshift
Kinesis
CDK
Airflow
Infrastructure as code
Cloudformation
Integration tests
DBT
Nextflow
VCF
Snakemake
Pulumi
Prefect
Dagster
Argo Workflows
Glue
Step Functions
Fastq
AWS Data Stack
Contract Tests
BAM/CRAM
Data-Quality Tests
WDL/Cromwell
Workflow Orchestrator
NGS Data Formats
Bioinformatics Workflow Engines
AWS HealthOmics

About the job

Our Purpose

Our mission is to build a healthier and more connected world with precision health and genealogy services. We empower individuals with actionable insights into their genetic makeup, fostering a deeper understanding of their ancestry, health, and wellness. By integrating the experience of Gene by Gene Laboratory Services, FamilyTreeDNA genealogy, and myDNA reporting services, we strive to deliver cutting-edge genetic testing and personalized solutions that inspire informed decisions and enhance quality of life. Our team is dedicated to advancing the field of genomics through innovation, research, and a commitment to excellence.

Our Values

All employees are expected to demonstrate our values of Innovate, One Team, and Integrity when carrying out the accountabilities and responsibilities of their role. This is how we show up every day for ourselves, our colleagues and our customers and strategic partners to deliver our vision and strategic goals.

Position Overview

We are seeking a Senior Data Pipeline Engineer/Developer to design, build, and own the data pipelines that move genomic and operational data from lab instruments and LIMS through to analytics, products, and clinical reporting. In this senior individual-contributor role you will set technical direction for our data platform, build production pipelines that are reliable and reproducible at scale, and provide architectural and code-level guidance to existing engineering teams. You will own data quality, lineage, and observability end to end, and operate in a regulated environment where reproducibility and auditability are non-negotiable. This is a Python-first, full-stack engineering role. This is a contractor position, slated for a 12-month term.

Accountabilities and Responsibilities

Pipeline Engineering & Integration

  • Builds and operates production ETL and ELT for high-volume genomic data as well as operational and business data.
  • Integrates data across LIMS, lab instruments, internal applications, and third-party sources through robust, well-tested interfaces.
  • Develops across the stack in Python (data services, internal APIs, and supporting application code) and provides architectural and code-level guidance to engineering teams on data-layer integration.

Architecture & Technical Leadership

  • Sets technical direction for batch and streaming data pipelines by evaluating frameworks, orchestration, storage, and processing patterns, making recommendations, and leading adoption.
  • Models and tunes data stores (Microsoft SQL Server and PostgreSQL, plus a cloud warehouse or lake) for performance and scale.
  • Defines and enforces engineering standards for testing, CI/CD, infrastructure as code, code review, and architecture decision records.

Data Quality & Pipeline Observability

  • Owns data quality, lineage, and observability, including freshness, completeness, schema-drift detection, cost-per-job, and SLAs.
  • Builds pipelines for reproducibility and audit-readiness, incorporating versioned data and code, lineage, decision logging, access controls, and evidence collection.
  • Partners with security and compliance on data privacy, PII and PHI handling, and regulatory requirements across the data lifecycle.

Regulatory Compliance and Security Governance

  • Apply deep understanding of CAP/CLIA, HIPAA, GDPR, and GxP regulations specifically to data pipeline architecture.
  • Oversee protected health information (PHI) handling, data lineage, retention, and comprehensive audit logging.
  • Produce and maintain the critical operational evidence required to carry the data platform successfully through compliance audits.
  • Adhere to strict data-governance controls necessary for securely handling sensitive genomic data across global regions.
  • Enforce least-privilege access and ensure zero data egress to personal or unapproved infrastructure.
  • Utilize exclusively de-identified or synthetic data within development environments.
  • Maintain strict compliance with international data-residency and localization requirements.

Position Requirements

Skills and Knowledge

  • Strong SQL proficiency on Microsoft SQL Server and PostgreSQL, including schema design, query tuning, and performance troubleshooting.
  • Strong Python proficiency across the stack (data pipelines, backend services, and APIs). This is a Python-first role.
  • Production experience with a workflow orchestrator (Argo Workflows, Step Functions, Prefect, Dagster, Airflow, or similar).
  • Production experience with containers (Docker) and Kubernetes, which the lead orchestrator (Argo Workflows) runs on.
  • Strong testing discipline, including unit, integration, and data-quality or contract tests.
  • Excellent written and verbal communication, including the ability to explain data tradeoffs to technical and compliance stakeholders.
  • Genomics or NGS data formats and handling (FASTQ, BAM/CRAM, VCF).
  • Bioinformatics workflow engines (Nextflow, WDL/Cromwell, or Snakemake).
  • AWS data stack proficiency (S3, Glue, EMR, Batch, Lambda, Redshift) and/or AWS HealthOmics.
  • Distributed processing and streaming frameworks (Spark, Kafka, Kinesis).
  • Data warehouse and transformation tooling (Snowflake, dbt).
  • Infrastructure as code practices (Terraform, CDK, CloudFormation, or Pulumi).
  • Familiarity with .NET (C#) and/or C++ as supporting languages for integrating with existing services.
  • Knowledge of LIMS integration and on-premises plus hybrid data architectures.

Experience

  • 8+ years of professional software or data engineering experience.
  • 4+ years building and operating production data pipelines at scale.
  • 3+ years of production cloud experience (AWS preferred).
  • Demonstrable experience working in regulated environments. Compliance is a hard requirement for this role.
  • Experience supporting HIPAA or GxP audits.

Education

  • Bachelor's degree in Computer Science, Bioinformatics, or a related field, or equivalent professional experience.
  • An advanced degree in a quantitative field is preferred.

Why Join Us

At Gene by Gene, you’ll join a mission-driven team advancing the science of genetics and discovery. You’ll have the opportunity to shape meaningful campaigns, tell compelling brand stories, and collaborate with talented professionals who share your passion for creativity, curiosity, and impact.

About the company

GenebyGene.com company logo

GenebyGene.com

Actively Hiring
Genetic testing for ancestry, health, wellness, and COVID-19 tests 51-200 Employees
Company Size
51-200
Company Industries
Genealogy Market Worldwide
Learn more about GenebyGene.com image

Founders

Max Blankfeld
Founder
Houston
image
View the team image

Similar Jobs

Archesys company logo
Archesys
Improving the government services that impact everyday lives
Archesys company logo
Archesys
Improving the government services that impact everyday lives
Archesys company logo
Archesys
Improving the government services that impact everyday lives