Avatar for Hexplora
Hexplora
Actively Hiring
Data platform for healthcare orgs to improve care and cut costs

Senior Data Engineer with Pyspark

Posted: 1 month ago
Job Location
Remote Work Policy

In office

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Data Warehousing
Relational Databases
ETL
Performance Tuning
Azure
Data Modeling
Troubleshooting
SSIS
AWS
Pyspark
GCP
Data Pipeline Monitoring

About the job

Job Title: Senior Data Engineer (PySpark, ETL SSIS) -W2 only

Location: Rocky Hill,CT

Job Description

We are looking for an experienced and motivated Data Engineer with expertise in PySpark to join our dynamic team. As a key member of our data engineering team, you will play a crucial role in designing, building, and maintaining scalable data pipelines that enable efficient data processing and analytics within the healthcare domain.

This role will combine both development and administrative activities, making it essential that the candidate has experience not only in building robust data pipelines but also in overseeing their operational aspects to ensure performance, reliability, and optimization.

Key Responsibilities

  • Design, develop, and maintain scalable ETL pipelines using PySpark to process large datasets.
  • Collaborate with cross-functional teams (data scientists, analysts, business stakeholders) to understand data requirements and deliver high-quality solutions.
  • Work on administrative tasks, including monitoring, troubleshooting, and optimizing data pipelines and infrastructure.
  • Manage data integration across healthcare systems, ensuring compliance with relevant standards.
  • Leverage SSIS for ETL development and ensure smooth data movement across different environments.
  • Integrate and transform data from multiple sources, ensuring data quality and consistency.
  • Handle and resolve data processing issues, ensuring minimal disruption to operations.
  • Document best practices, processes, and workflows to maintain pipeline efficiency and scalability.
  • Work with both relational and non-relational databases, ensuring smooth data flow and optimized performance.

Required Skills & Qualifications

  • Strong experience in Data Engineering, with expertise in designing, building, and maintaining ETL pipelines.
  • Strong proficiency in PySpark for large-scale data processing and transformation.
  • Experience with ETL tools, particularly SSIS (SQL Server Integration Services).
  • Solid understanding of data modeling, relational databases, and data warehousing principles.
  • Experience working with cloud-based data storage and processing technologies (AWS, GCP, or Azure).
  • Familiarity with healthcare data standards, such as HL7 and FHIR, is highly desirable.
  • Proven ability in data pipeline monitoring, troubleshooting, and performance tuning.
  • Strong communication skills and the ability to work collaboratively with cross-functional teams.

Infowave Systems is an equal opportunity employer that is committed to diversity and inclusion in the workplace.

About the company

Hexplora company logo

Hexplora

Actively Hiring
Data platform for healthcare orgs to improve care and cut costs51-200 Employees
Learn more about Hexplora image

Similar Jobs

Astranis company logo
Astranis
Building next-generation internet satellites to get the world online
Neuralink company logo
Neuralink
Ultra-high bandwidth brain-machine interfaces to connect humans and computers
Wynd Labs company logo
Wynd Labs
Making AI Data Accessible. Building a suite of products powered by Grass
EliseAI company logo
EliseAI
Building AI agents that transform complex healthcare and housing systems
Ursa Major company logo
Ursa Major
Powering the future of defense and aerospace. Fly More. Fly Faster
Biofire company logo
Biofire
Biofire: Building the future of firearms safety
Postman company logo
Postman
Postman is the world’s leading collaboration platform for API development