Avatar for SourceIQ
SourceIQ
Actively Hiring
SourceIQ is a cloud-based platform that automates supplier management and sourcing
  • Top 10% of responders
    SourceIQ is in the top 10% of companies in terms of response time to applications
  • Responds within two weeks
    Based on past data, SourceIQ usually responds to incoming applications within two weeks
  • B2B
  • +1

Data Engineering Intern (Python + Databricks + ETL Pipelines)

  • No equity
  • |Remote (
    Austin • 
    +18)
  • |4 years of exp
  • |Internship
Reposted: 1 month ago• Recruiter recently active
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

Preferred Timezones
Central Time, Eastern Time
Collaboration Hours
8:00 AM - 6:00 PM Central Time
RelocationNot Allowed
Skills
Python
Data Mining
Data Analysis
noSQL
MongoDB
Pipeline Management
ETL
Engineering
Entity Resolution/Deduplication
Data Management
ETL Testing
Deduplication
Python Flask
Etl Development
Data Mapping
Elt
Inductive/Deductive Logic
Python/Django/Flask
Machine Learning Data Science Python R
Database: Mysql, Nosql (Mongo Db, H Base, Redis, Postgre Sql, Graph Based: Gremin)
Pyspark
Apache/Spark/Databricks
Machine Learning Data Science Python
NoSQL DBs - MongoDB Cassandra Redis Neo4J
MongoDB
PySpark (Core Spark and Spark SQL)
data deduplication
Databases (SQL and NoSQL)
databricks
Pyspark dataframes
noSQL Databases
Extraction Loading & Transformation (ELT)
ETL/ELT
Apache Spark/Databricks
Python (web Scraping with Beautiful Soup, Numpy, Pandas, Pyspark, Matplotlib, Nltk)
Azure Databricks
Pandas, Numpy, Spacy, NLTK and PySpark, PyTest, Seaborn
Microsoft Azure DataBricks
ELT Processes

About the job

Company: SourceIQ
Location: Remote
Duration: 3-6 months (with potential for conversion to full-time)
Compensation: Unpaid (Academic credit available)
Start Date: Flexible (Rolling applications) Winter

About SourceIQ
SourceIQ is an AI-powered supplier management and sourcing platform transforming how enterprises discover, evaluate, onboard, and collaborate with suppliers. We automate compliance, analytics, and sourcing workflows from end to end, helping companies streamline procurement operations.

We're building the infrastructure layer for enterprise procurement—a platform designed to consolidate a fragmented $50B market by becoming the centralized intelligence and data layer that integrates with all procurement tools.

The Role
As a Data Engineering Intern, you'll build and maintain the data infrastructure that powers SourceIQ's AI models and analytics. You'll work with Databricks, Apache Spark, Python, and Azure to create ETL pipelines, process large datasets, and ensure data quality across our platform.

This is hands-on data engineering where you'll contribute directly to production data pipelines used by enterprise customers including JP Morgan and Humana. You'll work closely with our Head of Technology, ML engineers, and backend team to build the data foundation for our AI-powered platform.

You'll learn modern data engineering practices including lakehouse architecture, Delta Lake, PySpark, and data quality frameworks while working on real business problems.

What You'll Do

  • ETL Pipeline Development (Primary Focus)
  • Build and maintain ETL pipelines using Python and PySpark on Databricks
  • Process supplier data from multiple sources: CSV imports from conferences and events State/local government databases (150K+ suppliers) Association databases ERP systems (Coupa, SAP Ariba, QuickBooks) API integrations (real-time data ingestion)
  • Implement schema mapping to standardize non-standard data formats
  • Create automated data refresh pipelines (e.g., marketplace data updated every 30 days)

Data Quality & Validation

  • Design and implement data quality checks and validation rules
  • Build data profiling and completeness scoring systems
  • Create deduplication logic to identify duplicate suppliers across sources
  • Implement data cleansing (address normalization, phone validation, etc.)
  • Monitor data quality metrics and set up alerting for anomalies

Lakehouse Architecture (Databricks)

  • Work with Delta Lake (Bronze → Silver → Gold layers) Bronze: Raw data ingestion from all sources Silver: Cleaned, deduplicated, standardized data Gold: Aggregated features for ML models and analytics
  • Design and optimize data models for analytics and ML workloads
  • Implement data versioning and time travel for auditability
  • Work with Unity Catalog for data governance and access control

Analytics & Reporting Support

  • Create aggregated data tables for dashboards and BI tools
  • Build data marts for spend analytics, supplier performance, and engagement metrics
  • Optimize query performance for reporting workloads
  • Support ad-hoc data analysis requests from product and business teams

Data Pipeline Orchestration

  • Implement Databricks Jobs for scheduled data processing
  • Build idempotent, fault-tolerant pipelines with retry logic
  • Monitor pipeline health and performance
  • Create data lineage documentation

Collaboration & Process

  • Follow Jira tasks and sprint cycles
  • Participate in code reviews focused on data quality and pipeline reliability
  • Document data schemas, transformations, and pipeline dependencies
  • Work cross-functionally with ML, backend, and product team

Required Qualifications

Technical Skills (Must Have)

  • ✅ Strong Python programming skills (data manipulation with Pandas, NumPy)
  • ✅ SQL proficiency (joins, aggregations, subqueries, window functions)
  • ✅ Understanding of ETL concepts and data pipeline design
  • ✅ Experience working with data (cleaning, transformation, analysis)
  • ✅ Database knowledge (relational or NoSQL - PostgreSQL, MongoDB, etc.)
  • ✅ Git/GitHub proficiency for version control
  • ✅ Basic understanding of data modeling and schema design

Nice to Have

  • Experience with Apache Spark or PySpark
  • Familiarity with Databricks or cloud data platforms
  • Experience with Azure, AWS, or GCP cloud services
  • Knowledge of data warehousing concepts (fact/dimension tables, star schema)
  • Understanding of streaming data and event-driven architectures
  • Experience with data quality frameworks (Great Expectations, dbt)
  • Familiarity with orchestration tools (Airflow, Databricks Workflows)
  • Knowledge of data visualization tools (Tableau, Power BI, Looker)
  • Understanding of distributed computing concepts

Mindset & Approach

  • Detail-oriented: Data quality matters—you care about edge cases and validation
  • Ownership mentality: You build it, you monitor it, you fix it
  • Systems thinker: You understand how data flows through the entire platform
  • Pragmatic problem solver: Balance perfect solutions with shipping quickly
  • Collaborative: Work well with engineers, data scientists, and business stakeholders

What You'll Learn

Data Engineering Skills

  • Lakehouse architecture: Modern data platform combining data lakes and warehouses
  • Databricks ecosystem: Industry-leading platform for data engineering and ML
  • PySpark: Distributed data processing at scale (handle millions of records)
  • Delta Lake: ACID transactions, time travel, schema evolution
  • ETL best practices: Idempotency, error handling, monitoring, testing

Data Quality & Governance

  • Data validation frameworks: Automated quality checks and anomaly detection
  • Schema design: Optimal data models for analytics and ML
  • Data lineage: Track data from source to consumption
  • Access control: Role-based permissions and data security
  • Compliance: GDPR, CCPA requirements for enterprise data

Cloud & Infrastructure

  • Azure services: Blob Storage, SQL Database, Cognitive Search
  • Scalable architecture: Design systems that handle growing data volume
  • Cost optimization: Efficient data storage and compute usage
  • Monitoring: Pipeline health, data quality metrics, performance tracking

Why Join SourceIQ?
✨ Real-world data engineering: Build pipelines processing 150K+ suppliers and millions in procurement spend
🚀 Hands-on mentorship: Work directly with our Head of Technology and senior data engineers
📈 Career acceleration: Exceptional interns convert to full-time data engineering roles
💡 Modern data stack: Databricks, PySpark, Delta Lake—the tools used by leading tech companies
🌍 Mission-driven work: Build data infrastructure that creates economic opportunities for diverse suppliers
🎯 High-impact work: Your pipelines will power AI models and analytics from day one
📊 Rich data environment: Work with real enterprise data from Fortune 500 companies

Our Data Tech Stack

Data Platform (Your Primary Focus):

  • Databricks (unified data + ML platform)
  • Delta Lake (lakehouse storage)
  • PySpark (distributed processing)
  • Databricks Jobs (orchestration)
  • Unity Catalog (governance)
  • Databricks SQL (analytics)
  • Apache Spark (distributed data processing)
  • Python (data engineering scripts, transformation logic)

Data Storage:
MongoDB (NoSQL, application database)
Azure SQL / PostgreSQL (relational data)
Azure Blob Storage (raw files, documents)
Delta Lake (unified data storage with ACID)

Data Sources:
CSV imports (conferences, events, leads)
State/local databases (API and bulk downloads)
Association databases
ERP integrations (Coupa, SAP Ariba, QuickBooks)
Real-time events (user interactions, transactions)

Search & Analytics:
Azure Cognitive Search (full-text and vector search)
Databricks SQL (analytics queries)
Redis (caching layer)

ML Integration:
MLflow (feature engineering, model training)
Databricks Feature Store (reusable features)
TensorFlow / scikit-learn (ML model consumption)

DevOps:
Git/GitHub (version control)
GitHub Actions / Azure DevOps (CI/CD)
Docker (containerization)

Application Requirements

Please submit the following:

  1. Resume/CV highlighting data/analytics coursework, projects, and experience
  2. GitHub profile with code samples (Python, SQL preferred)
  3. Portfolio showcasing data work: Data analysis projects (Jupyter notebooks) ETL scripts or pipeline code SQL queries demonstrating complexity Data visualization dashboards

Brief cover letter (200-300 words) answering:

  • Why are you interested in data engineering at SourceIQ?
  • Describe a data quality issue you've encountered and how you solved it
  • What excites you most about working with large-scale data pipelines?

Note: During the interview process, you may be asked to:
Walk through a data project you've completed
Write SQL queries to solve a business problem
Discuss how you would design an ETL pipeline for a given scenario

Commitment & Expectations

  • Time commitment: 20-40 hours per week (flexible based on your schedule)
  • Duration: Minimum 3 months, with potential to extend to 6 months
  • Work style: Remote-first with regular video standups and async collaboration
  • Academic credit: Available for students whose universities offer it
  • Mentorship: Weekly 1:1s with senior engineers, regular code reviews, pair programming

Ideal Candidate Profile

You might be a great fit if you:
🎓 Are pursuing or recently completed a degree in Computer Science, Data Science, Information Systems, or related field
💻 Have completed coursework in databases, data structures, or data analysis
📊 Have worked with real data (messy, incomplete, from multiple sources)
🔧 Enjoy solving data quality problems and building robust pipelines
🚀 Are excited about building infrastructure that powers AI and analytics
🌱 Want to learn modern data engineering tools and best practices
⚡ Care about performance and scalability of data systems

Ready to Apply?
If you're passionate about data engineering, eager to build production pipelines, and excited about transforming enterprise procurement with clean, reliable data, we want to hear from you!
To apply: Send your resume, GitHub profile, data portfolio, and cover letter to [[email protected]] with the subject line "Data Engineering Intern - [Your Name]"
Questions? Reach out to our team at [[email protected]]

SourceIQ is building the AI-powered infrastructure layer for enterprise procurement. Join us in building the data foundation for a $50B industry transformation.

About the company

SourceIQ company logo

SourceIQ

Actively Hiring
SourceIQ is a cloud-based platform that automates supplier management and sourcing11-50 Employees
Company Size
11-50
Company Type
Technology Provider
Company Type
Artificial Intelligence
Company Type
Enterprise Software Company
Company Type
Marketplace
Company Type
Small And Medium Business
Company Industries
Artificial Intelligence / Machine Learning
  • Top 10% of responders
    SourceIQ is in the top 10% of companies in terms of response time to applications
  • Responds within two weeks
    Based on past data, SourceIQ usually responds to incoming applications within two weeks
  • B2B
  • Early Stage
    Startup in initial stages
Learn more about SourceIQ image

Similar Jobs

Archesys company logo
Archesys
Improving the government services that impact everyday lives
Archesys company logo
Archesys
Improving the government services that impact everyday lives
Voreas Laboratories company logo
Voreas Laboratories
Cyberattack attribution for high-value organizations
Snout company logo
Snout
Pet wellness plans that actually work