
- Top 10% of respondersSourceIQ is in the top 10% of companies in terms of response time to applications
- Responds within two weeksBased on past data, SourceIQ usually responds to incoming applications within two weeks
- B2B
- +1
Data Engineering Intern (Python + Databricks + ETL Pipelines)
- No equity
- |Remote (Austin •+18)
- |4 years of exp
- |Internship
Remote only
Not Available
About the job
Company: SourceIQ
Location: Remote
Duration: 3-6 months (with potential for conversion to full-time)
Compensation: Unpaid (Academic credit available)
Start Date: Flexible (Rolling applications) Winter
About SourceIQ
SourceIQ is an AI-powered supplier management and sourcing platform transforming how enterprises discover, evaluate, onboard, and collaborate with suppliers. We automate compliance, analytics, and sourcing workflows from end to end, helping companies streamline procurement operations.
We're building the infrastructure layer for enterprise procurement—a platform designed to consolidate a fragmented $50B market by becoming the centralized intelligence and data layer that integrates with all procurement tools.
The Role
As a Data Engineering Intern, you'll build and maintain the data infrastructure that powers SourceIQ's AI models and analytics. You'll work with Databricks, Apache Spark, Python, and Azure to create ETL pipelines, process large datasets, and ensure data quality across our platform.
This is hands-on data engineering where you'll contribute directly to production data pipelines used by enterprise customers including JP Morgan and Humana. You'll work closely with our Head of Technology, ML engineers, and backend team to build the data foundation for our AI-powered platform.
You'll learn modern data engineering practices including lakehouse architecture, Delta Lake, PySpark, and data quality frameworks while working on real business problems.
What You'll Do
- ETL Pipeline Development (Primary Focus)
- Build and maintain ETL pipelines using Python and PySpark on Databricks
- Process supplier data from multiple sources: CSV imports from conferences and events State/local government databases (150K+ suppliers) Association databases ERP systems (Coupa, SAP Ariba, QuickBooks) API integrations (real-time data ingestion)
- Implement schema mapping to standardize non-standard data formats
- Create automated data refresh pipelines (e.g., marketplace data updated every 30 days)
Data Quality & Validation
- Design and implement data quality checks and validation rules
- Build data profiling and completeness scoring systems
- Create deduplication logic to identify duplicate suppliers across sources
- Implement data cleansing (address normalization, phone validation, etc.)
- Monitor data quality metrics and set up alerting for anomalies
Lakehouse Architecture (Databricks)
- Work with Delta Lake (Bronze → Silver → Gold layers) Bronze: Raw data ingestion from all sources Silver: Cleaned, deduplicated, standardized data Gold: Aggregated features for ML models and analytics
- Design and optimize data models for analytics and ML workloads
- Implement data versioning and time travel for auditability
- Work with Unity Catalog for data governance and access control
Analytics & Reporting Support
- Create aggregated data tables for dashboards and BI tools
- Build data marts for spend analytics, supplier performance, and engagement metrics
- Optimize query performance for reporting workloads
- Support ad-hoc data analysis requests from product and business teams
Data Pipeline Orchestration
- Implement Databricks Jobs for scheduled data processing
- Build idempotent, fault-tolerant pipelines with retry logic
- Monitor pipeline health and performance
- Create data lineage documentation
Collaboration & Process
- Follow Jira tasks and sprint cycles
- Participate in code reviews focused on data quality and pipeline reliability
- Document data schemas, transformations, and pipeline dependencies
- Work cross-functionally with ML, backend, and product team
Required Qualifications
Technical Skills (Must Have)
- ✅ Strong Python programming skills (data manipulation with Pandas, NumPy)
- ✅ SQL proficiency (joins, aggregations, subqueries, window functions)
- ✅ Understanding of ETL concepts and data pipeline design
- ✅ Experience working with data (cleaning, transformation, analysis)
- ✅ Database knowledge (relational or NoSQL - PostgreSQL, MongoDB, etc.)
- ✅ Git/GitHub proficiency for version control
- ✅ Basic understanding of data modeling and schema design
Nice to Have
- Experience with Apache Spark or PySpark
- Familiarity with Databricks or cloud data platforms
- Experience with Azure, AWS, or GCP cloud services
- Knowledge of data warehousing concepts (fact/dimension tables, star schema)
- Understanding of streaming data and event-driven architectures
- Experience with data quality frameworks (Great Expectations, dbt)
- Familiarity with orchestration tools (Airflow, Databricks Workflows)
- Knowledge of data visualization tools (Tableau, Power BI, Looker)
- Understanding of distributed computing concepts
Mindset & Approach
- Detail-oriented: Data quality matters—you care about edge cases and validation
- Ownership mentality: You build it, you monitor it, you fix it
- Systems thinker: You understand how data flows through the entire platform
- Pragmatic problem solver: Balance perfect solutions with shipping quickly
- Collaborative: Work well with engineers, data scientists, and business stakeholders
What You'll Learn
Data Engineering Skills
- Lakehouse architecture: Modern data platform combining data lakes and warehouses
- Databricks ecosystem: Industry-leading platform for data engineering and ML
- PySpark: Distributed data processing at scale (handle millions of records)
- Delta Lake: ACID transactions, time travel, schema evolution
- ETL best practices: Idempotency, error handling, monitoring, testing
Data Quality & Governance
- Data validation frameworks: Automated quality checks and anomaly detection
- Schema design: Optimal data models for analytics and ML
- Data lineage: Track data from source to consumption
- Access control: Role-based permissions and data security
- Compliance: GDPR, CCPA requirements for enterprise data
Cloud & Infrastructure
- Azure services: Blob Storage, SQL Database, Cognitive Search
- Scalable architecture: Design systems that handle growing data volume
- Cost optimization: Efficient data storage and compute usage
- Monitoring: Pipeline health, data quality metrics, performance tracking
Why Join SourceIQ?
✨ Real-world data engineering: Build pipelines processing 150K+ suppliers and millions in procurement spend
🚀 Hands-on mentorship: Work directly with our Head of Technology and senior data engineers
📈 Career acceleration: Exceptional interns convert to full-time data engineering roles
💡 Modern data stack: Databricks, PySpark, Delta Lake—the tools used by leading tech companies
🌍 Mission-driven work: Build data infrastructure that creates economic opportunities for diverse suppliers
🎯 High-impact work: Your pipelines will power AI models and analytics from day one
📊 Rich data environment: Work with real enterprise data from Fortune 500 companies
Our Data Tech Stack
Data Platform (Your Primary Focus):
- Databricks (unified data + ML platform)
- Delta Lake (lakehouse storage)
- PySpark (distributed processing)
- Databricks Jobs (orchestration)
- Unity Catalog (governance)
- Databricks SQL (analytics)
- Apache Spark (distributed data processing)
- Python (data engineering scripts, transformation logic)
Data Storage:
MongoDB (NoSQL, application database)
Azure SQL / PostgreSQL (relational data)
Azure Blob Storage (raw files, documents)
Delta Lake (unified data storage with ACID)
Data Sources:
CSV imports (conferences, events, leads)
State/local databases (API and bulk downloads)
Association databases
ERP integrations (Coupa, SAP Ariba, QuickBooks)
Real-time events (user interactions, transactions)
Search & Analytics:
Azure Cognitive Search (full-text and vector search)
Databricks SQL (analytics queries)
Redis (caching layer)
ML Integration:
MLflow (feature engineering, model training)
Databricks Feature Store (reusable features)
TensorFlow / scikit-learn (ML model consumption)
DevOps:
Git/GitHub (version control)
GitHub Actions / Azure DevOps (CI/CD)
Docker (containerization)
Application Requirements
Please submit the following:
- Resume/CV highlighting data/analytics coursework, projects, and experience
- GitHub profile with code samples (Python, SQL preferred)
- Portfolio showcasing data work: Data analysis projects (Jupyter notebooks) ETL scripts or pipeline code SQL queries demonstrating complexity Data visualization dashboards
Brief cover letter (200-300 words) answering:
- Why are you interested in data engineering at SourceIQ?
- Describe a data quality issue you've encountered and how you solved it
- What excites you most about working with large-scale data pipelines?
Note: During the interview process, you may be asked to:
Walk through a data project you've completed
Write SQL queries to solve a business problem
Discuss how you would design an ETL pipeline for a given scenario
Commitment & Expectations
- Time commitment: 20-40 hours per week (flexible based on your schedule)
- Duration: Minimum 3 months, with potential to extend to 6 months
- Work style: Remote-first with regular video standups and async collaboration
- Academic credit: Available for students whose universities offer it
- Mentorship: Weekly 1:1s with senior engineers, regular code reviews, pair programming
Ideal Candidate Profile
You might be a great fit if you:
🎓 Are pursuing or recently completed a degree in Computer Science, Data Science, Information Systems, or related field
💻 Have completed coursework in databases, data structures, or data analysis
📊 Have worked with real data (messy, incomplete, from multiple sources)
🔧 Enjoy solving data quality problems and building robust pipelines
🚀 Are excited about building infrastructure that powers AI and analytics
🌱 Want to learn modern data engineering tools and best practices
⚡ Care about performance and scalability of data systems
Ready to Apply?
If you're passionate about data engineering, eager to build production pipelines, and excited about transforming enterprise procurement with clean, reliable data, we want to hear from you!
To apply: Send your resume, GitHub profile, data portfolio, and cover letter to [[email protected]] with the subject line "Data Engineering Intern - [Your Name]"
Questions? Reach out to our team at [[email protected]]
SourceIQ is building the AI-powered infrastructure layer for enterprise procurement. Join us in building the data foundation for a $50B industry transformation.
About the company

SourceIQ
- Top 10% of respondersSourceIQ is in the top 10% of companies in terms of response time to applications
- Responds within two weeksBased on past data, SourceIQ usually responds to incoming applications within two weeks
- B2B
- Early StageStartup in initial stages
Employees joined from
Similar Jobs





