Avatar for 3GIMBALS
3GIMBALS
Actively Hiring
Artificial Intelligence for Geospatial Intelligence
  • Early Stage
    Startup in initial stages

Data Engineer – Analytic Platform & Data Pipelines

Posted: 7 days ago• Recruiter recently active
Hires remotely in
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
SQL
MongoDB
PostgreSQL
JSON
Azure
Spark
Docker
GeoJson
ElasticSearch
AWS
Parquet
Kubernetes
Avro
Airflow
Dask
GCP
Prefect
Dagster
AWS GovCloud
Azure Government

About the job

Role Overview

3GIMBALS is seeking a Data Engineer to design, build, and maintain the data pipelines and infrastructure that power our unclassified PAI/CAI-based analytic platform. This role is responsible for ingesting, transforming, and curating large volumes of structured and unstructured data from diverse open and commercial sources; building resilient, automated ETL/ELT workflows; and ensuring data is high-quality, well-governed, and analysis-ready for the downstream analytics, knowledge graph, and modeling teams. The ideal candidate is comfortable working with messy, multi-source data at scale within secure development environments.

Key Responsibilities

Data Pipeline Development & Ingestion

  • Design and build scalable batch and streaming pipelines to ingest structured and unstructured data from PAI/CAI sources, APIs, and third-party feeds
  • Develop ETL/ELT workflows to normalize, enrich, and transform heterogeneous data into standardized schemas
  • Build and maintain automated ingestion connectors for web, document, geospatial, and tabular data sources
  • Manage data orchestration and scheduling using tools such as Airflow, Dagster, or Prefect

Data Modeling & Storage

  • Design and maintain data models, schemas, and storage layers across relational, NoSQL, and object stores
  • Build and maintain data lakes/lakehouses and curated, analysis-ready data marts
  • Optimize partitioning, indexing, and query performance for large datasets
  • Support entity resolution and data linking in coordination with the knowledge graph and modeling teams

Data Quality, Governance & Lineage

  • Implement data validation, quality checks, and monitoring across pipelines
  • Establish data lineage, cataloging, and metadata management
  • Enforce data governance, provenance tracking, and source attribution appropriate for PAI/CAI data
  • Document datasets, schemas, and pipeline logic for downstream consumers

Security & Compliance

  • Ensure pipelines and data stores meet security requirements for operation in sensitive environments
  • Implement encryption, access control, and secure data-handling practices
  • Support Authority to Operate (ATO) processes and compliance frameworks

Required Qualifications

Technical Expertise

  • 4+ years of data engineering experience building and operating production data pipelines
  • Strong programming skills in Python and SQL (Scala or Java a plus)
  • Experience with distributed data processing frameworks (Spark, Dask, or similar)
  • Hands-on experience with workflow orchestration tools (Airflow, Dagster, Prefect)
  • Proficiency with relational and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch, etc.)
  • Experience with cloud data platforms and services (AWS, Azure, or GCP)

Data & Infrastructure

  • Experience designing data models, warehouses, and lakehouse architectures
  • Familiarity with data formats and serialization (Parquet, Avro, JSON, GeoJSON)
  • Understanding of data quality, lineage, and governance practices
  • Experience with containerization (Docker) and CI/CD for data workflows

Domain Knowledge

  • Experience working with large-scale, heterogeneous, or open-source datasets
  • Understanding of data provenance and source-attribution requirements

Preferred Qualifications

  • Active security clearance or ability to obtain one
  • Experience in government, defense, or intelligence contracting environments
  • Familiarity with PAI/CAI (publicly and commercially available information) data sources
  • Experience with geospatial data processing (PostGIS, GDAL, or similar)
  • Knowledge of graph data structures and preparing data for knowledge graphs
  • Experience with streaming platforms (Kafka, Kinesis)
  • Familiarity with federal compliance frameworks (FedRAMP, FISMA, NIST 800-53)

Technical Environment

  • Languages: Python, SQL (Scala/Java a plus)
  • Processing: Spark, Airflow/Dagster/Prefect, streaming frameworks
  • Storage: PostgreSQL, Elasticsearch, object storage / data lake, Parquet
  • Infrastructure: Docker, Kubernetes, cloud platforms (AWS GovCloud, Azure Government)
  • Security: Encryption at rest and in transit, RBAC, secure data handling

This role is central to the platform: the data engineering team delivers the clean, trustworthy, well-documented data that every analytic, knowledge graph, and risk-modeling capability depends on.

Salary: $135000 - $185000 per year

About the company

3GIMBALS company logo

3GIMBALS

Actively Hiring
Artificial Intelligence for Geospatial Intelligence11-50 Employees
  • Early Stage
    Startup in initial stages
Learn more about 3GIMBALS image

Founders

Terry Dyess
Founder
Miami
image
View the team image

Similar Jobs