
- Top 5% of respondersStackGen is in the top 5% of companies in terms of response time to applications
- Responds within a dayBased on past data, StackGen usually responds to incoming applications within a day
- B2B
- +1
Senior Data Engineer – Data & Context Platform
- $175k – $215k • 0.15% – 0.35%
- |
- |6 years of exp
- |Full Time
In office
Not Available
About the job
About StackGen
StackGen delivers an Agentic Infrastructure Platform powered by Aiden, its AI agent that enables platform engineering, DevOps, and SRE teams to move from manual processes to intent-driven automation. Our platform enables autonomous infrastructure across four pillars: building, governing, healing, while maintaining compliance and security standards across multiple cloud environments. StackGen serves enterprise and fast-growing customers and is based in the San Francisco Bay Area, with a globally distributed team.
The Role
Agents are only as good as the context they can reason over. Today that context is scattered across cloud provider APIs, IaC state, Kubernetes clusters, telemetry systems, CI/CD pipelines, and ticketing tools. Each speaking a different dialect, changing on its own schedule, and describing the same underlying resource in a different way.
We are hiring a Senior Data Engineer to build the layer that solves this problem: the ingestion pipelines, the canonical data model, and the graph and retrieval system that Aiden's agents query. This is a foundational build, not maintenance of something that already works, and the architectural decisions made here will shape the platform for years.
This is a hands-on individual contributor role. You will design the system and write the code that runs it.
What you will do
Data ingestion and pipelines
Build ingestion pipelines across heterogeneous infrastructure sources: cloud provider APIs, Terraform/OpenTofu state, Kubernetes, observability backends, source control, and ticketing systems.
Handle the operational reality of production pipelines: incremental sync, backfill, rate limits, retries, and graceful degradation when an upstream source is unavailable.
Make the batch-versus-streaming call per source rather than applying one pattern everywhere.
Design the canonical schema that unifies how infrastructure, services, ownership, and events are represented across sources. Design for schema evolution so that new sources and entity types can be absorbed without a migration crisis each time.
Context graph and retrieval
Model the relationships between infrastructure, services, ownership, deployments, and incidents so agents can traverse them to answer real operational questions.
Build the retrieval layer on top: graph queries, vector search over unstructured artifacts, and whatever hybrid approach proves out in practice.
Evaluate and select the storage technologies, with the tradeoffs argued explicitly rather than assumed.
Data quality, freshness, and observability
Define and enforce staleness contracts. Agents acting on stale or incorrect context is worse than agents with no context at all.
Instrument the layer so data quality, coverage, and freshness are measurable rather than anecdotal, and build the validation and alerting that catches drift early.
Multi-tenancy, scale, and security
Build tenant isolation into the storage and query paths from the start, not as a later retrofit.
Apply access control and data-handling practices appropriate to customer infrastructure metadata, working with security to keep controls enforceable.
Cross-functional collaboration
Partner with the agent, platform, and product teams to understand what context agents actually need, and make the layer usable enough that other engineers build on it without needing you in the loop.
Write design docs and tradeoff memos that let the team engage with the architecture, and raise the bar through code review and shared practice.
What we're looking for
- 6+ years building production data systems, with substantial experience owning data pipelines end to end: ingestion, transformation, orchestration, and the 3am operational reality.
- Hands-on experience with data normalization, entity resolution, or master data management in messy multi-source environments where the same entity has different names and different levels of completeness in each system.
- Strong data modeling judgment. You can explain why you would choose a graph, relational, document, or hybrid store for a given problem, and you have been wrong about one before and learned from it.
- Proficiency in Python (Go is a plus) and SQL well beyond the basics.
- Experience designing systems from a blank page, including making decisions with incomplete information and revisiting them when reality disagrees.
- Working familiarity with cloud and DevOps ecosystems — AWS/GCP/Azure resource models, Terraform, Kubernetes APIs. You should find the infrastructure domain genuinely interesting, not just the pipelines.
- Clear written communication. Design documents, tradeoff memos, and code that explains why rather than what.
- Comfort in a startup environment with ambiguity and high ownership.
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
Nice to have
- Graph databases in production (Neo4j, Neptune, Dgraph, Memgraph, or similar), including a clear view of where they stop being the right answer.
- Vector databases and retrieval systems, particularly RAG patterns serving agents rather than chat interfaces.
- Streaming or CDC experience (Kafka, Debezium, Flink) and the judgment to know when a batch is the better call.
- Multi-tenant SaaS data architecture.
- Data quality, lineage, or catalog tooling (Great Expectations, OpenLineage, DataHub, or similar).
- Exposure to agentic AI frameworks and an understanding of how context quality affects agent behavior.
- Experience with infrastructure or CMDB-style data models (CAI, OCSF, or internal equivalents).
Why StackGen
- High ownership. You will set the standard for how data and context work at StackGen, as the engineer who owns this layer.
- Build the foundation of an AI-driven infrastructure platform from the ground up, on a greenfield problem with real enterprise customers already depending on the outcome.
- Work directly with engineering leadership and influence architecture across the platform.
- A collaborative engineering culture that values continuous learning, deep technical work, and short paths from problem to fix.
About the company

StackGen
- Top 5% of respondersStackGen is in the top 5% of companies in terms of response time to applications
- Responds within a dayBased on past data, StackGen usually responds to incoming applications within a day
- B2B
- Early StageStartup in initial stages
Similar Jobs








