
Collinear.ai
Actively Hiring
Collinear solves the fundamental problem of LLM customization; Elevate your LLM game today
- B2B
- Early StageStartup in initial stages
- Growing fastShowed strong hiring growth in the past month
MTS - Engineering (Data Infrastructure)
- |Full Time
Posted: yesterday• Recruiter recently active
Job Location
Visa Sponsorship
Not Available
RelocationNot Allowed
Hiring contact
Nick Shelton
Employee

About the job
About the role
As a Member of Technical Staff - Engineering (Data Infrastructure), you will own the systems that turn large, real-world datasets into data Collinear can build on. Our environments are grounded in terabytes of data, spread across archives, spreadsheets, email, PDFs, and scanned documents. The pace at which we process these determines how fast we deliver to frontier labs.
This is a hands-on role at the intersection of algorithms, distributed systems, and data quality. Many of our hardest problems, such as linking related records across millions of files, don't split up neatly, and you will define how we solve them at scale.
What you'll do
- Build pipelines that process multi-terabyte datasets in parallel across archives, spreadsheets, email, PDFs, and scanned documents
- Design graph-based systems that link related records, such as the same person or company appearing across millions of files
- Build fast string and pattern search over large, heterogeneous datasets
- Profile and remove bottlenecks, and decide how to split work that doesn't parallelize neatly
- Transform data for downstream use, including consistently replacing sensitive fields across files
- Define how we measure data quality, and build review tools so the team can catch and fix errors without reprocessing everything
- Assess new datasets and filter out low-quality data before it reaches our environments
About you
- You have 5+ years of experience building data-intensive systems in production
- You have processed large datasets in parallel with frameworks such as Apache Spark, Ray, or Dask, and know when to design your own
- You have a strong command of graph algorithms, and experience using them to transform large amounts of data
- You have built efficient string matching and pattern search at scale, such as fuzzy matching or indexing
- You make sound tradeoffs between accuracy, speed, and cost, and can explain them clearly
Nice to have
- Experience with entity resolution or record linkage
- Experience with NLP or LLM-based information extraction
- OCR or document processing experience, including poor scans and handwriting
- Experience in a systems language such as Rust, C++, or Go
- Experience with regulated or sensitive data, such as financial or healthcare records
About the company
11-50
Artificial Intelligence / Machine Learning
- B2B
- Early StageStartup in initial stages
- Growing fastShowed strong hiring growth in the past month
Similar Jobs

Oklo
Making reactors people want

LiveReach AI
Easy to Use Cloud Camera Systems

Astranis
Building next-generation internet satellites to get the world online

Neuralink
Ultra-high bandwidth brain-machine interfaces to connect humans and computers

Ursa Major
Powering the future of defense and aerospace. Fly More. Fly Faster

IonQ
Building the world’s best quantum computers to solve the world’s most ecomplex problems

Pyka
Autonomous Electric Airplanes

Ayar Labs
It's time for optical I/O

Path Robotics, Inc (We're Hiring!)
We develop and advance state-of-the-art methods to solve general problems