Avatar for Thirdbase
Thirdbase
Actively Hiring
Cindr is the open-core, developer first security engine redefining email threat detection
  • B2B
  • Early Stage
    Startup in initial stages

CTO / Lead Engineer, Thirdbase Capital

  • Remote (
    Everywhere
    ) • +1
  • |5 years of exp
  • |Full Time
Posted: 1 month ago• Recruiter recently active
Job Location
Remote Work Policy

Onsite or remote

Hires remotely in
Everywhere
Visa Sponsorship

Available

RelocationAllowed
Skills
Python
Entity Framework
Machine Learning Data Science Python
Large Language Models (LLMs)
Hiring contact
Nitya Singh
Employee
image

About the job

About Thirdbase Capital

Thirdbase Capital is a growth-stage venture platform investing in AI, defense, frontier tech, robotics, energy, and fintech. Our portfolio includes some of the most sought-after private companies in AI and deep tech. We are now building the internal technology that will define how we source and underwrite the next decade of investments.

The Role

  • We are hiring a founding engineer to build Thirdbase Intelligence, a proprietary system that helps us identify exceptional growth companies earlier than conventional investors.
  • The system will measure where top talent is moving, how product and AI usage is accelerating, and whether that activity is converting into durable revenue. The north-star question it must answer: "What private company is becoming dramatically better, faster than the market realizes, and what objective evidence proves it?"
  • This is a data-engineering-first mandate. You will begin with the longitudinal data model, entity resolution, ingestion pipelines, and scoring methodology, and add a natural-language interface once the data foundation is solid. The full technical scope will be shared with shortlisted candidates.

What You'll Work On

  • Designing a longitudinal data model that tracks companies, people, roles, and career transitions over time
  • Building reliable ingestion pipelines for licensed workforce, hiring, web, and developer-activity datasets
  • Entity resolution across messy real-world data at the scale of millions of companies and profiles
  • Quantitative scoring and ranking methodologies, with backtesting against historical investment outcomes
  • A secure, read-only connector framework for structured diligence on companies raising capital
  • An evidence-grounded query and alerting layer on top of the data platform, built with modern LLM tooling

What We're Looking For

  • 5+ years building data-intensive systems: pipelines, entity resolution, time-series infrastructure, and graph data models at scale
  • Strong backend engineering (Python and/or a systems language; SQL; modern orchestration and warehouse tooling)
  • Experience integrating third-party data APIs and building monitored, production-grade ingestion for imperfect datasets
  • Practical experience with LLM application development, including retrieval, structured outputs, and agentic workflows
  • Ability to design scoring and ranking methodologies and validate them empirically
  • A security-conscious mindset for handling sensitive third-party data under strict access controls and licensing terms
  • Self-direction, with the ability to own architecture end-to-end and work directly with the investment team

Nice to Have

  • Prior work on people or workforce data, alternative data for investing, or data platforms at growth or hedge funds
  • Experience building secure customer-facing data connectors (OAuth, read-only scopes, audit logging)
  • Background in quantitative research, recommender systems, or anomaly detection
  • Why This Role
  • You will own a greenfield build with direct impact on real investment decisions, work closely with the fund's partners, and create a system that compounds in value with every week of data collected. This is a chance to build the kind of proprietary infrastructure most funds only talk about.

How to Apply

Send your CV or portfolio along with brief answers to the following to [email protected]:

  • Describe the most complex data system you have built: its architecture, scale, and the hardest problem you solved.
  • If you had to merge person-level records from three overlapping datasets with inconsistent identifiers, how would you approach entity resolution and measure its accuracy?

Similar Jobs

GVOS  company logo
GVOS
An Edge Cloud for Autonomous Driving
Enigma Technologies company logo
Enigma Technologies
Enigma provides trusted business data built on unparalleled entity resolution
Coast company logo
Coast
Coast is re-imagining B2B payments, beginning with fleet and fuel
Vise company logo
Vise
Technology-Powered Asset Manager for customized, intelligent investing
Deepgram company logo
Deepgram
AI speech API for transcription with human-level understanding
Solace company logo
Solace
Solace is a healthcare advocacy marketplace
Arusto company logo
Arusto
Arusto automates the creation of adult learning assets in minutes & at min cost
TrialSpark company logo
TrialSpark
Our mission is to bring new treatments to patients faster and more efficiently
Pulley company logo
Pulley
Helping project teams break ground faster