Avatar for Wynd Labs
Wynd Labs
Actively Hiring
Making AI Data Accessible. Building a suite of products powered by Grass
  • Early Stage
    Startup in initial stages
  • Growing fast
    Showed strong hiring growth in the past month

Machine Learning Engineer

Reposted: 7 days ago• Recruiter recently active
Job Location
Visa Sponsorship

Not Available

RelocationAllowed
Hiring contact
Christopher Nguyen
Founder
New York City
image

About the job

Who We Are:

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.

We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

The Role:

We are looking for a Machine Learning Engineer with strong skills and significant experience developing machine learning models. You will join a small, innovative team and lead efforts to advance our capabilities, drive model development, and support our vision for a future where Grass is transformative in the internet's evolution.

Please note: This role requires a work schedule that sufficiently overlaps with EST business hours to collaborate effectively with the team.

Who You Are:

  • Bachelor’s, Master’s, or Doctoral degree in Data Science, Computer Science, Statistics, or a related field.
  • A minimum of 3 years of work or research experience dealing with large datasets.
  • Experience working with large-scale text datasets, NLP pipelines, or data preparation for LLM training is highly preferred.
  • Strong coding skills in Python or other object-oriented programming languages.
  • Experience with text deduplication, dataset filtering, corpus curation, or data distillation is a strong plus.
  • Graduate-level knowledge of statistics, including but not limited to hypothesis testing, regression analysis, and probability.
  • Excellent work ethic and the ability to thrive in a fast-paced startup environment.
  • Strong problem-solving skills and attention to detail.
  • Good communication skills, with the ability to articulate complex data concepts to non-technical stakeholders.
  • Experience working in a high-output team.

What You'll Be Doing:

  • Developing data processing pipelines and machine learning solutions for large-scale NLP and LLM applications, including improving the quality, filtering, and preparation of training datasets.
  • Designing and implementing pipelines for processing and analyzing large datasets.
  • Analyzing and interpreting complex time series data to provide actionable insights and solutions.
  • Designing, implementing, and maintaining data-driven models and algorithms.
  • Developing techniques for dataset curation to improve the quality and efficiency of AI training data.
  • Building scalable pipelines for filtering, deduplicating, and improving large-scale text datasets used for LLM training.
  • Collaborating with cross-functional teams to understand data needs and deliver timely solutions.
  • Ensuring data quality and integrity throughout all processes.
  • Utilizing Optical Character Recognition (OCR) technology to convert different types of documents into editable and searchable data.
  • Continuously researching and implementing best practices in data science and machine learning.
  • Contributing to the development and improvement of internal data processing tools and infrastructure.

Why Work With Us:

  • Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.
  • Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better.
    We prioritize low ego and high output. This is a fully remote team.
  • Compensation. You’ll receive a competitive salary, benefits and equity package.

About the company

Wynd Labs company logo

Wynd Labs

Actively Hiring
Making AI Data Accessible. Building a suite of products powered by Grass11-50 Employees
  • Early Stage
    Startup in initial stages
  • Growing fast
    Showed strong hiring growth in the past month
Learn more about Wynd Labs image

Funding

AMOUNT RAISED
$4.5M
FUNDED OVER
2 rounds
Rounds
S
$3500000
Seed - Jan 2024+1

Founders

Christopher Nguyen
Founder
New York City
image
View the team image

Similar Jobs

Albeado company logo
Albeado
Breakthrough causal AI predictions, optimizations and interventions - in real time
Voicera.io company logo
Voicera.io
Believe what you hear. Trust what you see
Wynd Labs company logo
Wynd Labs
Making AI Data Accessible. Building a suite of products powered by Grass
Deepgram company logo
Deepgram
AI speech API for transcription with human-level understanding
Scale AI company logo
Scale AI
Accelerate the development of AI applications
TrialSpark company logo
TrialSpark
Our mission is to bring new treatments to patients faster and more efficiently
MOLOCO company logo
MOLOCO
Machine learning-powered performance advertising solutions
Rocket Money company logo
Rocket Money
The Money App That Works for You