Avatar for akkodis
akkodis
Actively Hiring
Digital engineering solutions for industry innovation

Speech Data Engineer

Posted: 1 week ago• Recruiter recently active
Hires remotely in
Remote Work Policy

Remote only

Company Location
Zug
Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
Machine Learning
Git
Gitlab
Deep Learning
Gerrit
PyTorch
CNNs
RNNs
FFT
MFCC
Transformers
LSTMs
Noise Filtering
Mel Spectrograms
Audio Segmentation
Audio Alignment

About the job

AI/ML Engineer, Speech Data Scientist

Position Overview

We are seeking a hands-on AI/ML Engineer specializing in speech technologies to support the development and improvement of production-scale speech synthesis solutions. This individual will work across the full speech machine learning lifecycle, with a primary focus on speech data sourcing, dataset cleanup, audio processing, training-set preparation, and model evaluation.

The ideal candidate has experience working directly with Text-to-Speech (TTS) or speech synthesis systems and understands how the quality, structure, and diversity of speech data impact model performance. This position will collaborate with data collection, engineering, and machine learning teams to improve speech datasets, support model training, evaluate model quality, and troubleshoot performance issues.

This is a highly collaborative and hands-on role for someone who enjoys solving complex speech and audio data problems in a fast-paced product development environment.

Rate Range: $60/hour to $64.98/hour; The rate may be negotiable based on experience, education, geographic location, and other factors.

Employment Details

  • Employment Type: Contract
  • Location: Remote
  • Schedule: Full-time
  • Experience Level: Mid-to-senior level

Key Objectives

  • Improve the quality and effectiveness of datasets used to train and evaluate speech synthesis models.
  • Build and maintain scalable processes for speech data sourcing, cleanup, filtering, augmentation, and preparation.
  • Support the training and evaluation of TTS models, mel-spectrogram models, and vocoders.
  • Identify data or model-related issues that affect speech quality, accuracy, bias, scalability, or performance.
  • Develop reliable evaluation methodologies for production speech AI systems.

Responsibilities

Speech Data Sourcing and Preparation

  • Source, collect, organize, and prepare speech datasets for model training and evaluation.
  • Work with data collection and data harvesting teams to improve the quality, diversity, and usability of speech data.
  • Develop and improve processes for audio cleanup, segmentation, filtering, augmentation, normalization, and quality control.
  • Prepare TTS training datasets, including transcript validation, speaker balancing, alignment, and phoneme-related processing.
  • Evaluate available speech datasets and recommend appropriate datasets for specific training and testing objectives.
  • Identify issues involving noise, poor recording quality, transcription errors, speaker imbalance, pronunciation, or insufficient language coverage.

Model Training and Optimization

  • Train, fine-tune, and optimize speech synthesis models using PyTorch.
  • Support the training of mel-spectrogram models, acoustic models, and neural vocoders.
  • Collaborate with engineering teams to improve model performance across different hardware platforms and deployment environments.
  • Analyze unsuccessful training outcomes and determine whether issues originate from the model, training configuration, or underlying data.
  • Recommend improvements to datasets, preprocessing workflows, training methods, and evaluation strategies.

Model Evaluation

  • Develop, maintain, and improve TTS model evaluation systems.
  • Measure and benchmark speech model performance using objective and subjective evaluation methods.
  • Evaluate naturalness, intelligibility, speaker similarity, pronunciation, latency, throughput, and overall audio quality.
  • Perform detailed error analysis and identify patterns affecting model performance.
  • Analyze model quality and potential bias across speakers, languages, accents, and use cases.
  • Translate evaluation findings into actionable recommendations for data and model improvements.
  • Characterize performance and quality metrics across platforms for speech AI components.

Cross-Functional Collaboration

  • Collaborate with machine learning engineers, software engineers, researchers, linguists, and data collection teams.
  • Participate in code reviews, design reviews, use-case reviews, and test-plan reviews.
  • Document dataset decisions, preprocessing workflows, evaluation results, and recommended corrective actions.
  • Troubleshoot technical issues in a collaborative environment and help determine root causes.
  • Contribute to new product capabilities and improvements to existing speech AI solutions.

Required Qualifications

  • Master’s degree, PhD, or equivalent professional experience in Computer Science, Electrical Engineering, Artificial Intelligence, Applied Mathematics, Linguistics, Computational Linguistics, or a related discipline.
  • Approximately 3 to 5 or more years of relevant industry experience in machine learning, speech technologies, speech data science, or a related field.
  • Hands-on experience with Text-to-Speech, speech synthesis, speech recognition, voice technologies, or speech-to-speech systems.
  • Experience sourcing, cleaning, processing, filtering, augmenting, or evaluating speech and audio datasets.
  • Practical experience preparing datasets for speech model training and evaluation.
  • Experience training or fine-tuning speech models rather than only integrating third-party speech APIs.
  • Advanced Python programming skills.
  • Hands-on experience using PyTorch.
  • Strong knowledge of machine learning and deep learning concepts, tools, and techniques.
  • Understanding of neural network architectures such as CNNs, RNNs, LSTMs, and Transformers.
  • Familiarity with speech digital signal processing and feature extraction techniques, including:
  • FFT
  • MFCC
  • Mel spectrograms
  • Audio segmentation
  • Audio alignment
  • Noise filtering
  • Experience measuring, benchmarking, or evaluating speech model performance.
  • Experience using version control and code review tools such as Git, GitLab, or Gerrit.
  • Strong analytical and problem-solving skills, including the ability to diagnose technical issues independently.
  • Ability to work effectively in an agile environment with evolving priorities.
  • Strong collaboration, communication, and interpersonal skills.

Preferred Qualifications

  • Experience with modern TTS architectures or neural vocoders, such as:
  • FastSpeech or FastSpeech 2
  • Tacotron or Tacotron 2
  • VITS
  • XTTS
  • HiFi-GAN
  • WaveGlow
  • Experience working with publicly available or proprietary speech datasets, such as:
  • LibriTTS
  • LibriSpeech
  • LJSpeech
  • Mozilla Common Voice
  • VCTK
  • VoxCeleb
  • Experience with voice cloning, personalized speech synthesis, multi-speaker TTS, or speaker embeddings.
  • Experience developing multilingual or code-switched TTS systems.
  • Experience with cross-lingual voice cloning.
  • Familiarity with Grapheme-to-Phoneme systems, phonemization, pronunciation modeling, or forced alignment.
  • Experience with Weighted Finite-State Transducers, text normalization, or inverse text normalization.
  • Knowledge of linguistics, phonetics, phonology, accents, dialects, or language technologies.
  • Experience deploying machine learning models in production, cloud, data center, or embedded environments.
  • Familiarity with GPU technologies such as CUDA, cuDNN, or TensorRT.
  • Experience improving model quality through data cleanup, speaker balancing, noise filtering, duration alignment, phonemization, or dataset augmentation.
  • Strong C++ programming skills.
  • Native or near-native fluency in an additional language, including Spanish, Mandarin, German, Japanese, Russian, French, Arabic, Hindi, Korean, Italian, Portuguese, or UK English.

What Success Looks Like

A successful candidate will be able to:

  • Explain how speech datasets are sourced, cleaned, validated, and prepared for TTS training.
  • Identify data quality issues that negatively affect speech model performance.
  • Build repeatable speech data processing and evaluation workflows.
  • Work directly with TTS training pipelines and understand how data affects output quality.
  • Evaluate synthesized speech using appropriate quality, similarity, intelligibility, and performance metrics.
  • Diagnose whether a model issue is caused by training data, preprocessing, architecture, or configuration.
  • Recommend specific, technically sound improvements based on evaluation findings.
  • Collaborate effectively across data, engineering, research, and product teams.

Equal Opportunity Employer/Veterans/Disabled

Benefit offerings available for our associates include medical, dental, vision, life insurance, short-term disability, additional voluntary benefits, an EAP program, commuter benefits, and a 401K plan. Our benefit offerings provide employees the flexibility to choose the type of coverage that meets their individual needs. In addition, our associates may be eligible for paid leave including Paid Sick Leave or any other paid leave required by Federal, State, or local law, as well as Holiday pay where applicable. Disclaimer: These benefit offerings do not apply to client-recruited jobs and jobs that are direct hires to a client.

To read our Candidate Privacy Information Statement, which explains how we will use your information, please visit https://www.akkodis.com/en/privacy-policy.

The Company will consider qualified applicants with arrest and conviction records in accordance with federal, state, and local laws and/or security clearance requirements, including, as applicable:

· The California Fair Chance Act

· Los Angeles City Fair Chance Ordinance

· Los Angeles County Fair Chance Ordinance for Employers

· San Francisco Fair Chance Ordinance

About the company

akkodis company logo

akkodis

Actively Hiring
Digital engineering solutions for industry innovation5000+ Employees
Company Location
Zug
Company Size
5000+
Learn more about akkodis image