
Senior AI Engineer
- ₹60L – ₹80L • 0.0% – 1.0%
- |
- |5 years of exp
- |Full Time
In office
Not Available
About the job
Hiring: Senior Audio Research Engineer
We're building India's most powerful custom SLMs and real-time Speech-to-Speech voice agents that already handle millions of calls across banking, healthcare, and new economy brands.
We want to take our platform to even crazier scale. We're building things that real people use. Come train and improve models you're proud of.
If you're a Data Scientist / ML Engineer who thrives on pushing the limits of what's possible and wants to own the model training and fine-tuning for an early-stage AI startup, this is your seat at the table.
The Mission
As our Senior Audio Research Engineer, you will own the training and fine-tuning of the models behind our voice AI platform. You'll be responsible for training and improving custom audio and speech models, building high-quality training pipelines, running experiments, and turning research ideas into models that perform reliably in the real world. Your goal is to continuously improve model quality, performance, and the overall product experience.
What You'll Do
- Train and fine-tune models for custom SLMs and real-time Speech-to-Speech voice agents
- Build and improve training pipelines for large-scale speech and audio datasets
- Experiment with model architectures, training objectives, and fine-tuning techniques to continuously improve model quality
- Implement ideas from research papers and evaluate whether they work on our real-world problems
- Analyze model failures and performance and improve them through better data, training, and experimentation
- Work with GPUs and distributed training to efficiently train models at scale
What We're Looking For
Model Training Depth: You've actually trained or fine-tuned ML models. You understand the training process, data, objectives, evaluation, and trade-offs. You're comfortable going beyond simply using or deploying existing Hugging Face models.
Audio / Speech Experience: You've worked with speech, audio, TTS, ASR, Speech-to-Speech, generative audio, or related ML problems. Experience with diffusion-based audio models is a strong plus.
Research & Experimentation: You can take an idea from a research paper, implement it, design experiments, and determine what actually improves the model. You don't necessarily need to invent new architectures, but you should be comfortable going deep into how models work.
Battle-Tested Experience: You have 4-5+ years of serious ML engineering or research experience, ideally in the "trenches" of early-stage startups or high-velocity product teams.
Owners, Not Employees: You don't wait for a ticket. You identify what needs to improve, run the experiments, and ship models that drive the product forward. You measure success by better models, better performance, and happier users.
The Rewards
- Impact: Your models will directly power millions of voice interactions. You're building the core engine, not a side experiment.
- High-Growth Environment: We're angel-backed and scaling fast. This is a front-row seat to building a category-defining voice AI company.
- Culture: A high-intensity, high-trust team. No politics, no fluff—just building and winning.
Interested?
If you've spent the last few years training models, fine-tuning them, breaking them, figuring out why they failed, and making them better, let's talk.
About the company
Similar Jobs








