Kalpa Labs Ai careers
"Kalpa Labs is building the GPT-3 moment for speech.
Today's audio AI is fragmented — separate models for STT, TTS, voice cloning, dubbing, conversational agents, and music. This ""Curse of Specialization"" creates brittle workflows, no context carryover across tasks, and zero system-prompt steerability.
We're training a single generalist speech foundation model with the steerability of an LLM and the flexibility of in-context learning. One model that can transcribe, generate, clone, dub, edit, sing, and reason across audio — directed with natural-language prompts the way you'd brief a sound engineer.
What's already working:
• Pretrained models from 800M to 4.8B parameters on 2M hours of mixed-domain audio.
• Architectural wins that make audio nearly as cheap to train on as text — our 800M model cost under $1,000 to train.
• Emergent capabilities: contextual accent switching from spoken context, prosodic stress and emotion without explicit tags, disfluency handling, broad voice diversity.
• Demos competitive with ElevenLabs V3 at a fraction of the training cost.
What we're building toward:
• Speech-in / speech-out with system prompts and audio-in-context examples.
• Cross-task generalism: ""sing this verse in my voice,"" ""explain this chart in a calm tone,"" ""dub this clip preserving the speaker's emotion.""
• Real-time inference with graceful latency/quality trade-offs.
Founded in 2025 by Prashant Shishodia (ex-Google, full-stack ML lead for Google Assistant; scaled Gemini variants to billions of queries/month) and Gautam Jha (ex-QRT and Squarepoint Capital, built nanosecond-latency infrastructure for high-frequency trading).
Backed by Y Combinator (F25) and Nexus Venture Partners. Based in San Francisco."
Today's audio AI is fragmented — separate models for STT, TTS, voice cloning, dubbing, conversational agents, and music. This ""Curse of Specialization"" creates brittle workflows, no context carryover across tasks, and zero system-prompt steerability.
We're training a single generalist speech foundation model with the steerability of an LLM and the flexibility of in-context learning. One model that can transcribe, generate, clone, dub, edit, sing, and reason across audio — directed with natural-language prompts the way you'd brief a sound engineer.
What's already working:
• Pretrained models from 800M to 4.8B parameters on 2M hours of mixed-domain audio.
• Architectural wins that make audio nearly as cheap to train on as text — our 800M model cost under $1,000 to train.
• Emergent capabilities: contextual accent switching from spoken context, prosodic stress and emotion without explicit tags, disfluency handling, broad voice diversity.
• Demos competitive with ElevenLabs V3 at a fraction of the training cost.
What we're building toward:
• Speech-in / speech-out with system prompts and audio-in-context examples.
• Cross-task generalism: ""sing this verse in my voice,"" ""explain this chart in a calm tone,"" ""dub this clip preserving the speaker's emotion.""
• Real-time inference with graceful latency/quality trade-offs.
Founded in 2025 by Prashant Shishodia (ex-Google, full-stack ML lead for Google Assistant; scaled Gemini variants to billions of queries/month) and Gautam Jha (ex-QRT and Squarepoint Capital, built nanosecond-latency infrastructure for high-frequency trading).
Backed by Y Combinator (F25) and Nexus Venture Partners. Based in San Francisco."
Kalpa Labs Ai hasn't added any jobs yet
Get notified when Kalpa Labs Ai posts new jobs.
Prashant Shishodia
Valuation
$0
Funded over
1 round
Latest round
Seed (Sep 2025)
Smart Skin Technologies • Fredericton • 3 days ago
Innova • Irvine • 3 days ago
Betterworks • Menlo Park • 5 days ago
adRise • Old Toronto • 1 week ago
Movable Ink • New York City • $123k – $160k • 1 week ago


