
- Recently fundedRaised funding in the past six months
Applied Research Engineer / Research Intern
- $15k – $20k • 1.0% – 5.0%
- |Remote ()
- |2 years of exp
- |Full Time
Remote only
Not Available
About the job
Research Engineer Intern – Voice AI and Conversational Architectures
**Please note this is a remote role (Preferably Dallas or India) ideally for internship position, but open to full time consulting for the rught person.
About us
We are VoiceMonk (voicy.ai), a foundational research and intellectual property firm building the core architectures for real-time voice AI. With 14 granted patents in the United States, 2 granted patents in India, and 20 pending applications, our conversational AI research spans over a decade. We invented several of the core dialogue, streaming, and latency-reduction technologies driving today's mass adoption of multimodal voice agents.
The role
You will sit at the critical boundary between state-of-the-art voice AI research and patent engineering. You will read the latest literature on full-duplex Spoken Language Models (FD-SLMs), speech-to-speech (S2S) architectures, and real-time streaming, reproduce what matters, and build functional reference implementations of our proprietary architectures.
This is not a standard product-engineering role. You will act as the technical engine for our licensing and IP strategy—translating cutting-edge latency reduction and turn-taking models into rigorous technical specifications, evidentiary prototypes, and infringement analyses.
What you'll work on
- Reference Architectures & Evidentiary Prototypes: Design and deploy functional testbeds of our patented claims (e.g., real-time dialogue management, interruption/barge-in handling, and context retrieval). These prototypes serve as direct evidentiary validation for our licensing campaigns and defend against Alice invalidation challenges.
- Full-Duplex & Speech-to-Speech (S2S) Systems: Implement ultra-low latency voice-to-voice pipelines, streaming ASR-LLM-TTS cascades, and Voice Activity Detection (VAD) optimizations to benchmark against current industry standards.
- Technical Infringement Analysis: Analyze and reverse-engineer state-of-the-art enterprise voice agents and LLM telephony stacks to map commercial implementations directly to our existing patent claims.
- Continuation Patent Engineering: Translate architectural breakthroughs into rigorous technical whitepapers. You will work directly with our patent counsel to provide the engineering specifications necessary to file future-proof continuation patents mapping to modern LLM architectures.
- Real-Time Inference and Benchmarking: Measure Time to First Audio (TTFA), end-to-end latency across edge-cloud boundaries, and orchestration timing to prove the technical superiority of our proprietary conversational loops.
Minimum qualifications
You should meet one of the following:
- MS in Computer Science, Machine Learning, Audio Engineering, or a related field, plus 2+ years of research experience (industry research, a research lab, or published work).
- PhD in a related field, plus 1+ year of research experience.
- Currently 3+ years into a PhD program (for the internship track), with a track record of independent research.
You should also have:
- Strong Python: You should write clean, reproducible, and tested code to build robust reference architectures, not just disposable research scripts.
- Audio/ML Ecosystem Experience: Hands-on experience with modern ML frameworks (PyTorch, JAX) and the current voice AI ecosystem (streaming STT/TTS, WebRTC, LLM orchestration).
- Demonstrated Research Output: Publications, preprints, a thesis, or a reproducible body of work in speech processing, conversational AI, or low-latency ML systems.
- Self-Directed Execution: The ability to go from an academic paper (or a legal patent claim) to a working prototype independently. We will give you a technical or legal objective, not a detailed spec.
Nice to have
- Experience with full-duplex Spoken Language Models, audio codecs, or native audio-in/audio-out architectures.
- Work on inference optimization specific to voice: KV-cache management, speculative decoding, or streaming chunk orchestration.
- Familiarity with enterprise telephony infrastructure (SIP trunks, WebRTC streaming).
- Experience interfacing with legal, patent, or compliance teams on technical matters.
- Prior startup experience, or evidence that you thrive in an unstructured, highly autonomous environment.
What we offer
- Direct collaboration with the founder to shape the technical and architectural direction of a massive IP portfolio.
- A unique hybrid role combining bleeding-edge ML research with high-stakes technology licensing and patent strategy.
- Highly competitive cash compensation, supplemented by significant performance bonuses directly tied to the successful technical validation and licensing of our IP.
- Support and resources for publishing technical whitepapers and open research.
- Interns get a defined architectural project, direct mentorship, and a serious shot at a full-time offer.
About the company
Similar Jobs





