
AI / Machine Learning Engineer (Voice AI, Agentic Systems & Telephony)
- €55k – €84k • 0.0% – 2.0%
- |Remote (+5) •
- |3 years of exp
- |Full Time
Onsite or remote
Available
About the job
About Menodi (Menodi.se)
Menodi (https://menodi.se) is an AI receptionist and voice-automation platform built for Swedish and Nordic service businesses. Menodi replaces missed calls and manual telephone scheduling with autonomous, human-level AI receptionists that handle incoming calls 24/7, integrate directly into business booking systems, and converse naturally in Swedish.
We are engineering low-latency conversational AI and robust telephony infrastructure to modernize how local companies, clinics, trades, and service providers handle customer calls.
Role Overview:
As an AI / Machine Learning Engineer at Menodi.se, you will lead the core intelligence, streaming pipelines, and agent orchestration powering our AI phone receptionist platform. You will build and scale the real-time agentic loop—connecting automated speech recognition (ASR), multi-step LLM function calling, and zero-latency text-to-speech (TTS) to real-world phone networks.
Key Responsibilities:
Real-Time Voice & Agent Architecture: Develop ultra-low-latency agentic loops (tool use, function calling, stateful dialog flow) engineered specifically for voice telephony.
Telephony & Real-Time Streaming: Integrate LLM agent workflows with SIP/telephony backends, WebSockets, and real-time streaming media pipelines (FreeSWITCH, Twilio, WebRTC).
Nordic Language & Swedish Voice Optimization: Benchmark, fine-tune, and optimize speech models and frontier LLMs for Swedish conversational flow, regional accents, and domain-specific terminology.
Autonomous Scheduling & System Integration: Build resilient agent tools that execute zero-shot calendar bookings, customer lookup, and CRM sync during live customer calls.
Evaluation & Latency Engineering: Reduce end-to-end voice roundtrip latency below conversational thresholds while hardening guardrails, deterministic fallbacks, and multi-turn accuracy.
Requirements:
3+ years of backend and systems engineering experience with significant focus on shipping LLM/AI applications to production.
Deep production proficiency in Python alongside Rust or TypeScript/Node.js.
Strong hands-on background in Agentic AI: tool calling, dynamic context injection, structured outputs, and real-time state machines.
Experience architecting asynchronous event-driven backends, WebSockets, and streaming architectures.
Working knowledge of speech-to-text (STT), text-to-speech (TTS), and audio streaming pipelines.
**Nice to Have:*
Familiarity with SIP, VoIP, WEBrtc, FreeSWITCH, Asterisk, or programmable voice APIs.
Experience deploying voice agents or conversational bots in Scandinavian languages (Swedish).
Experience with open-source LLM hosting, local inference engines (vLLM), and context caching strategies.
About the company
