Software Quality Engineer
- $130k – $160k
- |
- |3 years of exp
- |Full Time
In office
Not Available
About the job
ABOUT GUAVA
Guava is the voice platform built for regulated industries: the calls that have to be right. Trained on 10B+ live agent minutes and in regulated production since 2013, we give healthcare, banking, insurance, and BPO customers a single compliant system for every second of a call. We train our own ASR, LLM, and TTS models in house and ship them as a developer SDK. The team comes out of Stanford, MIT, and NASA, and we are growing fast.
THE ROLE
You will test systems that do not behave like ordinary software. A voice bot does not return a value you can assert against. It answers differently every time, over a real audio channel, at conversational latency, while a caller talks over it. Deciding what "correct" means for that and building the infrastructure that proves it is the job.
Three areas:
Voice bots and spoken dialog systems. Design and own the eval sets that catch conversational regressions before customers do. Test turn-taking, barge-in, endpointing, and latency under real load. Drive automated calls through real telephony, not mocked audio. Probe what happens when the model is wrong, the caller is hostile, or the line is bad.
The SDK. We ship client libraries in several languages. Test them across languages, runtimes, and platforms. Validate docs and quickstarts by building with them. Be the first person to find out when the developer experience is bad, and say so.
Operations support. Reproduce what customers report, including the issues that only appear on one carrier at high concurrency, and turn them into permanent coverage.
WHAT WE'RE LOOKING FOR
– 3+ years in software quality, test engineering, or SDET work on production systems you were accountable for
– Strong Python, plus the ability to pick up whatever language an SDK is written in
– You have built test infrastructure other engineers depend on, not just written test cases
– Rigor about evidence. You are skeptical of a green suite, you know flaky from broken, and you can separate a model regression from a bug
– Comfort with ambiguity. Much of what you test has no obvious pass/fail line, and you will be the one who draws it
– A good ear and real patience for listening to calls. The audio is the best signal in this product
Nice to have: experience testing ML or LLM systems, familiarity with ASR/TTS metrics, telephony or real-time audio (SIP, WebRTC, codecs, packet loss), load testing for latency-sensitive systems, or prior work against a compliance bar.
WHY THIS ROLE
This is one of the few jobs where you evaluate frontier speech models every day without having trained one first. You will sit next to the ML engineers who build them, and your evals will shape what ships. Several of the hardest open problems in voice AI are measurement problems, and there is no settled playbook for any of this. You will write it.
The tradeoff is that the ground moves. Models change under you, the product ships fast, and nobody is going to hand you a test plan. If that sounds like the fun part, we should talk.
Similar Jobs









