
Realtime Engineer, Multi-Modal AI
- $60k – $120k • 1.0% – 3.0%
- |Remote (Everywhere)
- |5 years of exp
- |Full Time
Remote only
Not Available
About the job
Realtime Engineer, Multi-Modal AI
T4R Labs | Palo Alto, CA (or Remote) | Full-time
About T4R Labs
T4R Labs is a deep tech lab building at the frontier of AI. We work on hard, unsolved problems where real-time performance, multi-modal understanding, and safety all have to hold up under production load — not just in a demo.
The Role
We're looking for an experienced realtime engineer to help us build systems that process and respond to audio, video, and text simultaneously, with the latency budgets of a live conversation rather than a batch job. You'll work at the intersection of low-latency infrastructure and multi-modal AI models, owning the pipelines that turn raw sensory streams into model inputs and model outputs into something a user experiences as instant.
This is a hands-on, high-ownership role. You'll be contributing to architecture decisions, not just implementing someone else's spec.
What You'll Do
- Design and build low-latency, high-throughput pipelines for streaming audio, video, and text into and out of multi-modal models
- Optimize end-to-end latency across the stack — networking, serialization, model inference, and rendering
- Work closely with ML engineers to co-design model interfaces that are realtime-friendly (streaming inference, chunked generation, interruption handling, etc.)
- Debug and eliminate jitter, dropped frames, and tail latency in production systems
- Build the monitoring and tooling needed to keep a realtime system observable and reliable
- Make pragmatic tradeoffs between quality, latency, and cost as the product scales
What We're Looking For
- 5+ years of experience building realtime or low-latency systems (e.g., video/voice conferencing, live streaming, gaming netcode, trading systems, or similar)
- Strong systems programming skills (C++, Rust, Go, or similar) and comfort reasoning about performance at the network, OS, and hardware level
- Experience with streaming protocols (WebRTC, RTP, gRPC streaming, or similar)
- Familiarity with the practical constraints of serving ML models in production — batching, quantization, streaming inference
- A track record of shipping systems that hold up under real user load, not just in benchmarks
- Comfort working in a fast-moving, early-stage environment with ambiguous specs
Bonus points for:
- Direct experience integrating multi-modal (audio/video/text) AI models into a live product
- Experience with GPU-accelerated inference serving (Triton, TensorRT, vLLM, or similar)
- Prior work on voice assistants, live translation, or interactive AI agents
Why T4R Labs
- Work directly with the founding team on problems that don't have off-the-shelf solutions
- High autonomy, high impact — your decisions ship
- We're seed stage: this role is equity only for now, with salary starting once we close funding
How to Apply
Reach out with your resume and a note on the most interesting realtime system you've built.
About the company

T4R Labs
Similar Jobs



