Multimodal Modeling: Select, fine-tune, distill, and deploy VLMs and other multimodal models across video, audio, image, and text, and make them fast enough to be usable inside a real-time editor.
Harness Engineering: Own the layer between model and product — tool schemas, structured output, constrained decoding, retries, sandboxing, and failure recovery for agentic editing features.