
Machine Learning Engineer, Multimodal
- $190k – $210k • 0.1% – 0.2%
- |
- |4 years of exp
- |Full Time
In office
Not Available
About the job
About Wizard
Wizard is an AI-native creative suite — a video editing platform built by editors, for editors. Our team comes from Apple, Disney, Netflix, Adobe, Sony, Google, and ILM, and we’re building the tools we always wished we had. Small team, high bar, ships constantly. We’re reimagining professional editing around what AI now makes possible, and we’re doing it fast.
The Role
As a Machine Learning Engineer at Wizard, you will design, train, evaluate, and ship the
models behind every intelligent feature in the product. You will work directly with the
founding ML engineer and own problems end to end, from architecture decisions through
to what actually runs on a user's laptop.
The scope is deliberately broad. One day you are profiling a vision encoder on Metal to hit
frame budget on an M-series machine. The next you are rebuilding the tool-calling harness
so the editing agent stops inventing timeline operations that don't exist. The week after
that you are designing the evaluation suite that tells us whether either change actually
helped. We are not hiring someone to own a single model. We are hiring someone to own
the path from research to product.
Multimodal depth is the core requirement, not a nice-to-have. Everything in a video editor
is multimodal by definition — frames, waveforms, transcripts, and user intent all have to be
reasoned about together. Much of the work lives inside vision-language models: how they
encode frames, where they lose temporal information, and what it takes to make them
reliable enough to drive an edit. Candidates whose experience is text-only LLM work will
not be a fit.
Key Reponsibilities
- Multimodal Modeling: Select, fine-tune, distill, and deploy VLMs and other multimodal models across video, audio, image, and text, and make them fast enough to be usable inside a real-time editor.
- Harness Engineering: Own the layer between model and product — tool schemas, structured output, constrained decoding, retries, sandboxing, and failure recovery for agentic editing features.
- Evaluation Infrastructure: Build o!ine evals, regression suites, trace capture and replay, and metrics that distinguish a real feature from a good demo, including for tasks where ground truth is subjective.
- On-Device Inference: Optimize inference on Metal — quantization, kernel-level profiling, memory pressure, and cold-start latency on consumer hardware.
- Model Interface Layer: Maintain a pluggable, swappable interface that presents the same contract whether a model runs locally or in the cloud.
- Data Pipelines: Build curation, labeling, and synthetic generation pipelines for video and audio.
- End-to-End Ownership: Take projects from prototype through production, monitoring, and the inevitable second rewrite.
- Editor Collaboration: Work directly with the editors on our team to translate professional creative workflows into model behavior.
Minimum Qualifications
- 3+ years building and deploying machine learning systems in production.
- Deep multimodal experience, especially with VLMs. You have trained, fine-tuned, evaluated, or shipped vision-language models and other systems that reason jointly over video, images, or audio alongside text — not an LLM with an image endpoint attached.
- Working knowledge of transformer internals. Attention variants and their cost, KV cache behavior, positional encoding schemes, what degrades as context grows, what changes when you swap the vision tower, and where the FLOPs actually go.
- Demonstrated depth on harnesses. You have built them, not just called them: agent loops, tool-calling scaffolds, structured output validation, and eval harnesses with reproducible traces. You can look at a failed run and identify precisely which layer broke.
- Strong Python and PyTorch, with comfort reading and debugging model code at the layer level.
- Experience designing evaluation methodologies, benchmarks, and performance metrics.
- Experience with inference optimization: quantization, distillation, batching, and latency budgeting.
- A genuine generalist streak. Metal inference on Monday, LLM tool calls on Tuesday, a data pipeline on Wednesday — and none of it feels like a detour.
Preferred Qualifications
- Video understanding, temporal consistency, or di"usion and video generation models.
- On-device and cross-platform inference runtimes: CoreML, MLX, ONNX Runtime, TensorRT, or hand-written compute kernels.
- Speech-to-text and ASR systems at production quality.
- Open-source contributions to inference runtimes, agent frameworks, or evaluation tooling.
- Experience in a startup or otherwise fast-moving environment.
- Real hours in Premiere, Resolve, or After E"ects, with opinions about all three.
Why Wizard?
- Solve Hard Problems: Real-time multimodal inference on consumer hardware, with a latency budget measured in frames rather than seconds.
- Ship to People Who Notice: Professional editors are the most demanding users in software. They will tell you immediately when a feature is fake.
- Ownership and Impact: Founding team, no layers, no bureaucracy. Your work is the product, not a service behind it.
- Build the Craft, Not Just the Model: You will sit next to people who have cut features, shows, and campaigns for a living, and learn what "good" actually means in this domain.
Similar Jobs









