Avatar for Wizard
Wizard
Actively Hiring
A new creative suite built by editors, for editors

Machine Learning Engineer, Multimodal

  • $190k – $210k • 0.1% – 0.2%
  • |
  • |4 years of exp
  • |Full Time
Posted: today• Recruiter recently active
Job Location
Remote Work Policy

In office

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
PyTorch
Multimodal Models
Visual Language Models (VLMs)

About the job

About Wizard
Wizard is an AI-native creative suite — a video editing platform built by editors, for editors. Our team comes from Apple, Disney, Netflix, Adobe, Sony, Google, and ILM, and we’re building the tools we always wished we had. Small team, high bar, ships constantly. We’re reimagining professional editing around what AI now makes possible, and we’re doing it fast.

The Role
As a Machine Learning Engineer at Wizard, you will design, train, evaluate, and ship the
models behind every intelligent feature in the product. You will work directly with the
founding ML engineer and own problems end to end, from architecture decisions through
to what actually runs on a user's laptop.

The scope is deliberately broad. One day you are profiling a vision encoder on Metal to hit
frame budget on an M-series machine. The next you are rebuilding the tool-calling harness
so the editing agent stops inventing timeline operations that don't exist. The week after
that you are designing the evaluation suite that tells us whether either change actually
helped. We are not hiring someone to own a single model. We are hiring someone to own
the path from research to product.

Multimodal depth is the core requirement, not a nice-to-have. Everything in a video editor
is multimodal by definition — frames, waveforms, transcripts, and user intent all have to be
reasoned about together. Much of the work lives inside vision-language models: how they
encode frames, where they lose temporal information, and what it takes to make them
reliable enough to drive an edit. Candidates whose experience is text-only LLM work will
not be a fit.

Key Reponsibilities

  • Multimodal Modeling: Select, fine-tune, distill, and deploy VLMs and other multimodal models across video, audio, image, and text, and make them fast enough to be usable inside a real-time editor.
  • Harness Engineering: Own the layer between model and product — tool schemas, structured output, constrained decoding, retries, sandboxing, and failure recovery for agentic editing features.
  • Evaluation Infrastructure: Build o!ine evals, regression suites, trace capture and replay, and metrics that distinguish a real feature from a good demo, including for tasks where ground truth is subjective.
  • On-Device Inference: Optimize inference on Metal — quantization, kernel-level profiling, memory pressure, and cold-start latency on consumer hardware.
  • Model Interface Layer: Maintain a pluggable, swappable interface that presents the same contract whether a model runs locally or in the cloud.
  • Data Pipelines: Build curation, labeling, and synthetic generation pipelines for video and audio.
  • End-to-End Ownership: Take projects from prototype through production, monitoring, and the inevitable second rewrite.
  • Editor Collaboration: Work directly with the editors on our team to translate professional creative workflows into model behavior.

Minimum Qualifications

  • 3+ years building and deploying machine learning systems in production.
  • Deep multimodal experience, especially with VLMs. You have trained, fine-tuned, evaluated, or shipped vision-language models and other systems that reason jointly over video, images, or audio alongside text — not an LLM with an image endpoint attached.
  • Working knowledge of transformer internals. Attention variants and their cost, KV cache behavior, positional encoding schemes, what degrades as context grows, what changes when you swap the vision tower, and where the FLOPs actually go.
  • Demonstrated depth on harnesses. You have built them, not just called them: agent loops, tool-calling scaffolds, structured output validation, and eval harnesses with reproducible traces. You can look at a failed run and identify precisely which layer broke.
  • Strong Python and PyTorch, with comfort reading and debugging model code at the layer level.
  • Experience designing evaluation methodologies, benchmarks, and performance metrics.
  • Experience with inference optimization: quantization, distillation, batching, and latency budgeting.
  • A genuine generalist streak. Metal inference on Monday, LLM tool calls on Tuesday, a data pipeline on Wednesday — and none of it feels like a detour.

Preferred Qualifications

  • Video understanding, temporal consistency, or di"usion and video generation models.
  • On-device and cross-platform inference runtimes: CoreML, MLX, ONNX Runtime, TensorRT, or hand-written compute kernels.
  • Speech-to-text and ASR systems at production quality.
  • Open-source contributions to inference runtimes, agent frameworks, or evaluation tooling.
  • Experience in a startup or otherwise fast-moving environment.
  • Real hours in Premiere, Resolve, or After E"ects, with opinions about all three.

Why Wizard?

  • Solve Hard Problems: Real-time multimodal inference on consumer hardware, with a latency budget measured in frames rather than seconds.
  • Ship to People Who Notice: Professional editors are the most demanding users in software. They will tell you immediately when a feature is fake.
  • Ownership and Impact: Founding team, no layers, no bureaucracy. Your work is the product, not a service behind it.
  • Build the Craft, Not Just the Model: You will sit next to people who have cut features, shows, and campaigns for a living, and learn what "good" actually means in this domain.

About the company

Wizard company logo

Wizard

Actively Hiring
A new creative suite built by editors, for editors1-10 Employees
Company Size
1-10
Company Type
Software Development
Learn more about Wizard image

Founders

Jason Carman
Founder
image
View the team image

Similar Jobs

Nextdoor company logo
Nextdoor
Nextdoor is the private social network for your neighborhood
Scale AI company logo
Scale AI
Accelerate the development of AI applications
Wonderschool company logo
Wonderschool
Quality in-home child care and preschools near you
Orchard Robotics company logo
Orchard Robotics
Securing America's food supply by building the AI farmer that automates our nation's farms
Rocket Money company logo
Rocket Money
The Money App That Works for You
SmileShape company logo
SmileShape
SmileShape is using the forefront of AI to better digital dentistry
Archesys company logo
Archesys
Improving the government services that impact everyday lives
tribe.ai company logo
tribe.ai
We embed elite AI engineers to ship real, production-grade AI for enterprises