
Embedding VC
Actively Hiring
An early-stage venture fund backing Generative AI startups
- Growing fastShowed strong hiring growth in the past month
Job Location
Hires remotely in
Visa Sponsorship
Not Available
RelocationNot Allowed
Hiring contact
Congxing Cai
Employee
Palo Alto
About the job
我们在训练自研视频生成基础模型(DiT / Flow Matching),需要一位既能搭起训练平台、又能把研究代码变成数百卡集群上稳定结果的工程师。你不只是用平台的人,更是建平台的人。
你会做
• 训练平台搭建:从作业调度、断点续训、监控告警到数据 / 权重流水线,把分散的脚本沉淀为团队可复用的训练基础设施。
• 数百卡规模的分布式训练:FSDP、张量并行、Context Parallel、Ulysses,把 MFU 推到合理水位。
• PB 级视频数据 pipeline:NVDEC 解码、VAE latent 缓存、变分辨率 bucket sampling。
• 显存与性能:FlashAttention、FP8 混合精度、Triton kernel、activation checkpoint 策略。
• 训练稳定性:loss spike 根因分析、断点秒级恢复、慢节点自动剔除。
硬性要求
• 精通 PyTorch distributed 与 CUDA 体系结构。
• 至少一个主流训练框架(Megatron / DeepSpeed / FSDP / TorchTitan)的源码级理解。
• ≥ 256 卡训练实战经验。
• 有从零或半程搭建训练平台 / 集群调度 / 训练工具链的经验。
加分
• DiT / Diffusion / 视频数据处理经验。
• 写过 Triton / CUTLASS kernel。
• 对 HunyuanVideo / Wan / CogVideoX 等开源项目有源码级了解。
我们提供
• 真实数百卡算力、把训练当工程问题的团队、合规边界内的开源 / 发表空间。
About the company
Similar Jobs

Archesys
Improving the government services that impact everyday lives

Astranis
Building next-generation internet satellites to get the world online

Scale AI
Accelerate the development of AI applications

Orchard Robotics
Securing America's food supply by building the AI farmer that automates our nation's farms

EliseAI
Building AI agents that transform complex healthcare and housing systems

Prime Intellect
open superintelligence and infra

NumeralHQ
Sales tax on autopilot for Ecommerce & SaaS ✨ Spend 5 mins or less per month on compliance

Postman
Postman is the world’s leading collaboration platform for API development

Mercor
Mercor is at the intersection of labor markets and AI research