Avatar for SDVentures
SDVentures
Actively Hiring

Lead ML Engineer - LLM Inference & NLP

  • Remote (
    Everywhere
    )
  • |Full Time
Posted: 6 days ago• Recruiter recently active
Hires remotely in
Everywhere
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

RelocationAllowed
Skills
Dpo
NLP
distributed inference
PyTorch
CV
Quantization
Transformers
Language Models
vLLM
Inference
LLM Inference
SGLang
RLHF
TensorRT-LLM
Validation Teams
Post-Training
Fine-Tuning LLMs
KV Cache
Tensor Parallelism
Pipeline Parallelism
Moe
Batching
Speculative Decoding
Model Quality
GPU Profiling
KV Caching
GPU Servers
Prefix Caching
Agent Harnesses
Expert Parallelism
Attention Kernels
Multi-Node Setups
Training LLMs
Multi-GPU Setups
Content Teams
Serving Code
Chat Algorithm
ML Roadmap
Dataset Preparation Teams
Training of Large Models
Multi-Node GPU Clusters

About the job

Social Discovery Group (SDG) is a group of social discovery companies. SDG solves the problems of loneliness, isolation, and disconnection - transforming virtual intimacy into the new normal. SDG’s products redefine the way people interact and connect with one another.

Our portfolio includes social entertainment platforms designed to connect people online across different cultures and regions of the world.

We bring together a team of like-minded people and IT professionals who specialize in creating and developing globally impactful social discovery products. Our international team of digital nomads works remotely from all over the world.

We’re proud to be a two-time “Great Place to Work” winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025).

We are looking for Lead ML Engineer -LLM Inference & NLP.

*Your main tasks will be:*

  • Speed up and scale LLM inference in production: SGLang, KV and prefix caching, batching, quantization, speculative decoding
  • Run distributed inference for very large models (up to 1T+ parameters) across multi-GPU and multi-node setups
  • Benchmark new GPU servers and hardware, bring them into production and adapt our serving code to them
  • Lead the NLP and CV teams technically: review experiments, set direction, and step in early when something is heading the wrong way
  • Train and fine-tune the language models, and improve the agent harnesses and chat algorithm that run on them
  • Track cutting-edge research and open-source work in inference and post-training, and turn it into the ML roadmap
  • Collaborate closely with the validation, content, and dataset preparation teams to design experiments and measure model quality

*We expect from you:*

  • Deep hands-on experience optimizing LLM inference in production with SGLang, vLLM, or TensorRT-LLM
  • Experience with distributed inference or training of large models: MoE, tensor/expert/pipeline parallelism, multi-node GPU clusters
  • Strong understanding of what makes inference fast: KV cache, attention kernels, batching, quantization, GPU profiling
  • Experience training and fine-tuning LLMs, including post-training (RLHF, DPO, or similar)
  • Proven technical leadership: you've guided engineers through reviews, mentoring and technical decisions while still writing code yourself
  • Proficiency with PyTorch, transformers, and related libraries
  • Experience at AI-focused startups or companies (Character AI, OpenAI, and similar is a strong plus)
  • Backend engineering experience (Python, Go, C#) and knowledge of scalable deployment systems is a significant advantage
  • Advanced English or Russian

Nice to have:

  • CUDA or Triton kernel development
  • A computer vision background. We also welcome strong CV leads who have accelerated large generative image or video models
  • Experience with multimodal LLMs
  • First-author papers or notable open-source work, e.g. contributions to SGLang, vLLM, or post-training libraries
  • A degree in CS, math, or physics from a strong program (MSc or PhD)

*What do we offer:*

  • *REMOTE OPPORTUNITY* to work full-time;
  • The *initial pay level* or *pay range* for this role will be shared with candidates during the recruitment process and before the commencement of employment;
  • *Vacation* 28 calendar days per year;
  • *7 wellness days per year* (time off) that can be used to deal with household issues, to lie down and recover without taking sick leave;
  • *Bonuses up to $5000* for recommending successful applicants for positions in the company;
  • 50% *payment for professional training, international conferences, and meetings;*
  • Corporate discount for *English lessons;*
  • ​*Health benefits.* According to the paychecks, if you are not eligible for corporate medical insurance, the company will compensate you with up to $ 1,000 gross per year per employee. This can be spent on self-purchase of health insurance or on doctor’s fees for yourself and close relatives (spouse, children);
  • ​*Workplace organization.* The company provides all employees with an equipped workplace and all the necessary equipment (table, armchair, wifi, etc.) in our offices or co-working locations. In the other locations, the company provides reimbursement of workplace costs up to $ 1000 gross once every 3 years, according to the paychecks. This money can be spent on the rent of the co-working room, on equipping the working place at home (desk, chair, Internet, etc.) during those 3 years;
  • *Internal gamified gratitude system:* receive bonuses from colleagues and exchange them for our merchandise, team building activities, massage certificates, etc.

*Sounds good? Join us now!*

Similar Jobs

Scale AI company logo
Scale AI
Accelerate the development of AI applications
Rocket Money company logo
Rocket Money
The Money App That Works for You
Tatari company logo
Tatari
Buy and measure ads for brands across linear & streaming TV
Kepler company logo
Kepler
The ground truth platform for AI
Standard Bots company logo
Standard Bots
Propel human productivity through the world’s most accessible robots
CookUnity company logo
CookUnity
Signature meals from award-winning chefs, delivered to your door
Merciv company logo
Merciv
The Future of Enterprise Intelligence. Read your data’s past. Write your company’s future
Crosby company logo
Crosby
AI-native law firm that combined AI & lawyers to redefine the speed of legal services
Alloy company logo
Alloy
Alloy helps banks and fintech companies make safe and seamless fraud, credit