Avatar for Loft Labs
Loft Labs
Actively Hiring
Open-Source Developer Tools & Platform Engineering Software For Kubernetes
  • B2B
  • Growth Stage
    Expanding market presence

Senior Inference Engineer

Posted: 4 weeks ago
Job Location
Remote Work Policy

In office - WFH flexibility

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
Caching
routing
Golang
Quantization
vLLM
SGLang
TensorRT-LLM
Batching

About the job

As a Senior Inference Engineer at vCluster Labs, you are the first engineer we're hiring to own inference. You'll partner directly with our CTO to build the platform's inference layer from the ground up, taking models and turning them into a production-grade, query-to-response pipeline running at scale. From there, you will help lead the engineering direction of inference at vCluster, partnering with Product to shape what we build next as the space evolves.

As a Senior Inference Engineer, Your Role Will Include

  • Deploying models to production: Take LLMs and put them into production across one or more machines on GPU infrastructure, owning the full pipeline from a customer's query to the served response.
  • Serving frameworks: Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
  • Optimizing for scale: Apply quantization, batching, caching, and routing to keep latency and cost in check as traffic grows.
  • Programming, not just configuring: Build real infrastructure in Python or Golang — this is an engineering role, not a research or data-science one.
  • Owning the roadmap: Build the first iteration alongside our CTO, then take the lead on the inference platform and partner with Product to decide what we build next.

This role could be a fit for you if you bring:

  • Production LLM serving experience: You've deployed and served LLMs using vLLM, SGLang, or TensorRT-LLM, ideally at a company built around inference at scale.
  • Inference optimization know-how: Hands-on experience with quantization, batching, caching, and routing, not just familiarity with the terms.
  • Hands-on programming experience: Strong engineering skills in Python or Golang, with real production code experience.
  • Communication: Strong communication skills, explaining technical concepts clearly to both engineers and non-technical stakeholders.

Bonus Points For

  • Familiarity with containerized environments (Docker, Kubernetes)
  • Hands-on generative AI experience with common ML frameworks (PyTorch, Transformers)
  • Good understanding of the GPU stack: CUDA, NCCL, drivers, and related libraries
  • Knowledge of model architectures and fine-tuning approaches
  • Experience with NVIDIA Dynamo

About VCluster Labs

We are a venture-backed tech startup and the company pioneering Kubernetes virtualization for the AI era. We raised +$30M from top-tier VCs such as Khosla Ventures (first investor in OpenAI, GitLab, Stripe, Doordash) and are in a hyper-growth phase looking for motivated people to complement our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe and we have a remote-first work culture.

We are the leading platform for operating GPU infrastructure, enabling AI Cloud providers to deliver a hyperscaler-like experience to their customers and AI factories that need to build that same experience for their internal teams. Our platform delivers the full operational stack operators need to run their GPU data centers — managed Kubernetes, fast isolated tenant provisioning, and automated node provisioning and lifecycle management — enabling them to accelerate time to value, reduce operational burden, and maximize the ROI of every GPU.

We're the company behind vCluster, an open-source technology for virtualizing Kubernetes (10k+ GitHub stars, 40M+ virtual clusters created since 2021). Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI — a Kubernetes-native framework purpose-built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.

Benefits

We offer the following benefits:

  • Competitive Salary: We offer a competitive compensation package, including equity.
  • Platinum-Level Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).
  • Flexible Working Schedule: You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.
  • Workplace Flexibility: We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.

Culture & Values

At VCluster Labs, We Value And Stand For

  • Make it Happen: We have a relentless bias for action and the grit to push through obstacles. We do whatever it takes to figure it out, put in the work, and ruthlessly prioritize the actions that drive measurable impact for the business.
  • Own the Outcome: We understand that our responsibility doesn't end when a task is checked off; it ends when the value is delivered. We connect our daily individual actions to the broader success of the company and our customers.
  • Create Wow: We measure success by the experience we generate, both inside and outside the company. For our customers, this means impressive speed and intuitive experiences. For our team, this means going the extra mile to support one another and to continuously drive each other to new heights.
  • Open Source, Open Mind: We are actively contributing to and maintaining open-source projects. Internally, we foster meritocracy — the strongest ideas win, no matter who or where they come from.
  • Build Tomorrow’s Standards, Intentionally: We don't just ship software; we define the state-of-the-art of tomorrow. We are fearless in tearing down old approaches to build something better, but we are disciplined in how we do it because we know our users rely on our technology to run mission-critical infrastructure platforms.

Compensation Range: $190K - $230K

About the company

Loft Labs company logo

Loft Labs

Actively Hiring
Open-Source Developer Tools & Platform Engineering Software For Kubernetes51-200 Employees
  • B2B
  • Growth Stage
    Expanding market presence

Employees joined from

Learn more about Loft Labs image

Funding

AMOUNT RAISED
$4.6M
FUNDED OVER
1 round
Round
S
$4600000
Seed - Sep 2021

Perks

Platinum-Level Insurance
Health, Dental, Vision, Life Insurance including plans for eligible dependents (depends on country)
Equity Incentive Plan
We know that you want to make an impact on the success of our company and nothing connects your success and the company's success better than making stock options a part of your total compensation.
Workplace Flexibility
If you want to work from home, that’s great! If you’d prefer to work in a coworking space, we can make that happen as well. If you want to relocate to our HQ in San Francisco, let’s discuss that. We’re flexible on all
Flexible Working Schedule
You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.

Founders

Lukas Gentele
Founder
San Francisco
image
View the team image

Similar Jobs

Archesys company logo
Archesys
Improving the government services that impact everyday lives
Astranis company logo
Astranis
Building next-generation internet satellites to get the world online
Scale AI company logo
Scale AI
Accelerate the development of AI applications
Orchard Robotics company logo
Orchard Robotics
Securing America's food supply by building the AI farmer that automates our nation's farms
EliseAI company logo
EliseAI
Building AI agents that transform complex healthcare and housing systems
Postman company logo
Postman
Postman is the world’s leading collaboration platform for API development
Mercor company logo
Mercor
Mercor is at the intersection of labor markets and AI research