
Local LLM / Inference Engineer — llama.cpp / Rust
- $120k – $160k • No equity
- |Remote ()
- |4 years of exp
- |Contract
Remote only
Not Available
About the job
We are an early-stage AI startup building a new kind of privacy-first personal AI platform designed to run primarily on the user's own computer.
We're looking for an experienced engineer with deep, hands-on expertise working with local, open-weight large language models. This is not primarily a model-training, cloud AI, or LLM API integration role. We're looking for someone who understands the practical realities of running LLMs locally and getting the best possible performance, quality, and reliability from them.
Our ideal candidate has substantial real-world experience with llama.cpp / llama-server and enjoys working close to the inference stack.
We're flexible about the working relationship and are open to full-time, part-time, or contract arrangements with the right person. Finding someone with the right expertise and enthusiasm for the problem is more important to us than the particular arrangement.
San Francisco Bay Area candidates are strongly preferred. We work primarily remotely, but we'd like someone local enough to collaborate in person periodically.
Similar Jobs









