Avatar for Bynario
Autonomous Application Security for software, from code to runtime

AI Engineer - Model Adaptation & Evaluation

  • €40k – €90k • No equity
  • |Remote ()
  • |2 years of exp
  • |Full Time
Posted: 3 months ago• Recruiter recently active
Hires remotely in
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

Preferred Timezones
Azores, Coordinated Universal Time, Central European Time, Eastern European Time, Turkey Time, Dubai Time
Collaboration Hours
9:30 AM - 4:30 PM Central European Time
RelocationAllowed
Skills
Remote Working
Large Language Models (LLMs)
LLM Fine Tuning/ Lora/ QLora

About the job

Bynario is building autonomous security systems that identify unknown vulnerabilities, prioritize real risk, and deploy fixes without human intervention. We focus on compiled binaries because that's what actually runs in production.

Built by renowned security researchers and academics, our platform delivers deep software understanding to give organizations security independence.

About the role

We are looking for an LLM Fine-Tuning Engineer to help adapt, evaluate, and deploy open-weight language models for specialized operational use cases.

The role focuses on the full post-training lifecycle: dataset design, supervised fine-tuning, parameter-efficient fine-tuning, synthetic data generation, model evaluation, and deployment support. You will work closely with engineering and domain teams to transform foundation models into reliable, task-specific systems that can operate effectively in constrained, local, or secure environments.

This is not a generic prompt engineering role. We are looking for someone with hands-on experience in model adaptation, training data quality, evaluation methodology, and practical deployment constraints.

What You'll Do

  • Fine-tune and adapt open-weight LLMs for specialized internal or customer-facing use cases.
  • Design, run, and compare different post-training approaches, including supervised fine-tuning (SFT), parameter-efficient fine-tuning such as LoRA or QLoRA, preference tuning approaches such as DPO where appropriate, and full fine-tuning when justified by the task and infrastructure constraints.
  • Build, clean, and improve high-quality datasets for training, validation, and evaluation.
  • Generate and validate synthetic data to expand coverage, improve robustness, and accelerate iteration cycles.
  • Define and implement evaluation methods that go beyond generic benchmark scores.
  • Analyze model regressions, failure modes, and behavioral changes after fine-tuning.
  • Work with engineering teams to package, serve, and operationalize tuned models in local, private, or restricted environments.
  • Collaborate with domain experts to translate real-world needs into measurable model requirements and evaluation criteria.

What We're Looking For

  • Strong practical experience fine-tuning LLMs or adjacent foundation models in production, applied research, or serious experimental environments.
  • Hands-on experience with multiple post-training techniques, ideally including SFT, LoRA, and QLoRA.
  • Experience building, curating, and maintaining datasets for model training and evaluation.
  • Experience generating, filtering, and validating synthetic data for model improvement.
  • Strong Python skills and practical experience with PyTorch, Hugging Face Transformers, tokenization, preprocessing, and training pipelines.
  • Good understanding of GPU constraints, memory/performance trade-offs, quantization-aware workflows, and training optimization.
  • Strong debugging mindset, with the ability to understand why a model improved, regressed, overfit, or failed.
  • Ability to work independently across research, data, and engineering concerns.

Nice to Have

  • Experience with secure, air-gapped, on-premise, or otherwise restricted deployment environments.
  • Experience deploying or serving fine-tuned models in production-like settings.
  • Familiarity with inference optimization, quantization, model serving, or local LLM runtimes.
  • Background in cybersecurity, software security, vulnerability research, or other technical domains requiring specialized model behavior.
  • Experience with distributed training or multi-GPU training frameworks.
  • Familiarity with evaluation harnesses, red-teaming, model safety testing, or task-specific benchmarks

How We Think About This Role

We are looking for someone who can think critically about when fine-tuning is the right
approach, and when prompting, retrieval, orchestration, or tooling may be more
appropriate.
The ideal candidate understands that successful model adaptation is not just about running
training jobs. It requires creating data that improves model behavior rather than simply
inflating metrics, evaluating outputs in ways that reflect real operational value,
understanding trade-offs between model quality, cost, latency, and deployment constraints,
and making open-weight models reliable, maintainable, and useful in real-world
environments

Compensation

  • 40k-90k EUR

Benefits

  • 1k allowance for home office setup
  • AI max plan of your choice
  • Company retreats in Italy

About the company

Bynario company logo
Autonomous Application Security for software, from code to runtime1-10 Employees
Company Size
1-10
Company Type
Artificial Intelligence
Company Type
Software Development
Company Type
Information Security
Company Type
Vulnerability Management
Learn more about Bynario image

Funding

AMOUNT RAISED
Undisclosed amount
FUNDED OVER
1 round
Round
PRE
Undisclosed amount
Pre-Seed - Jan 2026