Avatar for quadric.io
quadric.io
Actively Hiring
  • B2B
  • Growth Stage
    Expanding market presence
  • Top Investors
    This company has received a significant amount of investment from top investors

AI Inference Engineer

Posted: 8 months ago
Job Location
Remote Work Policy

In office - WFH flexibility

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
C++
System Architecture
Profiling
Benchmarking
PyTorch
ONNXRuntime
Model Quantization
vLLM
LlamaCPP
PTQ
QAT
AI Model Algorithms
AI Toolchains
Huggingface-Transformer
Neural-Compressor
Model Accuracy Measures
Model Inference Performance Profiling
Deployment Optimization

About the job

Quadric has created an innovative general purpose neural processing unit (GPNPU) architecture. Quadric's co-optimized software and hardware is targeted to run neural network (NN) inference workloads in a wide variety of edge and endpoint devices, ranging from battery operated smart-sensor systems to high-performance automotive or autonomous vehicle systems. Unlike other NPUs or neural network accelerators in the industry today that can only accelerate a portion of a machine learning graph, the Quadric GPNPU executes both NN graph code and conventional C++ DSP and control code.

Role:

The AI Inference Engineer in Quadric is the key bridge between the world of AI/LLM models and Quadric unique platforms. The AI Inference Engineer at Quadric will [1] port AI models to Quadric platform; [2] optimize the model deployment for efficient inference; [3] profile and benchmark the model performance. This senior technical role demands deep knowledge of AI model algorithms, system architecture and AI toolchains/frameworks.

Responsibilities:

  • Quantize, prune and convert models for deployment
  • Port models to Quadric platform using Quadric toolchain
  • Optimize inference deployment for latency, speed
  • Benchmark and profile model performance and accuracy
  • Develop tools to scale and speed up the deployment
  • Make Improvement to SDK and runtime
  • Provide technical support and documents to customers and developer community

Requirements:

  • Bachelor’s or Master’s in Computer Science and/or Electric Engineering.
  • 5+ years of experience in AI/LLM model inference and deployment frameworks/tools
  • experience with model quantization (PTQ, QAT) and tools
  • experience with model accuracy measures
  • experience with model inference performance profiling
  • experience with at least one of the following frameworks: onnxruntime, Pytorch, vLLM, huggingface-transformer, neural-compressor, llamacpp
  • Proficiency in C/C++ and Python
  • Demonstrate good capability in problem solving, debug and communication

  • Health Care Plan (Medical, Dental & Vision)

  • Retirement Plan (401k, IRA)

  • Life Insurance (Basic, Voluntary & AD&D)

  • Paid Time Off (Vacation, Sick & Public Holidays)

  • Family Leave (Maternity, Paternity)

  • Short Term & Long Term Disability

  • Training & Development

  • Work From Home

  • Free Food & Snacks

  • Stock Option Plan

About the company

quadric.io company logo

quadric.io

Actively Hiring
51-200 Employees
  • B2B
  • Growth Stage
    Expanding market presence
  • Top Investors
    This company has received a significant amount of investment from top investors
Learn more about quadric.io image

Funding

AMOUNT RAISED
$27.3M
FUNDED OVER
3 rounds
Rounds
B
$10000000
Series B - Dec 2022+2

Founders

Veerbhan Kheterpal
CEO
San Francisco
image
Daniel Firu
Co-founder
San Francisco
image
Nigel Drego
CTO • 9 years
image
View the team image