Avatar for PolarGrid
PolarGrid
Actively Hiring
PolarGrid is building the future of AI inference
  • Growing fast
    Showed strong hiring growth in the past month

Inference Optimization Engineer

Posted: 1 week ago• Recruiter recently active
Hires remotely in
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

Preferred Timezones
Pacific Time, Mountain Time, Central Time, Eastern Time, Atlantic Time
RelocationAllowed
Hiring contact
Henry C
Founder • 7 months
Ottawa
image

About the job

About the Role

PolarGrid is building the infrastructure layer for real-time AI inference. We're looking for an Inference Optimization Engineer to squeeze every bit of performance out of our stack. You'll work directly on the systems that serve inference to our customers, making them faster, cheaper, and more efficient.

This is a deep technical, systems-focused role. You'll own latency, throughput, and cost per token as real metrics you're responsible for improving. You'll build the repeatable benchmarking and optimization process that takes new models and hardware from an initial baseline to a validated production configuration.


What You'll Do

  • Profile and optimize inference pipelines end to end using representative customer workloads, from request handling and scheduling through distributed GPU execution
  • Tune serving frameworks such as vLLM, TensorRT-LLM, and SGLang for specific latency, throughput, and cost targets
  • Build automated benchmarking and performance regression tooling across models, frameworks, precisions, hardware, and workload profiles
  • Characterize customer workloads and translate TTFT, ITL, concurrency, and context-length requirements into deployment configurations
  • Implement and evaluate quantization strategies across model families, measuring both performance gains and model-quality regressions
  • Work with hardware teams to match model configurations and parallelism strategies to GPU topology, NVLink, and interconnect bandwidth
  • Benchmark new hardware such as RTX Pro 6000s and B300s, identifying the best engine, precision, parallelism, and deployment configuration for each workload
  • Bring new model architectures into production, including checkpoint conversion, framework support, distributed configuration, and correctness validation
  • Contribute to continuous batching, speculative decoding, KV-cache optimization, prefill/decode disaggregation, and request-scheduling work
  • Read, debug, and modify inference framework internals when configuration-level tuning is not enough
  • Work with the platform team to canary performance improvements, measure them under production traffic, and turn successful configurations into repeatable deployment recipes

What We're Looking For

  • Strong GPU systems fundamentals, with the ability to work across Python, C++, CUDA, or Triton when optimization requires going below framework configuration
  • Hands-on experience with at least one major inference serving framework such as vLLM, TGI, TensorRT-LLM, or SGLang
  • ⁠Deep understanding of transformer architecture and where inference bottlenecks actually live
  • Ability to read, debug, and modify inference framework internals rather than treating them as black boxes
  • Experience building benchmarking, load-generation, or performance-regression infrastructure
  • Comfortable profiling with Nsight Systems, Nsight Compute, PyTorch Profiler, or similar tools
  • Experience with quantization and precision tradeoffs in production, including validating numerical correctness and model quality
  • Experience optimizing multi-GPU or multi-node inference across high-speed interconnects
  • Understanding of distributed inference, NCCL, GPU topology, and communication bottlenecks
  • ⁠You care about numbers: TTFT, ITL, P95/P99 latency, throughput, GPU utilization, and tokens/sec/dollar

Bonus Points

  • Experience writing custom CUDA or Triton kernels
  • Familiarity with speculative decoding, MoE routing optimizations, or prefill/decode disaggregation
  • Experience with inference request routing, scheduling, or admission control
  • ⁠Experience upstreaming performance improvements to vLLM, SGLang, TensorRT-LLM, or related projects
  • ⁠Open-source contributions to inference or ML systems projects

Why PolarGrid

You'll work on real hardware at scale, not toy benchmarks. The performance improvements you ship go directly to customers and directly affect our unit economics. Small team, real ownership.

About the company

PolarGrid company logo

PolarGrid

Actively Hiring
PolarGrid is building the future of AI inference11-50 Employees
Company Size
11-50
Company Industries
Artificial Intelligence / Machine Learning
  • Growing fast
    Showed strong hiring growth in the past month
Learn more about PolarGrid image

Perks

Equity benefits