Avatar for SuperQ Quantum Computing
Solving science and industry’s most challenging problems in a ChatGPT like intuitive experience

AI Engineer (GPU & LLM Optimisation) – Canada / UAE / Remote

Posted: today• Recruiter recently active
Hires remotely in
Remote Work Policy

Remote only

Company Location
Visa Sponsorship

Not Available

Preferred Timezones
Mountain Time
RelocationAllowed
Skills
Optimization
AI
GPU Computing
LLMs
Agentic AI

About the job

Location: Canada preferred, also possible in UAE or fully Remote
Experience: 4–6 years
Company: SuperQ Quantum Computing Inc.

About the Role
SuperQ is looking for an AI Engineer specializing in GPU and LLM optimization to maximize the performance, scale, and efficiency of the models powering our multi-agent architecture. You will be the driving force behind accelerating inference, reducing latency, and optimizing hardware utilization as we scale our hybrid compute platform across industries like healthcare, manufacturing, and finance.

Responsibilities

  • Accelerate Inference: Optimize LLM inference and serving pipelines using high-performance frameworks (e.g., vLLM, TensorRT-LLM, TGI, Triton Inference Server).
  • Model Compression: Implement state-of-the-art model compression techniques, including quantization (e.g., GPTQ, AWQ, FP8/INT8), pruning, and knowledge distillation to reduce memory footprints.
  • Hardware Profiling: Profile GPU performance to identify and eliminate bottlenecks, optimizing memory management techniques such as KV caching and PagedAttention.
  • Custom Compute: Develop and optimize custom CUDA or OpenAI Triton kernels for specialized, computationally heavy operations where standard libraries fall short.
  • Cross-functional Deployment: Collaborate with backend and platform engineering teams to deploy highly scalable, low-latency AI endpoints within our Kubernetes-based infrastructure.
  • System Monitoring: Ensure robust monitoring of GPU health, utilization metrics, latency, and throughput in production environments.

Requirements

  • Experience: 2–4 years of experience in ML engineering, AI infrastructure, or high-performance computing (HPC) with a strong focus on Large Language Models.
  • Core Languages: Strong programming skills in Python; proficiency in C++ and/or CUDA is highly preferred.
  • Frameworks: Hands-on experience with LLM serving architectures, distributed computing, and optimization libraries (e.g., DeepSpeed, Hugging Face Accelerate, Ray).
  • Hardware Knowledge: Deep understanding of GPU architectures (NVIDIA), memory hierarchies, and parallel computing paradigms (e.g., NCCL).
  • Infrastructure: Solid understanding of containerization and orchestration (Docker, Kubernetes) tailored for GPU-accelerated workloads.
  • Bonus: Experience with hardware-aware neural architecture search (NAS), ML Ops, or working at the intersection of classical AI hardware and quantum compute interfaces.

Benefits & Work Culture
Benefits:

  • Competitive salary plus bonus based on module delivery and platform impact.
  • Stock options / equity – participate in the growth of a foundational tech platform.
  • Global remote-friendly roles: Canada, UAE or anywhere remote; with occasional team meetups and hack-weeks.
  • Learning & development stipend: conferences, AI/LLM workshops, certifications (e.g., in ML Ops, prompt engineering).
  • “Innovation time” built into schedule: work on passion projects, internal hackathons, share learnings with the team.
  • Cross-domain exposure: work across industries, build modules for different verticals – variety and career growth built-in.

Work Culture:

  • Startup-scale agility within a mission-driven organisation: you’ll have impact and shape the direction of the platform.
  • Multi-disciplinary collaboration: you’ll work with quantum engineers, UI developers, domain experts – bridging tech and business.
  • Strong focus on autonomy, responsibility and ownership: you’ll own modules end-to-end, from conception through production and maintenance.
  • Culture of continuous learning: we encourage curiosity in new models, LLM architectures, AI trends, and provide time to experiment and publish/share.
  • Balanced remote-first mindset: flexible working hours, asynchronous collaboration but also moments of in-person/virtual team building.

Equal Opportunity Statement
SuperQ is dedicated to building a workplace where everyone can thrive. We are an equal opportunity employer and we make employment decisions based on merit, qualifications and business needs-without regard to race, colour, religion, sex, gender identity or expression, sexual orientation, national origin, age, disability, veteran status, or any other characteristic protected by law. We value diversity, equity and inclusion and strive to create a culture of belonging for all employees.

Interested in this role? Click here to apply.

About the company

SuperQ Quantum Computing company logo
Solving science and industry’s most challenging problems in a ChatGPT like intuitive experience11-50 Employees
Learn more about SuperQ Quantum Computing image

Similar Jobs

Coast company logo
Coast
Coast is re-imagining B2B payments, beginning with fleet and fuel
Vise company logo
Vise
Technology-Powered Asset Manager for customized, intelligent investing
Astranis company logo
Astranis
Building next-generation internet satellites to get the world online
Astranis company logo
Astranis
Building next-generation internet satellites to get the world online
Astranis company logo
Astranis
Building next-generation internet satellites to get the world online
Elfen Software company logo
Elfen Software
Make more, spend less through disciplined engineering and technology