Avatar for J&M Group
J&M Group
Actively Hiring
IT Staffing

Cloud DevOps Engineer

  • Old Toronto
  • |Contract
Posted: 3 days ago• Recruiter recently active
Job Location
Old Toronto
Remote Work Policy

In office

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Docker
Microsoft Azure
Kibana
Kubernetes
Terraform
Grafana
ARM Templates
Azure Monitor
Bicep
Application Insights

About the job

We are hiring for a Senior DevOps / Site Reliability Engineer AI Platform. The role focuses on building, operating, monitoring, and scaling an enterprise AI platform across Development, QA, and Production environments.

A strong fit will have advanced Azure and Kubernetes experience along with strong expertise in observability, dashboards, monitoring, alerting, and automated scaling.

What you bring

  • Strong professional experience in DevOps, Site Reliability Engineering, cloud infrastructure, or platform engineering.
  • Advanced hands-on experience with Microsoft Azure.
  • Strong experience deploying and operating Kubernetes environments.
  • Strong knowledge of Docker and containerization technologies.
  • Experience supporting containerized applications across Development, QA, and Production environments.
  • Advanced experience designing and building dashboards using Grafana, Kibana, Azure Monitor, Application Insights, or comparable tools.
  • Strong experience with monitoring, observability, logging, alerting, and operational health checks.
  • Experience with application load, infrastructure capacity, performance, and automated scaling.
  • Experience with CI/CD pipelines and automated application deployment.
  • Experience with Infrastructure as Code tools such as Terraform, Bicep, or ARM templates.
  • Strong troubleshooting skills across applications, containers, infrastructure, networking, and cloud services.

What you'll do

  • Design, deploy, configure, and maintain infrastructure within Microsoft Azure.
  • Deploy and operate containerized applications using Kubernetes.
  • Monitor container and cluster health, resource consumption, capacity, and performance.
  • Configure scaling policies and develop intelligent scaling approaches based on workload and resource utilization.
  • Design and build operational dashboards covering platform health, performance, capacity, errors, latency, and container health.
  • Implement monitoring and alerting across infrastructure, applications, containers, integrations, and AI platform services.
  • Establish actionable alerts, health checks, anomaly detection, and automated remediation where appropriate.
  • Support production platform stability, availability, and operational readiness.
  • Investigate platform, deployment, infrastructure, monitoring, and performance issues.
  • Participate in root-cause analysis and implement preventative improvements.
  • Create operational runbooks and troubleshooting guidance.

Nice to have

  • Experience supporting AI, machine learning, data, or high-compute platforms.
  • Experience monitoring AI models, inference services, token usage, GPU workloads, API consumption, queues, or model performance.
  • Experience implementing automated remediation, predictive monitoring, or AI-assisted platform operations.
  • Familiarity with AWS services and cloud operations.
  • Experience with Elasticsearch, Log Analytics, OpenTelemetry, Prometheus, or similar observability technologies.
  • Experience defining service-level indicators, service-level objectives, and reliability standards.
  • Experience with security, identity, secrets management, and cloud governance within Azure.

About the company

J&M Group company logo

J&M Group

Actively Hiring
IT Staffing51-200 Employees
Learn more about J&M Group image