Avatar for Motion Industries
Motion Industries
Actively Hiring

Site Reliability Engineer - ED&A - Global Industrial

Posted: 7 days ago• Recruiter recently active
Job Location
Remote Work Policy

In office - WFH flexibility

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Software Engineering
Python
Business Intelligence
Reporting
Databases
Integration
Customer Experience
Infrastructure
Authentication
Automation
Cloud Services
Operating Systems
Configuration Management
Governance
Security
Systems Engineering
Identity Management
Compliance
Reliability
Google Bigquery
Incident Response
Enterprise Integration Patterns
PowerShell
Scripting
Data Governance
cloud
Data Engineering
Cross-functional Collaboration
Data protection
Troubleshooting
Production support
Supportability
Data
Reliability Engineering
DataDog
Terraform
Dynatrace
authorization
Grafana
Google Cloud Storage
Recovery
Power BI
Infrastructure as code
Production Readiness
APIs
Azure DevOps
Observability
Power Automate
Power Platform
Power Apps
Access Management
Application
Root-Cause Analysis
Azure Monitor
Microsoft Fabric
Google Cloud IAM
Runbooks
Operational Maturity
Backup
Technical Guidance
Error Budgets
Google Cloud Monitoring
CI/CD Tooling
SRE Principles
Operational Dashboards
Microsoft Azure Networking
Microsoft Azure Monitor
Google Cloud Logging
Microsoft Connectors
Platform Standards
Microsoft Workspaces
Reliability Scorecards
Data Platform Operations
Change-Management Practices
Service-Level Objectives
Microsoft Platform Administration
Platform Performance Monitoring
Continuous Improvement Plans
Service-Level Indicators
Alerting Standards
Enterprise Data and Analytics Platforms
Production Support Documentation
Platform Reliability Monitoring
Platform Availability Monitoring
Network Dependencies
Post-Incident Improvement
Business Productivity Workloads
Google Cloud Networking Concepts
Google Cloud Workload Performance Troubleshooting
Google Cloud Job Performance Troubleshooting
Microsoft Azure Log Analytics
Microsoft Azure Identity and Access Management
Microsoft Azure Resource Management
Microsoft Azure Cloud-Native Administration
Microsoft Gateways
Microsoft Data Pipelines
Microsoft Semantic Models
Platform Capacity Monitoring
Jobs Monitoring
Queries Monitoring
Pipelines Monitoring
APIs Monitoring
Automation Flows Monitoring
Dependent Services Monitoring
Business Technology Teams
Standards Adoption

About the job

Site Reliability Engineer - ED&A

The Site Reliability Engineer – EDA is a technical contributor responsible for Site Reliability Engineering practices supporting the Enterprise Data & Analytics platform. This role performs reliability and operational engineering for data and analytics platforms, integrations, pipelines, and related services, including platforms such as Google BigQuery, Microsoft Fabric, Power BI, Power Platform, and associated Azure and Google Cloud services. The Site Reliability Engineer establishes service-level indicators, service-level objectives, error budgets, monitoring, alerting, dashboards, and runbooks while driving incident response, root-cause analysis, automation, capacity planning, performance tuning, resilience, and production readiness. This role provides technical guidance, engineering standards, and operational best practices for engineers supporting EDA services, and partners closely with data engineering, application, cloud, security, and governance teams to improve reliability, supportability, and business outcomes.

JOB DUTIES

  • Performs reliability and operational engineering for the Enterprise Data & Analytics platform, including Google BigQuery, Microsoft Fabric, Power BI, Power Platform, associated data pipelines, integrations, automation, reporting services, and dependent cloud services.
  • Establishes and governs service-level indicators, service-level objectives, error budgets, availability targets, performance baselines, monitoring standards, alerting practices, dashboards, runbooks, and operational health metrics for EDA services.
  • Analyzes telemetry from monitoring, logging, tracing, platform administration, pipeline execution, query performance, capacity, consumption, and cost-management tools to identify reliability, performance, security, scalability, and efficiency improvements.
  • Partners with data engineering, application, cloud, security, governance, analytics, and infrastructure teams to improve platform design, integration patterns, deployment practices, release readiness, supportability, resilience, and production operations.
  • Drives automation for provisioning, deployment, remediation, monitoring configuration, environment validation, job and pipeline health checks, alert enrichment, access reviews, incident response workflows, and operational reporting.
  • Guides capacity planning, performance tuning, resilience engineering, disaster recovery planning, backup and restore validation, service continuity planning, architecture reviews, and production readiness assessments for EDA services.
  • Troubleshoots and resolves incidents involving data and analytics platforms, workloads, integrations, pipelines, APIs, automation flows, connectors, workspaces, gateways, permissions, queries, semantic models, and platform dependencies.
  • Drives incident response, root-cause analysis, post-incident reviews, corrective actions, and reliability improvement plans to reduce recurrence, improve operational maturity, and strengthen customer experience.
  • Provides technical guidance, engineering standards, implementation patterns, peer support, operational reviews, documentation practices, and reliability expectations for engineers supporting the EDA platform.
  • Promotes secure, compliant, cost-effective, and well-governed data and analytics operations by supporting access controls, data protection practices, platform governance, resource utilization reviews, lifecycle management, and operational reporting.

EDUCATION & EXPERIENCE

Typically requires a bachelor's degree and five (5) to seven (7) years of experience in a technology and/or software engineering role or an equivalent combination.

KNOWLEDGE, SKILLS, ABILITIES

  • Advanced understanding of SRE principles, including reliability engineering, observability, automation, incident response, root-cause analysis, post-incident improvement, service-level indicators, service-level objectives, error budgets, and production readiness.
  • Experience performing reliability, operations, or engineering support for enterprise data and analytics platforms used for reporting, data engineering, integration, automation, business intelligence, and business productivity workloads.
  • Experience with Google Cloud data services such as BigQuery, Cloud Storage, Cloud Logging, Cloud Monitoring, IAM, networking concepts, and workload or job performance troubleshooting.
  • Experience with Microsoft Azure services and operational capabilities, including Azure Monitor, Log Analytics, Azure networking, identity and access management, resource management, and cloud-native administration.
  • Experience with Microsoft Fabric, Power BI, Power Platform, Power Automate, Power Apps, gateways, connectors, workspaces, data pipelines, semantic models, and platform administration concepts.
  • Ability to monitor, troubleshoot, and tune platform performance, capacity, reliability, availability, jobs, queries, pipelines, APIs, integrations, automation flows, and dependent services.
  • Knowledge of identity, access, security, data governance, compliance, data protection, backup, recovery, and change-management practices for enterprise data and analytics platforms.
  • Experience creating and governing operational dashboards, alerting standards, runbooks, production support documentation, platform standards, reliability scorecards, and continuous improvement plans.
  • Experience with automation, scripting, infrastructure as code, configuration management, and CI/CD tooling such as Azure DevOps, Terraform, PowerShell, Python, or similar tools.
  • Experience with observability and monitoring platforms such as Azure Monitor, Google Cloud Monitoring, Grafana, Datadog, Dynatrace, or similar tools.
  • Ability to provide technical guidance across infrastructure, security, data, application, cloud, governance, and business technology teams to improve reliability, supportability, standards adoption, operational maturity, and customer experience.
  • Strong troubleshooting skills across cloud services, network dependencies, APIs, databases, operating systems, authentication, authorization, and enterprise integration patterns.
  • A strong mix of software engineering, systems engineering, data platform operations, automation, production support, technical guidance, and cross-functional collaboration skills.

PHYSICAL DEMANDS:

LICENSES & CERTIFICATIONS:

SUPERVISORY RESPONSIBILITY:

BUDGET RESPONSIBILITY: No

COMPANY INFORMATION: Motion offers an excellent benefits package which includes options for healthcare coverage, 401(k), tuition reimbursement, vacation, sick, and holiday pay.

DISCLAIMER: This job description illustrates the general nature and level of work performed by employees within this job classification. It is not intended to contain or be interpreted as a comprehensive inventory of all duties, responsibilities and skills required. Management retains the right to add or modify duties at any time.

Not the right fit? Let us know you're interested in a future opportunity by joining our Talent Community on jobs.genpt.com or create an account to set up email alerts as new job postings become available that meet your interest!

GPC conducts its business without regard to sex, race, creed, color, religion, marital status, national origin, citizenship status, age, pregnancy, sexual orientation, gender identity or expression, genetic information, disability, military status, status as a veteran, or any other protected characteristic. GPC's policy is to recruit, hire, train, promote, assign, transfer and terminate employees based on their own ability, achievement, experience and conduct and other legitimate business reasons.

Similar Jobs

Voreas Laboratories company logo
Voreas Laboratories
Cyberattack attribution for high-value organizations
Braze company logo
Braze
Customer Engagement Platform