Avatar for Albertsons
Albertsons
Actively Hiring
Food and drug retailer with 2,200+ stores

Senior Engineer – ML

  • $109k – $156k
  • |
  • |6 years of exp
  • |Full Time
Posted: 2 weeks ago• Recruiter recently active
Job Location
Remote Work Policy

In office

Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
Metrics
Clustering
Classification
System Design
Numpy
REST APIs
Pandas
Anomaly Detection
Forecasting
scikit-learn
Prediction
Docker
Multi-agent Systems
Topology
Kubernetes
Feature Engineering
Prioritization
Causal inference
TensorFlow
Alerts
Xgboost
Microservices Architecture
PyTorch
APIs
Data Preprocessing
Recommendation
Ranking
Prompt Engineering
Causal ML
LangChain
AI Agents
Vector Stores
Embeddings
Langgraph
Logs
Traces
Time-Series Modeling
Software Engineering Fundamentals
Backend Services
Streaming Inference
RAG Architecture
Tool Calling
Scalable Architecture Patterns
Incidents
Event Deduplication
Memory Handling
Agent Evaluation
Graph-Based Reasoning
Model Pipelines
Batch Inference
Signal Correlation
Incident Prediction
Alert Classification
Dependency-Aware Analysis

About the job

Why choose us?

Are you ready to take the next step in your career? Join us for an exciting opportunity at Albertsons Companies, where innovation and customer service go hand-in-hand!

At Albertsons Companies, we are looking for someone who’s not just seeking a job, but someone who wants to make an impact. In this role, you’ll have the opportunity to lead, innovate, and contribute to the growth of a company that values great service and lasting customer relationships. This position offers the chance to work in a fast-paced, dynamic environment that’s constantly evolving.

This role is an individual contributor position responsible for designing, developing, fine-tuning, and operationalizing AI/ML capabilities for the AIOps platform within the Observability product. The candidate will work closely with the Lead Engineer, SRE teams, Platform team, Data Ingestion team, Platform DevOps team, Visualization team, and other portfolio teams.

As part of the AIOps Platform team, you will design and build intelligent systems that improve observability, incident response, and operational efficiency. This includes developing machine learning models for forecasting, anomaly prediction, alert classification, event intelligence, and causal analysis, as well as building AI agents and multi-agent workflows for RCA summarization, investigative assistance, and SRE productivity use cases.

This position will be based out of Phoenix, Arizona or Pleasanton CA.

Main Responsibilities

  • Design, develop, and productionize AI/ML capabilities for the AIOps platform to support intelligent observability and operational decision-making.
  • Build machine learning systems for time-series forecasting, anomaly prediction, incident prediction, alert classification, noise reduction, and event correlation.
  • Develop causal ML and statistical inference solutions to identify likely root causes, dependency impacts, and relationships across systems and services.
  • Create and fine-tune models for incident intelligence use cases such as forecasting service degradation, capacity risk prediction, alert prioritization, and anomaly explanation.
  • Design feature pipelines and model training workflows using telemetry, log, metric, trace, topology, and incident data.
  • Build intelligent RCA summarization capabilities using LLMs and agentic frameworks such as LangChain and LangGraph.
  • Develop AI agents and multi-agent systems for use cases such as Multi-Agent RCA, SRE Assistant, remediation guidance, incident triage, and operational knowledge retrieval.
  • Design prompt orchestration, reasoning workflows, retrieval pipelines, tool usage patterns, and memory/context handling for AI agents.
  • Integrate AI/ML services with observability platforms, event systems, knowledge bases, CMDB, incident management tools, and automation platforms.
  • Collaborate with platform and engineering teams to build scalable model-serving and agent-serving architectures.
  • Define and implement evaluation frameworks for model quality, agent effectiveness, hallucination reduction, relevance, and operational usefulness.
  • Ensure AI/ML systems are scalable, reliable, explainable, and aligned with enterprise security, governance, and responsible AI practices.
  • Build and maintain APIs and microservices for model inference, online scoring, batch predictions, and agent orchestration.
  • Partner with SREs, observability engineers, and product stakeholders to translate operational pain points into ML and AI-driven solutions.
  • Continuously improve model performance, feature quality, inference latency, agent reliability, and business impact through experimentation and monitoring.
  • Establish engineering best practices for ML development, prompt engineering, evaluation, model deployment, testing, versioning, and documentation.
  • Support production incident analysis for AI/ML services and drive root cause identification and remediation for model or agent failures.
  • Create technical documentation covering model design, feature logic, training pipelines, evaluation metrics, deployment architecture, and agent workflows.
  • Drive innovation in AI-enabled observability, causal intelligence, and agentic SRE workflows to enhance the value of the AIOps platform.

We Are Searching For Someone With The Following Skills

  • Strong experience designing and building AI/ML systems for real-world production use cases.
  • Solid hands-on experience with Python and common ML frameworks and libraries such as scikit-learn, XGBoost, PyTorch, TensorFlow, Pandas, and NumPy.
  • Experience building machine learning solutions for forecasting, anomaly detection, prediction, classification, clustering, ranking, or recommendation problems.
  • Strong understanding of time-series modeling techniques for forecasting and operational prediction use cases.Experience with alert classification, incident prediction, event deduplication, prioritization, or
  • signal correlation in observability or IT operations contexts.
  • Knowledge of causal ML, causal inference, graph-based reasoning, and dependency-aware analysis techniques for RCA and impact analysis.
  • Hands-on experience building AI applications using LLM frameworks such as LangChain and LangGraph.
  • Experience designing AI agents or multi-agent systems for reasoning, summarization, task orchestration, troubleshooting, or assistant workflows.
  • Strong understanding of prompt engineering, RAG architecture, embeddings, vector stores, tool calling, memory handling, and agent evaluation techniques.
  • Experience integrating LLM systems with enterprise tools, APIs, knowledge repositories, and operational systems.
  • Experience building backend services and APIs for AI/ML model inference and agent orchestration.
  • Good understanding of observability data such as logs, metrics, traces, topology, incidents, and alerts.
  • Experience with data engineering concepts including feature engineering, data preprocessing, model pipelines, and batch or streaming inference.
  • Familiarity with graph databases such as Neo4j and their use in dependency mapping, causal analysis, and knowledge-driven AI systems.
  • Experience with REST APIs, microservices architecture, Docker, Kubernetes, and cloud-native deployment patterns.
  • Familiarity with CI/CD, MLOps, model lifecycle management, experiment tracking, and model versioning practices.
  • Knowledge of OpenTelemetry, monitoring systems, and observability platforms is highly desirable.
  • Strong understanding of software engineering fundamentals, system design, and scalable architecture patterns.
  • Strong analytical and problem-solving skills, with the ability to convert ambiguous operational problems into measurable AI/ML solutions.
  • Excellent communication and collaboration skills to work with SREs, platform engineers, product owners, and business stakeholders.
  • Self-driven mindset with strong curiosity, innovation, and the ability to learn and apply emerging AI techniques effectively.

We believe the successful candidate has these qualifications and experience:

  • Bachelor’s degree in computer science, Information Systems, Engineering, Data Science, Artificial Intelligence, or a related field, or equivalent practical experience.
  • 6 to 10 plus years of overall experience in software engineering, machine learning, or AI system development.
  • 3 plus years of hands-on experience building and deploying machine learning systems in production.
  • Strong experience in Python-based AI/ML development is required.
  • Experience working on observability, monitoring, or AIOps-related platforms is strongly preferred.
  • Experience building LLM-powered applications, AI agents, or multi-agent workflows for enterprise use cases is highly preferred.
  • Experience in AIOps, Observability, SRE, IT operations, or incident management domains.
  • Experience applying AI/ML to RCA, anomaly explanation, incident summarization, service health prediction, or remediation recommendations.
  • Familiarity with knowledge graphs and graph-based ML techniques for dependency-aware intelligence.
  • Experience using vector databases and retrieval frameworks for enterprise search and agentic applications.
  • Experience integrating AI services with tools such as ServiceNow, Grafana, Prometheus, Splunk, AppDynamics, or similar platforms.
  • Familiarity with MCP-based client or agent integrations is a plus.

We Also Provide a Variety Of Benefits Including

  • Competitive wages paid weekly
  • Access to up to 50% of your earned wages before payday, via our partnership with Stream
  • Associate discounts
  • Health and financial well-being benefits for eligible associates (Medical, Dental, 401k and more!)
  • Time off (vacation, holidays, sick pay). For eligibility requirements please visit myACI Benefits
  • Leaders invested in your training, career growth and development
  • An inclusive work environment with talented colleagues who reflect the communities we serve

Our Values – Click below to view video: ACI Values

A copy of the full job description can be made available to you.

About the company

Albertsons company logo

Albertsons

Actively Hiring
Food and drug retailer with 2,200+ stores5000+ Employees
Learn more about Albertsons image

Funding

AMOUNT RAISED
$24.6B
FUNDED OVER
1 round
Round
S
$24600000000
Seed - Oct 2022

Similar Jobs

IonQ company logo
IonQ
Building the world’s best quantum computers to solve the world’s most ecomplex problems
IonQ company logo
IonQ
Building the world’s best quantum computers to solve the world’s most ecomplex problems
IonQ company logo
IonQ
Building the world’s best quantum computers to solve the world’s most ecomplex problems