Jobs · Engineering · Arizona

Senior Engineer – ML

Albertsons Companies · Phoenix, AZ · 1 wk ago
EngineeringFull-time

Main Responsibilities

  • Design, develop, and productionize AI/ML capabilities for the AIOps platform to support intelligent observability and operational decision-making.
  • Build machine learning systems for time-series forecasting, anomaly prediction, incident prediction, alert classification, noise reduction, and event correlation.
  • Create causal ML and statistical inference solutions to identify likely root causes, dependency impacts, and relationships across systems and services.
  • Develop feature pipelines and model training workflows using telemetry, log, metric, trace, topology, and incident data.
  • Create intelligent RCA summarization capabilities using LLMs and agentic frameworks such as LangChain and LangGraph.
  • Develop AI agents and multi-agent systems for use cases such as Multi-Agent RCA, SRE Assistant, remediation guidance, incident triage, and operational knowledge retrieval.
  • Design prompt orchestration, reasoning workflows, retrieval pipelines, tool usage patterns, and memory/context handling for AI agents.
  • Integrate AI/ML services with observability platforms, event systems, knowledge bases, CMDB, incident management tools, and automation platforms.
  • Collaborate with platform and engineering teams to build scalable model-serving and agent-serving architectures.
  • Define and implement evaluation frameworks for model quality, agent effectiveness, hallucination reduction, relevance, and operational usefulness.
  • Ensure AI/ML systems are scalable, reliable, explainable, and aligned with enterprise security, governance, and responsible AI practices.
  • Build and maintain APIs and microservices for model inference, online scoring, batch predictions, and agent orchestration.
  • Partner with SREs, observability engineers, and product stakeholders to translate operational pain points into ML and AI-driven solutions.
  • Continuously improve model performance, feature quality, inference latency, agent reliability, and business impact through experimentation and monitoring.
  • Create technical documentation covering model design, feature logic, training pipelines, evaluation metrics, deployment architecture, and agent workflows.
  • Drive innovation in AI-enabled observability, causal intelligence, and agentic SRE workflows to enhance the value of the AIOps platform.

Requirements

  • Strong experience designing and building AI/ML systems for real-world production use cases.
  • Solid hands-on experience with Python and common ML frameworks and libraries such as scikit-learn, XGBoost, PyTorch, TensorFlow, Pandas, and NumPy.
  • Experience building machine learning solutions for forecasting, anomaly detection, prediction, classification, clustering, ranking, or recommendation problems.
  • Strong understanding of time-series modeling techniques for forecasting and operational prediction use cases.
  • Experience with alert classification, incident prediction, event deduplication, prioritization, or signal correlation in observability or IT operations contexts.
  • Knowledge of causal ML, causal inference, graph-based reasoning, and dependency-aware analysis techniques for RCA and impact analysis.
  • Hands-on experience building AI applications using LLM frameworks such as LangChain and LangGraph.
  • Experience designing AI agents or multi-agent systems for reasoning, summarization, task orchestration, troubleshooting, or assistant workflows.
  • Strong understanding of prompt engineering, RAG architecture, embeddings, vector stores, tool calling, memory handling, and agent evaluation techniques.
  • Experience integrating LLM systems with enterprise tools, APIs, knowledge repositories, and operational systems.
  • Experience building backend services and APIs for AI/ML model inference and agent orchestration.
  • Good understanding of observability data such as logs, metrics, traces, topology, incidents, and alerts.
  • Experience with data engineering concepts including feature engineering, data preprocessing, model pipelines, and batch or streaming inference.
  • Familiarity with graph databases such as Neo4j and their use in dependency mapping, causal analysis, and knowledge-driven AI systems.
  • Experience with REST APIs, microservices architecture, Docker, Kubernetes, and cloud-native deployment patterns.
  • Familiarity with CI/CD, MLOps, model lifecycle management, experiment tracking, and model versioning practices.
  • Knowledge of OpenTelemetry, monitoring systems, and observability platforms is highly desirable.
  • Strong understanding of software engineering fundamentals, system design, and scalable architecture patterns.
  • Strong analytical and problem-solving skills, with the ability to convert ambiguous operational problems into measurable AI/ML solutions.
  • Excellent communication and collaboration skills to work with SREs, platform engineers, product owners, and business stakeholders.
  • Self-driven mindset with strong curiosity, innovation, and the ability to learn and apply emerging AI techniques effectively.

Qualifications

  • Bachelor’s degree in computer science, Information Systems, Engineering, Data Science, Artificial Intelligence, or a related field, or equivalent practical experience.
  • 6 to 10 plus years of overall experience in software engineering, machine learning, or AI system development.
  • 3 plus years of hands-on experience building and deploying machine learning systems in production.
  • Strong experience in Python-based AI/ML development is required.
  • Experience working on observability, monitoring, or AIOps-related platforms is strongly preferred.
  • Experience building LLM-powered applications, AI agents, or multi-agent workflows for enterprise use cases is highly preferred.
  • Experience in AIOps, Observability, SRE, IT operations, or incident management domains.
  • Experience applying AI/ML to RCA, anomaly explanation, incident summarization, service health prediction, or remediation recommendations.
  • Familiarity with knowledge graphs and graph-based ML techniques for dependency-aware intelligence.
  • Experience using vector databases and retrieval frameworks for enterprise search and agentic applications.
  • Experience integrating AI services with tools such as ServiceNow, Grafana, Prometheus, Splunk, AppDynamics, or similar platforms.
  • Familiarity with MCP-based client or agent integrations is a plus.

Similar jobs

Senior Engineer

City of Santa Fe SpringsProsper, TX· 2 mo ago
Engineering$98k–$127k/yrapply on governmentjobs.com

Senior Engineer

City of PascoPasco, WA· 2 mo ago
Engineering$120k–$144k/yrapply on governmentjobs.com

Senior Engineer

City of Las CrucesLas Cruces, NM· 2 mo ago
Engineeringapply on governmentjobs.com

Senior Engineer

City of Ormond BeachOrmond Beach, FL· 2 mo ago
Engineeringapply on governmentjobs.com

Senior Engineer

MANTECHCrane, IN· 1 mo ago
Engineeringapply on mantech.avature.net