Senior Engineer – ML
Albertsons Companies · Phoenix, AZ · 1 wk ago
EngineeringFull-time
Main Responsibilities
- Design, develop, and productionize AI/ML capabilities for the AIOps platform to support intelligent observability and operational decision-making.
- Build machine learning systems for time-series forecasting, anomaly prediction, incident prediction, alert classification, noise reduction, and event correlation.
- Create causal ML and statistical inference solutions to identify likely root causes, dependency impacts, and relationships across systems and services.
- Develop feature pipelines and model training workflows using telemetry, log, metric, trace, topology, and incident data.
- Create intelligent RCA summarization capabilities using LLMs and agentic frameworks such as LangChain and LangGraph.
- Develop AI agents and multi-agent systems for use cases such as Multi-Agent RCA, SRE Assistant, remediation guidance, incident triage, and operational knowledge retrieval.
- Design prompt orchestration, reasoning workflows, retrieval pipelines, tool usage patterns, and memory/context handling for AI agents.
- Integrate AI/ML services with observability platforms, event systems, knowledge bases, CMDB, incident management tools, and automation platforms.
- Collaborate with platform and engineering teams to build scalable model-serving and agent-serving architectures.
- Define and implement evaluation frameworks for model quality, agent effectiveness, hallucination reduction, relevance, and operational usefulness.
- Ensure AI/ML systems are scalable, reliable, explainable, and aligned with enterprise security, governance, and responsible AI practices.
- Build and maintain APIs and microservices for model inference, online scoring, batch predictions, and agent orchestration.
- Partner with SREs, observability engineers, and product stakeholders to translate operational pain points into ML and AI-driven solutions.
- Continuously improve model performance, feature quality, inference latency, agent reliability, and business impact through experimentation and monitoring.
- Create technical documentation covering model design, feature logic, training pipelines, evaluation metrics, deployment architecture, and agent workflows.
- Drive innovation in AI-enabled observability, causal intelligence, and agentic SRE workflows to enhance the value of the AIOps platform.
Requirements
- Strong experience designing and building AI/ML systems for real-world production use cases.
- Solid hands-on experience with Python and common ML frameworks and libraries such as scikit-learn, XGBoost, PyTorch, TensorFlow, Pandas, and NumPy.
- Experience building machine learning solutions for forecasting, anomaly detection, prediction, classification, clustering, ranking, or recommendation problems.
- Strong understanding of time-series modeling techniques for forecasting and operational prediction use cases.
- Experience with alert classification, incident prediction, event deduplication, prioritization, or signal correlation in observability or IT operations contexts.
- Knowledge of causal ML, causal inference, graph-based reasoning, and dependency-aware analysis techniques for RCA and impact analysis.
- Hands-on experience building AI applications using LLM frameworks such as LangChain and LangGraph.
- Experience designing AI agents or multi-agent systems for reasoning, summarization, task orchestration, troubleshooting, or assistant workflows.
- Strong understanding of prompt engineering, RAG architecture, embeddings, vector stores, tool calling, memory handling, and agent evaluation techniques.
- Experience integrating LLM systems with enterprise tools, APIs, knowledge repositories, and operational systems.
- Experience building backend services and APIs for AI/ML model inference and agent orchestration.
- Good understanding of observability data such as logs, metrics, traces, topology, incidents, and alerts.
- Experience with data engineering concepts including feature engineering, data preprocessing, model pipelines, and batch or streaming inference.
- Familiarity with graph databases such as Neo4j and their use in dependency mapping, causal analysis, and knowledge-driven AI systems.
- Experience with REST APIs, microservices architecture, Docker, Kubernetes, and cloud-native deployment patterns.
- Familiarity with CI/CD, MLOps, model lifecycle management, experiment tracking, and model versioning practices.
- Knowledge of OpenTelemetry, monitoring systems, and observability platforms is highly desirable.
- Strong understanding of software engineering fundamentals, system design, and scalable architecture patterns.
- Strong analytical and problem-solving skills, with the ability to convert ambiguous operational problems into measurable AI/ML solutions.
- Excellent communication and collaboration skills to work with SREs, platform engineers, product owners, and business stakeholders.
- Self-driven mindset with strong curiosity, innovation, and the ability to learn and apply emerging AI techniques effectively.
Qualifications
- Bachelor’s degree in computer science, Information Systems, Engineering, Data Science, Artificial Intelligence, or a related field, or equivalent practical experience.
- 6 to 10 plus years of overall experience in software engineering, machine learning, or AI system development.
- 3 plus years of hands-on experience building and deploying machine learning systems in production.
- Strong experience in Python-based AI/ML development is required.
- Experience working on observability, monitoring, or AIOps-related platforms is strongly preferred.
- Experience building LLM-powered applications, AI agents, or multi-agent workflows for enterprise use cases is highly preferred.
- Experience in AIOps, Observability, SRE, IT operations, or incident management domains.
- Experience applying AI/ML to RCA, anomaly explanation, incident summarization, service health prediction, or remediation recommendations.
- Familiarity with knowledge graphs and graph-based ML techniques for dependency-aware intelligence.
- Experience using vector databases and retrieval frameworks for enterprise search and agentic applications.
- Experience integrating AI services with tools such as ServiceNow, Grafana, Prometheus, Splunk, AppDynamics, or similar platforms.
- Familiarity with MCP-based client or agent integrations is a plus.