Lead Data Scientist-Deep Learning Specialist
Albertsons Companies · Pleasanton, CA · 3 wk ago
EngineeringFull-time
Main Responsibilities
- Design end-to-end deep learning model development, from problem framing and target definition through architecture selection, training strategy, evaluation, and iteration within a Databricks Lakehouse environment
- Architect and build end-to-end deep learning pipelines, from data ingestion and feature engineering to training, deployment, scaling, and monitoring
- Implement distributed training and large-scale data processing using Apache Spark
- Build scalable batch and real-time inference pipelines integrated with Databricks workflows
- Lead fine-tuning and adaptation of large models and foundation models using custom data, with checkpoints, experiments, and model artifacts tracked in MLflow and prepared for governed deployment
- Optimize data pipelines and model performance for scalability, latency, and cost efficiency
- Collaborate with cross-functional teams to productionize ML solutions on the Lakehouse
Required Qualifications
- Proven experience leading deep learning model development for complex business problems, including problem formulation, experimentation, evaluation, and productionization
- Strong hands-on expertise in PyTorch or TensorFlow and modern neural architectures, with experience scaling training using multi-GPU or distributed approaches
- Deep hands-on experience with Databricks, including Delta Lake, Spark, and MLflow, Unity Catalog, governance, and security
- Strong experience with distributed computing and large-scale data processing (Apache Spark)
- Proficiency in Python and ML/data ecosystems (NumPy, Pandas, Scikit-learn, PySpark)
- Strong understanding of feature engineering and data pipeline design in a Lakehouse architecture
- Expertise in distributed training and inference (multi-GPU, multi-node systems)
- Experience designing high-throughput, low-latency inference systems
- Experience building feature stores and reusable ML components within Databricks
Preferred Qualifications
- Experience deploying large-scale deep learning models (e.g., LLMs, recommendation systems) on Databricks
- Experience with cloud platforms (AWS, Azure, GCP) alongside Databricks
- Experience with streaming pipelines (Structured Streaming, Kafka integration)
- Experience with generative AI, LLM fine-tuning, or foundation models
- Background in retail, e-commerce, or supply chain analytics