Machine Learning Engineer, Sr.
Capital One is a leading financial services organization dedicated to transforming banking through innovative technology and responsible AI systems. With a long-standing reputation for leveraging machine learning to enhance customer experiences, Capital One invests heavily in cutting-edge infrastructure and attracts world-class talent to maintain its industry leadership. The company's commitment to responsible AI development ensures that its systems are reliable, scalable, and ethical, ultimately aiming to create a more human-centered banking experience. Initiatives include real-time fraud detection, personalized customer interactions, and intelligent decision-making tools, all designed to simplify banking and foster trust.
About the role
The Sr. Distinguished Machine Learning Engineer role at Capital One offers an exciting opportunity to lead the development of advanced AI and machine learning systems that power personalized customer experiences across multiple channels. This remote-eligible position involves defining and executing the technical strategy for the personalization platform, central to Capital One's efforts in hyper-personalization and real-time decisioning. The engineer will collaborate with cross-functional teams, including product managers, data scientists, and infrastructure engineers, to design scalable, high-performance ML systems.
Responsibilities
- Define and execute the technical strategy and roadmap for the personalization platform supporting real-time, multi-channel customer experiences
- Collaborate with product, data science, cloud infrastructure, and ML platform teams to develop advanced recommendation systems and algorithms
- Design and maintain a flexible rules engine for dynamic user segmentation, targeting, and real-time decisioning
- Develop and manage scalable ML infrastructure and pipelines, including feature extraction, model training, testing, deployment, and inference workflows
- Architect low-latency, event-driven systems that utilize streaming data for real-time personalization and decision-making
- Advance MLOps practices by creating automated deployment workflows, validation systems, and monitoring solutions
- Research and implement state-of-the-art LLM optimization techniques to enhance system scalability, latency, and cost-efficiency
- Leverage open-source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, and PyTorch
- Provide technical leadership, influence architectural standards, mentor engineers, and drive innovation across teams
Qualifications
- Bachelor's degree in Computer Science, Engineering, or a related field
- Minimum of 10 years of experience designing and building data-intensive solutions using distributed computing
- At least 7 years of programming experience in C, C++, Python, or Scala
- At least 4 years of experience managing the full ML development lifecycle in a business-critical environment
- Experience deploying scalable AI solutions on major cloud platforms (AWS, GCP, Azure)
- Proficiency with ML frameworks such as PyTorch and TensorFlow, and orchestration tools like Databricks, Airflow, or Kubeflow
- Strong foundation in engineering, mathematics, and AI optimization techniques
- Deep understanding of cloud-native engineering, containerization (Docker, Kubernetes), and CI/CD pipelines
- Excellent communication skills with the ability to articulate complex technical concepts
- Proven leadership in platform strategy and cross-functional collaboration
Benefits
- Competitive salary packages aligned with experience and location
- Performance-based incentives including cash bonuses and long-term incentives
- Comprehensive health, dental, and vision insurance plans
- Retirement savings options and financial wellness programs
- Flexible work arrangements, including remote work eligibility
- Generous paid time off and holiday policies
- Opportunities for professional development and continuous learning
- Inclusive workplace culture promoting diversity and inclusion