Data Engineer
About the role
As a Senior Data Scientist, you will partner with a diverse team of Engineers, Economists, Computer Scientists, Mathematicians, Physicists, Statisticians, and Actuaries tasked with mining industry-leading internal data to develop Machine Learning, AI, and agentic capabilities for our businesses. The role requires a rare combination of sophisticated analytical expertise, business acumen, strategic mindset, client relationship skills, problem-solving, and a passion for generating business impact. This is an exciting opportunity to be part of a strategic initiative that is evolving and growing over time.
You will bring excellent problem-solving, communication, and teamwork skills, along with agile ways of working, strong business insight, and a continuous learning focus.
Responsibilities
- Hands-on development of advanced Data Science, Machine Learning, and Agentic AI solutions as part of a portfolio led by the Lead Data Scientist.
- Perform data analysis, model development, training, testing, and deployment of Agentic AI systems.
- Write production-grade code and partner with machine learning engineers to deploy models into production, including traditional machine learning, statistical models, GenAI, and agentic solutions.
- Engineer Agentic AI systems, apply fine-tuning techniques (e.g., LoRA), deploy LLMs, and work with RAG, Agentic RAG, Strands, Claude Agents SDK, and related concepts.
- Partner with machine learning engineers to productionize models and agents, and with data engineers to build data pipelines.
- Collaborate with software engineers to integrate solutions with business platforms.
- Continuously research new methods, algorithms, modeling techniques, and data analytics approaches.
- Utilize cloud-based AI platforms like Bedrock and SageMaker AI.
Requirements
- Advanced degree (Master’s, Ph.D.) in Mathematics, Statistics, Engineering, Econometrics, Physics, Computer Science, Actuarial Science, Data Science, or comparable quantitative disciplines.
- Experience working on complex problems requiring in-depth analysis of situations or data.
- Ability to exercise judgment within broadly defined practices and policies to select methods, techniques, and evaluation criteria for results.
- Strong self-initiative to learn new skills and knowledge continuously.
- Excellent problem-solving, communication, and collaboration skills.
Skills
- Data Acquisition and Transformation: Acquiring data from disparate sources using APIs, semantic data models, and SQL; transforming data using SQL and Python; visualizing data using tools like Python.
- Database Management System: Knowledge of database structures, cloud/AWS environments, primary/foreign key relationships, table design, SQL (relational), NoSQL, Graph/ontology (Graph DB), and semantic data models.
- Data Analysis and Insights: Analyzing structured and unstructured data using visualization, manipulation, and statistical methods to identify patterns, anomalies, relationships, and trends.
- Statistics and Computing: Exceptional understanding of multivariable calculus, linear algebra, differential equations, applied probability, applied statistics, computer science (programming methodologies), and cloud computing. Knowledge of statistical techniques such as descriptive/inferential/Bayesian statistics, time series analysis, and experimentation.
- Machine Learning: Understanding of machine learning theory, including the mathematics underlying algorithms. Expertise in building, training, testing, interpreting, and monitoring supervised (regression/classification) and unsupervised (clustering/segmentation/anomaly detection) models.
- Generative AI & Natural Language Processing: Experience with text analysis, NLP, LLMs, and Generative AI. Familiarity with modern Gen AI technologies like RAG, LangChain, LangGraph, vector DBs, LangFuse, AgentCore, and Agents, particularly in the Retirement Strategies area.
- Model Deployment: Understanding of the model development lifecycle, A/B testing, CI/CD pipelines, and frameworks like AWS SageMaker and newer AWS/Azure Agentic AI infrastructure products.
- Data Wrangling: Preparing data for analysis, redefining and mapping raw data, and processing large datasets (structured and unstructured).
- DevSecOps: Knowledge of the project development lifecycle in an AWS environment, including development, QA, staging, and production deployment stages.
- Programming Languages: Proficiency in Python and SQL.
Schedule
Hybrid role based in Newark, NJ.