Data Scientist II
About the Role
The Applied Research team is a group of data scientists and content specialists who are experts in leveraging machine learning, natural language processing, and generative AI models to develop solutions that deliver value to our users and business. We act as a key driver for innovation, whether it’s in product surface experimentation, metadata generation, or model development. Along with Product and Engineering partners, we design solutions and collaborate in cross-functional squads to maximize business impact.
Our areas of impact include content enrichment, representation learning, recommendations, search, translation, and many others, applied to diverse media across text, image, and audio. We operate at a scale of hundreds of millions of documents, millions of users, and billions of user interactions.
We are seeking a curious and collaborative individual with an eye for simplicity, end-to-end visibility, and impact, who is excited about building models using massive amounts of data, leveraging language models, and deploying models.
Responsibilities
- Focus on a variety of content classification use cases, leveraging everything from traditional NLP to sophisticated LLMs and generative models.
- Investigate methods of solving our most challenging problems at Scribd, at scale.
- Collaborate with other Data Scientists, Machine Learning Engineers, and ML Data Engineers on cross-functional projects.
- Leverage any algorithm at your disposal: from classical Scikit-learn and NumPy models to custom Neural Networks in PyTorch to third-party LLM APIs.
- Process massive amounts of data with Python, SQL, and Spark.
- Align with stakeholders through written and verbal communications on the approaches and results of projects, while writing detailed, accurate, and concise project documentation.
Requirements
- 3+ years of post-qualification experience developing machine learning models, working with systems at scale, and deploying to production environments.
- Proficiency in Python.
- Hands-on experience building ML pipelines and working with distributed data processing frameworks like Apache Spark, Databricks, or similar.
- Intermediate level in at least three of these fields: classification algorithms, natural language processing, search, information retrieval, named entity recognition, deep learning, generative models.
- Intermediate level or greater experience with SQL or PySpark.
- Bachelors or Masters in a relevant quantitative discipline including but not limited to Statistics, Computer Science, Data Science, Artificial Intelligence, or another field with a strong quantitative focus.
Pay
At Scribd, your base pay is one part of your total compensation package and is determined within a range. Our pay ranges are based on the local cost of labor benchmarks for each specific role, level, and geographic location.
- California: $118,000 to $184,000
- United States (outside California): $97,000 to $175,000
- Canada: $123,000 CAD to $150,000 CAD
This position is also eligible for a competitive equity ownership and a comprehensive and generous benefits package.
Benefits
- Scribd Flex (flexible work model)
- Comprehensive health, dental, and vision coverage
- Mental health support and disability coverage
- Generous paid time off, including vacation, sick time, holidays, winter break, volunteer time, and sabbaticals
- Paid parental leave and family support benefits
- Retirement matching and employee equity
- Learning and development programs and professional growth opportunities
- Wellness and home office stipends
- Complimentary access to the Scribd suite of products
- Enterprise access to leading AI tools
Schedule
Scribd Flex empowers employees to choose the workstyle and location that support their best performance, while committing to intentional in-person moments that strengthen collaboration and culture. Occasional in-person attendance is required for all Scribd employees, regardless of location.
Employees must have their primary residence in or near one of the following cities (including surrounding metro areas or locations within a typical commuting distance):
- United States: Atlanta, Austin, Boston, Dallas, Denver, Chicago, Houston, Jacksonville, Los Angeles, Miami, New York City, Phoenix, Portland, Sacramento, Salt Lake City, San Diego, San Francisco, Seattle, Washington D.C.
- Canada: Ottawa, Toronto, Vancouver
- Mexico: Mexico City