Machine Learning Operations Engineer II
Mission
The MLOps team is the de facto ML platform team at Kensho. Our team’s mission is critical: empower our ML engineers with state-of-the-art processes, tooling, and infrastructure to iterate quickly, build reliably, and identify potential production issues early. We sit at the intersection of infrastructure and ML, and work closely with all our ML teams (ML Product teams, R&D, ...) and our infrastructure teams (Core Infra, SRE, Security).
What You’ll Do
- Iterate on Kensho’s ML processes to develop tools, services, and frameworks that make every stage of the ML workflow robust, auditable, and usable.
- Work closely with ML engineers to understand their unique processes, identify pain points, and form effective solutions.
- Empower engineers with the stable tooling necessary to rapidly experiment and actualize their research into demonstrable prototypes and mature products.
- Provide resources and training for ML teams on best practices, enabling them to efficiently productionize their work to be leveraged by high-value products and services.
- Evaluate, select and champion open source and third-party solutions, driving their adoption across teams and integrating into Kensho’s existing platform ecosystem.
- Ship scalable, efficient, and automated processes for model fine-tuning and reinforcement learning and for the evaluation of LLMs/Agents.
- Improve LLM and Agentic observability to help monitor agentic applications in production, detecting performance, decay and drift issues.
- Stay at the frontier by actively tracking emerging tools and frameworks, promote best practices and strengthen the technical expertise of the team with your unique skill set.
What You’ll Need
- 2+ years of experience in ML infra, ML Ops, ML Engineering or some similar skillset.
- Experience managing distributed systems with Kubernetes.
- Cloud Platform (AWS) understanding. We utilize tools like EKS and managed ML services like Bedrock and SageMaker.
- Python proficiency (we are a python shop mostly).
- Familiarity with distributed computing frameworks and workflow orchestration (ie. Ray, Airflow).
- Familiarity with software engineering best practices in an ML context.
- Some basic understanding of ML concepts, LLMs and agents.
- Excellent communication skills to drive adoption of new tools and best practices across multiple teams.
- Someone who’s very curious, driven, low-ego and eager to learn across a range of engineering disciplines, while being part of a fantastic team.
- Experience with Agentic AI systems, tools, frameworks and workflows.
- Experience with running workflows on Ray.
- Experience with MCP server patterns.
Technologies & Tools We Use
- Development: Python, Bash, LangGraph, PyTorch
- Infrastructure: Ray, Amazon EKS, Airflow, Jsonnet, Terraform
- Ops: Git, Github, AWS, LangFuse, Sentry, Prometheus, W&B
How To Really Get Our Attention
Tell us a funny joke about data quality in your application – make sure to include it at all costs.
Benefits
- Medical, Dental, and Vision insurance 100% company paid premiums
- Unlimited Paid Time Off
- 26 weeks of 100% paid Parental Leave (paternity and maternity)
- 401(k) plan with 6% employer matching
- Generous company matching on donations to non-profit charities
- Up to $20,000 tuition assistance toward degree programs, plus up to $4,000/year for ongoing professional education such as industry conferences
- Plentiful snacks, drinks, and regularly catered lunches
- Dog-friendly office (CAM office)
- Bike sharing program memberships
- Caregiver leave and elder care leave
- Mentoring and additional learning opportunities
- Opportunity to expand professional network and participate in conferences and events
Recruitment Fraud Alert
If you receive an email from a spglobalind.com domain or any other regionally based domains, it is a scam and should be reported to reportfraud@spglobal.com. S&P Global never requires any candidate to pay money for job applications, interviews, offer letters, “pre-employment training” or for equipment/delivery of equipment.