Senior Machine Learning Operations Engineer
BetMGM is a leading brand in sports betting and online gaming, with over 1,400 team members revolutionizing the industry in the United States and Canada. We combine technology with driven talent to create exceptional experiences for our players.
About the Role
The Senior MLOps Engineer treats ML systems as software systems and owns the path from a trained model to a production endpoint that meets its latency, cost, and reliability budgets. This includes both batch scoring (SageMaker Batch Transform, Snowflake Cortex / Snowpark ML, dbt-orchestrated scoring) and real-time inference (SageMaker real-time endpoints, Lambda + Bedrock, sub-second feature serving). The role involves building the platform that data scientists and ML engineers ship on, including a feature store with guaranteed online/offline parity, model registry, CI/CD for ML, drift and quality monitoring, and deployment scaffolding. A software-engineering-first mindset is essential, with distributed systems, observability, and on-call instincts as the foundation. ML literacy and GenAI integration experience are beneficial.
Responsibilities
- ML Production Platform: Stand up and operate BetMGM's ML platform on AWS (SageMaker Training, Model Registry, Pipelines, Endpoints, Batch Transform) and Snowflake (Snowpark ML, Cortex), with Terraform-managed infrastructure. Build self-service scaffolds for data scientists to ship models end-to-end without ticket queues, including CI, drift monitoring, alerting, IaC, and Snowflake connectivity.
- Batch and Real-Time Inference:
- Design and operate batch scoring pipelines (SageMaker Batch Transform, dbt-orchestrated scoring against Snowflake, Snowpark ML) with explicit freshness and cost SLAs.
- Design and operate real-time inference paths (SageMaker real-time endpoints, Lambda + Bedrock for GenAI, API Gateway) with stated latency budgets (typically sub-100ms) and graceful degradation under load.
- Own the feature store (SageMaker Feature Store, Tecton, or Feast) with guaranteed online/offline parity; treat training-serving skew as an incident.
- CI/CD and Deployment Patterns: Build CI/CD for ML, including model registry, automated retraining triggers, model versioning, and lineage tracking. Implement champion/challenger, shadow deployments, and canary releases as platform primitives.
- Monitoring, Drift & Reliability: Stand up drift detection, data quality, and model performance monitoring (Evidently, Arize, or SageMaker Model Monitor) with paging for incident response. Own MLOps incident response as SEV events with postmortems.
- Cost and Performance: Right-size endpoints, batch caching, request batching, and autoscaling to meet cost-per-prediction targets.
- GenAI Integration (Plus, Not Required):
- Integrate LLM APIs (Bedrock, Anthropic, OpenAI) into production paths, including RAG pipelines, agent eval frameworks, prompt versioning, and cost/latency observability.
- Partner with the Helix team on AI personalization workloads for March Madness 2027.
- AI in the Engineering Loop: Direct AI coding agents (Claude Code, Cursor, GitHub Copilot, dbt Copilot) to enhance infrastructure code, eval suites, and model-serving glue.
- Collaboration:
- Partner with data engineering on shared standards (Terraform modules, CI/CD patterns, observability, lineage).
- Work alongside data scientists and analytics partners to define interfaces between research and production.
- Coordinate with Entain India and contractor ML partners to consolidate workloads onto the BetMGM-owned platform.
Qualifications
BS or MS in Computer Science, Math, Statistics, Machine Learning, or other STEM field — or equivalent practical experience. Practical experience is prioritized over academic credentials.
- Must-Haves:
- 5+ years shipping software in production (Python, Docker, Kubernetes or ECS, CI/CD, distributed systems debugging), including on-call experience.
- 3+ years operating ML in production, owning models with stated latency and cost budgets and a written runbook.
- AWS depth: SageMaker (Training, Endpoints, Batch Transform, Model Registry, Pipelines) and supporting services (IAM, Lambda, ECS, S3, Secrets Manager, VPC).
- Snowflake fluency: Snowpark ML, Cortex, dbt-orchestrated batch scoring, RBAC for ML workloads.
- IaC for ML: Terraform + SageMaker Pipelines or equivalent; no manual console deployments to production.
- Feature store experience: SageMaker Feature Store, Tecton, or Feast, with ownership of online/offline parity.
- Champion/challenger, shadow, and canary deployment patterns as production muscle.
- Drift and model monitoring: Evidently, Arize, WhyLabs, or SageMaker Model Monitor, wired to a paging path.
- Software-engineering-first mindset: treat ML systems as systems, not notebooks.
- Nice-to-Haves:
- GenAI in production: Bedrock, Anthropic, or OpenAI APIs integrated into live systems; RAG pipelines; vector DBs (Snowflake Cortex Search, pgvector, Pinecone); evaluation frameworks (Langfuse or in-house).
- Snowflake-native ML: Snowpark Container Services, Cortex AISQL, Cortex Agents.
- Streaming feature engineering: Kafka, Flink, or Snowpipe Streaming for sub-second features.
- Fine-tuning experience: LoRA, QLoRA, instruction tuning, eval-driven iteration.
- Track record of shipping more with AI in the engineering loop.
- Regulated-industry experience (gaming, fintech, healthcare) with comfort in model risk, audit, and lineage requirements.
Benefits
- Medical, dental, vision, life, and disability insurance.
- 401(k) with company match.
- Pre-tax spending accounts (health care FSA and commuter savings).
- Flexible paid time off.
- Professional development reimbursement and ongoing skills training.
- Employee resource groups.
- Swag, ticket giveaways, and more.
BetMGM is committed to building a respectful, inclusive workplace that reflects our values and fosters a culture of belonging.
Pay
The annual salary range for this position is $135,000 to $170,000, with factors such as geography, skills, education, and experience influencing starting pay. This role is also eligible for a performance-based bonus plan.
Applicants must possess legal authorization to work in the U.S. without immigration sponsorship. This role is not eligible for H-1B, O-1, E-3, TN, OPT, or other immigration-related sponsorship.
Gaming Compliance & Licensing Requirements
As an online gaming company, BetMGM requires applicable employees to be licensed by jurisdictional agencies, which may include multiple agencies depending on the role. The licensing process involves comprehensive background checks, including criminal records, financial history, and personal background verification. Candidates must also comply with and support BetMGM's responsible gambling policies, procedures, and initiatives.