Jobs · Georgia

Senior AI Platform Engineer (DevOps)

Mastercard · Atlanta, GA · 3 days ago
Hybrid$115k–$184k/yrFull-time

Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.

About The Role

The AI Platform Engineering team is responsible for building, operating, and evolving Mastercard's enterprise AI platforms and capabilities. Our mission is to provide scalable, secure, and reliable AI infrastructure that enables teams across Mastercard to accelerate the development and deployment of AI-powered solutions.

As a Senior AI Platform Engineer, you will help design, implement, and operate the foundational platforms that support AI and machine learning workloads across the enterprise. You will work at the intersection of platform engineering, cloud infrastructure, MLOps, and AI operations to deliver enterprise-grade capabilities that enable teams to safely develop, deploy, observe, and scale AI solutions. This team functions as an enterprise AI Platform Engineering organization, delivering and operating shared AI capabilities that enable application teams to build AI-powered products at scale across public and private cloud environments.

This role provides the opportunity to influence and operate the foundational AI platforms that enable innovation across Mastercard. Rather than focusing solely on individual AI models or applications, you will help build and scale the enterprise platforms, tooling, operational practices, and cloud infrastructure that support the next generation of AI capabilities across the organization. You will work on challenging problems involving platform scalability, reliability, observability, governance, automation, and customer enablement while helping shape Mastercard's long-term AI platform strategy.

Responsibilities

  • Design, build, and operate enterprise AI platforms supporting machine learning, generative AI, and advanced analytics workloads.
  • Engineer scalable solutions across public and private cloud environments, ensuring security, reliability, availability, and performance.
  • Build and automate platform capabilities that simplify onboarding, deployment, operations, and lifecycle management for AI solutions.
  • Develop and maintain infrastructure, tooling, and services that support model training, evaluation, deployment, monitoring, and governance.
  • Implement and enhance MLOps capabilities that enable repeatable, scalable, and secure AI development workflows.
  • Design and maintain observability solutions, including telemetry, performance monitoring, logging, alerting, operational analytics, and drift detection.
  • Support production AI platforms and services, proactively identifying opportunities to improve reliability, scalability, efficiency, and customer experience.
  • Partner with internal engineering teams to understand requirements, enable platform adoption, and accelerate delivery of AI-powered products.
  • Collaborate with infrastructure, security, architecture, and governance teams to ensure alignment with enterprise standards, controls, and regulatory requirements.
  • Evaluate emerging AI technologies, platform capabilities, and industry trends to help shape the future direction of Mastercard's AI ecosystem.
  • Drive automation and engineering best practices through Infrastructure as Code, CI/CD, testing, and operational excellence initiatives.
  • Participate in troubleshooting, root cause analysis, operational support, and incident response activities to maintain highly available platforms.
  • Contribute to technical design discussions, architecture reviews, and long-term platform strategy.
  • Mentor peers and share knowledge across engineering teams while contributing to a culture of continuous improvement.

Requirements

  • Experience designing, building, and operating cloud-native systems in enterprise environments.
  • Strong experience working within both public and private cloud environments.
  • Experience deploying and managing containerized workloads using Kubernetes or OpenShift.
  • Strong software engineering and automation experience using Python.
  • Experience implementing CI/CD pipelines and modern DevOps practices.
  • Experience supporting production AI, machine learning, data platforms, or large-scale distributed systems.
  • Strong understanding of MLOps principles and machine learning lifecycle management.
  • Experience implementing monitoring, observability, telemetry, logging, and operational analytics solutions.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Ability to communicate effectively with technical and non-technical stakeholders.
  • Experience working in highly collaborative, cross-functional engineering environments.

Preferred Qualifications

  • Experience supporting generative AI platforms and large language model (LLM) workloads.
  • Experience implementing model evaluation, finetuning, guardrails, and model governance controls.
  • Experience with model observability, drift detection, telemetry monitoring, and operational analytics.
  • Experience with AI serving infrastructure and inference platforms.
  • Experience supporting GPU-based workloads and accelerated computing environments.
  • Experience with OpenShift, Kubernetes, Docker, Helm, GitOps, and Infrastructure as Code practices.
  • Experience with enterprise-scale platform engineering, developer enablement, and self-service platform capabilities.
  • Familiarity with vector databases, retrieval systems, AI gateways, agentic systems, or emerging AI platform technologies.
  • Experience working within regulated environments requiring strong security, governance, and compliance controls.

Benefits

  • Insurance (including medical, prescription drug, dental, vision, disability, life insurance).
  • Flexible spending account and health savings account.
  • Paid leaves (including 16 weeks of new parent leave and up to 20 days of bereavement leave).
  • 80 hours of Paid Sick and Safe Time, 25 days of vacation time, and 5 personal days (pro-rated based on date of hire).
  • 10 annual paid U.S. observed holidays.
  • 401k with a best-in-class company match.
  • Deferred compensation for eligible roles.
  • Fitness reimbursement or on-site fitness facilities.
  • Eligibility for tuition reimbursement.

Pay

  • O'Fallon, Missouri: $115,000 - $184,000 USD
  • Atlanta, Georgia: $115,000 - $184,000 USD
  • Austin, Texas: $115,000 - $184,000 USD

Similar jobs