Jobs · Massachusetts

Senior Principal AI Engineer

Vertex Pharmaceuticals · Boston, MA · 6 days ago
Hybrid$188k–$282k/yrFull-time

About the role

Vertex is seeking a Senior Principal AI Engineer to define and build the foundational enterprise AI platform that powers intelligent applications across the enterprise. This role will lead the design and implementation of scalable, secure, and reusable capabilities for agentic AI, with a strong focus on retrieval-augmented generation (RAG), orchestration frameworks, evaluation systems, and platform architecture. The engineer will own the vision and day-to-day operations of a centralized AI Gateway / Control Plane / Control Tower enabling agent monitoring, observability, policy enforcement, governance, and operational controls across AI solutions at Vertex.

Responsibilities

  • Own the strategy, architecture, implementation, and day-to-day operation of the centralized AI Gateway / Control Tower as the enterprise control point for model access, routing, governance, cost management, and operational oversight
  • Design and build the shared agentic AI platform, including architecture, services, and reusable capabilities for faster and safer AI-powered application development
  • Lead AI Control Tower operations: agent monitoring, observability, telemetry, policy enforcement, guardrails, usage analytics, and auditability
  • Implement intelligent model routing, provider abstraction, fallback, failover, rate limiting, and workload optimization across approved models
  • Develop AI FinOps capabilities: token and consumption visibility, budgeting, chargeback/showback, cost allocation, forecasting, and model-cost optimization
  • Establish centralized guardrails for content safety, prompt-injection defense, data-loss prevention, sensitive-data handling, and responsible AI policy enforcement
  • Define identity, access, security, privacy, regulatory compliance, and lifecycle governance controls for models, agents, tools, and AI interactions
  • Manage end-to-end observability, telemetry, quality monitoring, latency and reliability metrics, incident response, and operational health
  • Provide usage analytics, immutable audit trails, policy evidence, risk reporting, and executive-level transparency across the enterprise AI estate
  • Oversee model onboarding, approval, versioning, deprecation, resiliency, capacity management, and third-party provider governance
  • Build retrieval and knowledge systems: RAG pipelines, vector search, document retrieval, and grounding patterns for enterprise content
  • Manage agent lifecycle and quality: deployment, versioning, evaluation frameworks, reliability measurement, and continuous improvement in production
  • Define platform APIs, service contracts, architecture patterns, and engineering practices for other teams
  • Design and build reusable platform services accelerating development of safe, reliable, and scalable AI-powered applications
  • Lead architecture and implementation for RAG pipelines, knowledge retrieval systems, prompt workflows, tool use, and agent orchestration
  • Establish core frameworks for agent lifecycle management, including deployment, monitoring, observability, evaluation, and continuous improvement
  • Develop scalable infrastructure patterns for enterprise AI workloads, including model integration, data access, and orchestration services
  • Partner with product, engineering, data, security, and UX teams to deliver common AI platform capabilities supporting multiple use cases
  • Drive technical standards for platform APIs, service contracts, architecture patterns, and reusable components
  • Evaluate emerging technologies, frameworks, and vendors in the AI/agentic ecosystem and make strategic recommendations
  • Mentor engineers and influence cross-functional technical teams through architectural leadership and hands-on guidance
  • Ensure platform solutions align with enterprise requirements for scalability, resilience, security, and maintainability
  • Contribute to Vertex’s long-term AI strategy by identifying opportunities to expand platform capabilities and increase enterprise adoption

Requirements

  • Advanced degree in Computer Science, Engineering, Artificial Intelligence, Machine Learning, or a related technical field; or equivalent combination of education and experience
  • 10+ years of experience designing and building enterprise-grade AI/ML platforms and distributed systems
  • Deep expertise in agentic AI architectures, LLM-based applications, and platform engineering
  • Proven experience with retrieval-augmented generation (RAG) systems, vector search, document retrieval, and knowledge integration patterns
  • Strong experience with AI orchestration frameworks, workflow engines, and multi-step agent execution patterns
  • Demonstrated experience designing centralized operational platforms for monitoring, governance, observability, and control
  • Experience using AI-assisted software development and autonomous coding agents to design, generate, test, review, debug, optimize, and refactor code across complex enterprise systems
  • Deep understanding of AI-native software engineering practices and experience establishing standards, governance, and best practices for responsible use of AI coding assistants
  • Experience defining architecture, standards, and reusable services for large-scale enterprise environments
  • Strong understanding of AI system evaluation, quality measurement, and performance optimization
  • Excellent communication, leadership, and problem-solving skills
  • Ability to balance strategic architecture leadership with hands-on technical execution

Skills

  • Agentic AI platform architecture
  • Large Language Models (LLMs)
  • Retrieval-Augmented Generation (RAG)
  • AI orchestration and workflow design
  • Agent monitoring and observability
  • AI governance, policies, and controls
  • AI FinOps, token economics, budgeting, cost allocation, consumption forecasting, and model-cost optimization
  • AI Gateway / Control Tower architecture, including model routing, provider abstraction, fallback, rate limiting, and policy enforcement
  • Evaluation frameworks for AI systems
  • Platform engineering and reusable service design
  • Distributed systems architecture
  • API and service design
  • Knowledge retrieval systems
  • Telemetry, logging, and operational analytics
  • Scalability, reliability, and performance engineering
  • Security and enterprise controls for AI platforms

Preferred Skills

  • Experience building enterprise AI platforms in regulated or highly governed environments
  • Familiarity with human-in-the-loop workflows and responsible AI practices
  • Experience implementing policy engines, guardrails, and audit frameworks for AI applications
  • Knowledge of ML infrastructure, model serving, and production AI operations
  • Experience with cloud-native architectures and modern DevOps/MLOps practices
  • Exposure to user experience considerations for AI-powered applications and intelligent systems
  • Ability to translate complex technical capabilities into scalable enterprise adoption strategies
  • Experience leading technical teams through platform transformation initiatives

Pay

The pay range for this role is $188,000 - $282,000. This role is eligible for an annual bonus and annual equity awards. Some roles may also be eligible for overtime pay, in accordance with federal and state requirements. Actual base salary pay will be based on skills, competencies, experience, and other job-related factors permitted by law.

Benefits

  • Inclusive market-leading medical, dental, and vision benefits
  • Generous paid time off, including a week-long company shutdown in the Summer and Winter
  • Educational assistance programs, including student loan repayment
  • Generous commuting subsidy
  • Matching charitable donations
  • 401(k) retirement plan

Similar jobs