Senior Artificial Intelligence Engineer
About the role
High Performance Research Computing (HPRC) and the Center for Bioimage Informatics (CBI) at St. Jude Children's Research Hospital are seeking a Senior AI Engineer to lead efforts in advanced AI models, including large language models (LLMs), agentic AI systems, and multi-modal foundation models, and the secure computational infrastructure that powers them. This is a hands-on, high-ownership software systems role focused on evaluating, fine-tuning, deploying, and benchmarking AI models; designing safety guardrails and sandboxing for agentic systems; safeguarding data security and privacy; and optimizing GPU/HPC resource allocation. The role involves building shared AI architecture, secure environments, and best practices for CBI's image data scientists and software engineers. Deep bioimaging expertise is not required, though experience with biomedical research or imaging is a plus. This is an onsite role in Memphis, TN.
Responsibilities
- Assess institutional needs and implement AI-enabled services on high-performance AI computing platforms.
- Collaborate with scientists, business stakeholders, analysts, and IS professionals to design, develop, and implement AI solutions, including proof-of-concepts, pilots, and adoption lifecycle.
- Build innovative solutions leveraging data science and AI/ML to solve non-trivial problems.
- Develop AI demonstration use-cases, workshop materials, and provide training.
- Evaluate commercial and open-source AI/ML, data mining, and analytics approaches to solve business problems.
- Identify, develop, and implement standards and operating procedures for solutions and systems consistent with best practices.
- Document current and future state architecture roadmaps and reference architectures.
- Provide clear written and spoken communications to customers, teams, and vendors.
- Stay abreast of new and emerging technologies and assess their potential applicability.
- Perform other duties as assigned to meet departmental and institutional goals.
Requirements
- Master’s degree in computer science, computer engineering, data science, information technology, or related field required. PhD preferred.
- Minimum of 4 years of experience designing and developing solutions for large-scale AI/Machine Learning (ML) systems or building solutions for products with AI/ML features.
- Strong background in industry use cases built on deep learning and machine learning (unsupervised and supervised techniques).
- Proficiency with deep learning frameworks such as TensorFlow, Keras, PyTorch.
- Experience with time series analysis, anomaly detection, forecasting, predictive modeling, graph-based neural networks, Bayesian statistics, and text analytics.
Qualifications
- Hands-on experience training, fine-tuning, or adapting LLMs or multi-modal foundation models (e.g., PEFT/LoRA, instruction tuning, preference optimization), including debugging failure modes such as catastrophic forgetting or training instability.
- Experience diagnosing and resolving distributed/multi-GPU training or inference issues (e.g., NCCL communication hangs, CUDA out-of-memory errors, load-balancing across nodes) and scheduling AI workloads on HPC (e.g., Slurm) to maximize GPU utilization.
- Experience designing safety guardrails and sandboxing for agentic systems: tool-access scoping, prompt-injection defense, secrets management, audit logging, and containment of failures.
- Experience optimizing inference cost, latency, and resource usage (e.g., KV-cache management, quantization, batching, speculative decoding, high-throughput serving via vLLM/TensorRT-LLM/Triton) and judgment about when a simpler deterministic pipeline is a better fit than an agentic one.
- Experience safeguarding data security and privacy for AI systems handling sensitive research data, including access controls and institutional data-use/security policies.
- Experience building rigorous, reproducible benchmarking/evaluation frameworks that separate genuine model improvement from prompt overfitting, retrieval effects, or evaluator bias.
- Experience building and scaling AI/ML pipelines and workflows on HPC or cloud environments.
- Demonstrated ownership of a system beyond the prototype stage (observability, versioning, rollback, cost control, incident response).
- Contributions to open-source AI/ML infrastructure projects (e.g., vLLM, PyTorch, Ray, Hugging Face) or a public track record (GitHub, Hugging Face, papers) are a plus.
- Familiarity with biomedical research, imaging, or regulated health data is a plus.
- Demonstrated technical leadership: setting standards, mentoring, and cross-team collaboration.
Pay
A reasonable estimate of the current salary range is $86,320 - $154,960 per year for the role of Senior Artificial Intelligence Engineer.
Benefits
Explore our exceptional benefits.