AI Specialist
About the role
Join the IT team at SLAC National Accelerator Laboratory as an AI Specialist within our AI/Cloud Services team. We are seeking a highly skilled, collaborative, and motivated professional with experience designing and operationalizing artificial intelligence and machine learning solutions across Amazon Web Services (AWS), Google Cloud Platform (GCP), and hybrid on-premises environments.
Responsibilities
- Lead the end-to-end development and operationalization of AI and machine learning solutions across AWS, GCP, and hybrid environments, including problem definition, data ingestion, feature engineering, model development, evaluation, deployment, monitoring, and lifecycle management.
- Partner with researchers, business stakeholders, data scientists, software engineers, cybersecurity, and platform teams to understand requirements and translate them into scalable, secure, cost-effective, and supportable AI solutions.
- Design and implement data pipelines, orchestration workflows, model-training environments, evaluation processes, and deployment architectures using cloud-native services.
- Evaluate and optimize traditional machine learning and generative AI models for accuracy, reliability, latency, scalability, security, and cost-effectiveness.
- Establish MLOps and LLMOps capabilities, including source control, infrastructure as code, CI/CD, automated testing, model and prompt versioning, evaluation, observability, drift detection, logging, alerting, and rollback procedures.
- Design solutions that integrate cloud AI services with SLAC’s on-premises infrastructure, enterprise applications, scientific data sources, identity systems, networking, and security services.
- Apply cloud architecture and security best practices, including identity and access management, least-privilege access, encryption, secrets management, network segmentation, data protection, audit logging, compliance, resilience, and cost governance.
- Help develop reusable AI platforms, reference architectures, templates, APIs, and shared services that enable SLAC teams to innovate without creating unnecessary duplication or isolated solutions.
- Work with cybersecurity, privacy, legal, data owners, and governance stakeholders to assess data sensitivity, third-party model usage, information-sharing requirements, intellectual property considerations, and other AI-related risks.
- Promote responsible AI practices, including transparency, human oversight, explainability, fairness, accountability, privacy, security, and appropriate documentation of model limitations.
- Conduct technical evaluations and proofs of concept for emerging AI/ML technologies and provide clear recommendations based on business value, scientific value, risk, supportability, interoperability, and total cost of ownership.
- Troubleshoot complex technical issues spanning AI models, data pipelines, cloud services, APIs, networking, identity, security, and hybrid infrastructure.
- Provide technical leadership, mentoring, and knowledge sharing to team members who are developing their cloud, data, and AI skills.
- Create and maintain architecture diagrams, technical standards, operational runbooks, support procedures, model documentation, decision records, and service documentation.
- Prepare and deliver technical presentations, demonstrations, training workshops, and model-explainability reports for technical and non-technical audiences.
- Collaborate with cloud providers, consultants, vendors, Stanford University partners, and other external organizations while ensuring that SLAC retains the knowledge needed to operate and support its services.
- Stay current with developments across AWS, Google Cloud, open-source AI frameworks, foundation models, AI agents, data platforms, and responsible AI practices.
Requirements
Demonstrated experience designing, building, deploying, and supporting AI/ML solutions in production cloud environments. Substantial experience with AWS or GCP AI/ML and data services, along with the ability and willingness to develop proficiency across both platforms. Experience with relevant AWS technologies such as Amazon Bedrock, SageMaker, S3, Glue, Lambda, Athena, Redshift, and related services. Experience with relevant Google Cloud technologies such as Vertex AI, Gemini, BigQuery, Cloud Storage, Dataflow, Cloud Run, Cloud Functions, Pub/Sub, and related services. Strong programming skills in Python and experience with relevant languages or frameworks such as SQL, Java, R, Scala, PyTorch, TensorFlow, scikit-learn, Hugging Face, LangChain, or similar technologies. Experience developing generative AI applications using foundation models, APIs, embeddings, vector databases, retrieval-augmented generation, prompt engineering, model evaluation, and AI agent or workflow frameworks. Experience with data engineering, including data ingestion, cleansing, transformation, metadata, feature engineering, data quality, large-scale datasets, data warehouses, and distributed processing technologies such as Spark. Experience deploying and maintaining models in production through MLOps or LLMOps practices, including CI/CD, automated testing, monitoring, evaluation, drift detection, logging, and lifecycle management. Experience with infrastructure as code and automation technologies such as Terraform, CloudFormation, AWS CDK, or Google Cloud deployment tooling. Understanding of cloud architecture practices across networking, identity and access management, encryption, secrets management, observability, resilience, performance optimization, and cost management. Experience integrating cloud services with on-premises systems in a hybrid enterprise environment. Knowledge of data governance, privacy, cybersecurity, responsible AI, and risk-management principles applicable to enterprise and research environments. Strong analytical and troubleshooting skills, including the ability to diagnose issues that cross application, data, model, cloud-platform, security, and network boundaries. Strong written and verbal communication skills, with the ability to explain complex technical concepts, risks, limitations, and tradeoffs to both technical and non-technical audiences. Demonstrated ability to document solutions thoroughly and create operationally useful architecture diagrams, standards, procedures, and runbooks. Demonstrated ability to learn independently and adapt to rapidly changing AI, cloud, data, and security technologies.
Qualifications
A bachelor’s degree in information technology, computer science, data science, engineering, or a related field and ten years of increasingly responsible technical experience, or an equivalent combination of education and relevant experience.
Skills
Strong programming skills in Python and experience with relevant languages or frameworks such as SQL, Java, R, Scala, PyTorch, TensorFlow, scikit-learn, Hugging Face, LangChain, or similar technologies. Experience with relevant AWS technologies such as Amazon Bedrock, SageMaker, S3, Glue, Lambda, Athena, Redshift, and related services. Experience with relevant Google Cloud technologies such as Vertex AI, Gemini, BigQuery, Cloud Storage, Dataflow, Cloud Run, Cloud Functions, Pub/Sub, and related services. Strong programming skills in Python and experience with relevant languages or frameworks such as SQL, Java, R, Scala, PyTorch, TensorFlow, scikit-learn, Hugging Face, LangChain, or similar technologies. Experience with data engineering, including data ingestion, cleansing, transformation, metadata, feature engineering, data quality, large-scale datasets, data warehouses, and distributed processing technologies such as Spark. Experience deploying and maintaining models in production through MLOps or LLMOps practices, including CI/CD, automated testing, monitoring, evaluation, drift detection, logging, and lifecycle management. Experience with infrastructure as code and automation technologies such as Terraform, CloudFormation, AWS CDK, or Google Cloud deployment tooling. Understanding of cloud architecture practices across networking, identity and access management, encryption, secrets management, observability, resilience, performance optimization, and cost management. Experience integrating cloud services with on-premises systems in a hybrid enterprise environment. Knowledge of data governance, privacy, cybersecurity, responsible AI, and risk-management principles applicable to enterprise and research environments. Strong analytical and troubleshooting skills, including the ability to diagnose issues that cross application, data, model, cloud-platform, security, and network boundaries. Strong written and verbal communication skills, with the ability to explain complex technical concepts, risks, limitations, and tradeoffs to both technical and non-technical audiences. Demonstrated ability to document solutions thoroughly and create operationally useful architecture diagrams, standards, procedures, and runbooks. Demonstrated ability to learn independently and adapt to rapidly changing AI, cloud, data, and security technologies.
Benefits
SLAC offers a comprehensive benefits package including health insurance, retirement savings plans, paid time off, and more.
Pay
The salary range for this position is $80,000 - $120,000 annually, commensurate with experience.
Schedule
This position is full-time and works standard business hours.