Jobs · OTHR

AI Specialist

SLAC National Accelerator Laboratory · United States · 1 wk ago
RemoteRemoteOTHRFull-time

Join the IT team at SLAC National Accelerator Laboratory as an AI Specialist within our AI/Cloud Services team. We are seeking a highly skilled, collaborative, and motivated professional with experience designing and operationalizing artificial intelligence and machine learning solutions across Amazon Web Services (AWS), Google Cloud Platform (GCP), and hybrid on-premises environments. In this role, you will help establish and expand SLAC’s enterprise AI capabilities in support of scientific research, accelerator operations, and administrative functions.

Responsibilities

  • Lead the end-to-end development and operationalization of AI and machine learning solutions across AWS, GCP, and hybrid environments, including problem definition, data ingestion, feature engineering, model development, evaluation, deployment, monitoring, and lifecycle management.
  • Partner with researchers, business stakeholders, data scientists, software engineers, cybersecurity, and platform teams to understand requirements and translate them into scalable, secure, cost-effective, and supportable AI solutions.
  • Design and implement data pipelines, orchestration workflows, model-training environments, evaluation processes, and deployment architectures using cloud-native services.
  • Use AWS services such as Amazon Bedrock, SageMaker, EC2, S3, Glue, Lambda, Athena, Redshift, Step Functions, and related analytics, security, and monitoring services.
  • Use Google Cloud services such as Vertex AI, Gemini, BigQuery, Cloud Storage, Dataflow, Dataproc, Cloud Run, Cloud Functions, Pub/Sub, and related analytics, security, and monitoring services.
  • Develop and support generative AI solutions, including retrieval-augmented generation, enterprise search, prompt management, model routing, model evaluation, guardrails, and AI agents and workflows.
  • Evaluate and optimize traditional machine learning and generative AI models for accuracy, reliability, latency, scalability, security, and cost-effectiveness.
  • Establish MLOps and LLMOps capabilities, including source control, infrastructure as code, CI/CD, automated testing, model and prompt versioning, evaluation, observability, drift detection, logging, alerting, and rollback procedures.
  • Design solutions that integrate cloud AI services with SLAC’s on-premises infrastructure, enterprise applications, scientific data sources, identity systems, networking, and security services.
  • Apply cloud architecture and security best practices, including identity and access management, least-privilege access, encryption, secrets management, network segmentation, data protection, audit logging, compliance, resilience, and cost governance.
  • Help develop reusable AI platforms, reference architectures, templates, APIs, and shared services that enable SLAC teams to innovate without creating unnecessary duplication or isolated solutions.
  • Work with cybersecurity, privacy, legal, data owners, and governance stakeholders to assess data sensitivity, third-party model usage, information-sharing requirements, intellectual property considerations, and other AI-related risks.
  • Promote responsible AI practices, including transparency, human oversight, explainability, fairness, accountability, privacy, security, and appropriate documentation of model limitations.
  • Conduct technical evaluations and proofs of concept for emerging AI/ML technologies and provide clear recommendations based on business value, scientific value, risk, supportability, interoperability, and total cost of ownership.
  • Troubleshoot complex technical issues spanning AI models, data pipelines, cloud services, APIs, networking, identity, security, and hybrid infrastructure.
  • Provide technical leadership, mentoring, and knowledge sharing to team members who are developing their cloud, data, and AI skills.
  • Create and maintain architecture diagrams, technical standards, operational runbooks, support procedures, model documentation, decision records, and service documentation.
  • Prepare and deliver technical presentations, demonstrations, training workshops, and model-explainability reports for technical and non-technical audiences.
  • Collaborate with cloud providers, consultants, vendors, Stanford University partners, and other external organizations while ensuring that SLAC retains the knowledge needed to operate and support its services.
  • Stay current with developments across AWS, Google Cloud, open-source AI frameworks, foundation models, AI agents, data platforms, and responsible AI practices.

Requirements

  • A bachelor’s degree in information technology, computer science, data science, engineering, or a related field and ten years of increasingly responsible technical experience, or an equivalent combination of education and relevant experience.
  • Demonstrated experience designing, building, deploying, and supporting AI/ML solutions in production cloud environments.
  • Substantial experience with AWS or GCP AI/ML and data services, along with the ability and willingness to develop proficiency across both platforms.
  • Experience with relevant AWS technologies such as Amazon Bedrock, SageMaker, S3, Glue, Lambda, Athena, Redshift, and related services.
  • Experience with relevant Google Cloud technologies such as Vertex AI, Gemini, BigQuery, Cloud Storage, Dataflow, Cloud Run, Pub/Sub, and related services.
  • Strong programming skills in Python and experience with relevant languages or frameworks such as SQL, Java, R, Scala, PyTorch, TensorFlow, scikit-learn, Hugging Face, LangChain, or similar technologies.
  • Experience developing generative AI applications using foundation models, APIs, embeddings, vector databases, retrieval-augmented generation, prompt engineering, model evaluation, and AI agent or workflow frameworks.
  • Experience with data engineering, including data ingestion, cleansing, transformation, metadata, feature engineering, data quality, large-scale datasets, data warehouses, and distributed processing technologies such as Spark.
  • Experience deploying and maintaining models in production through MLOps or LLMOps practices, including CI/CD, automated testing, monitoring, evaluation, drift detection, logging, and lifecycle management.
  • Experience with infrastructure as code and automation technologies such as Terraform, CloudFormation, AWS CDK, or Google Cloud deployment tooling.
  • Understanding of cloud architecture practices across networking, identity and access management, encryption, secrets management, observability, resilience, performance optimization, and cost management.
  • Experience integrating cloud services with on-premises systems in a hybrid enterprise environment.
  • Knowledge of data governance, privacy, cybersecurity, responsible AI, and risk-management principles applicable to enterprise and research environments.
  • Strong analytical and troubleshooting skills, including the ability to diagnose issues that cross application, data, model, cloud-platform, security, and network boundaries.
  • Strong written and verbal communication skills, with the ability to explain complex technical concepts, risks, limitations, and tradeoffs to both technical and non-technical audiences.
  • Demonstrated ability to document solutions thoroughly and create operationally useful architecture diagrams, standards, procedures, and runbooks.
  • Demonstrated ability to learn independently and adapt to rapidly changing AI, cloud, data, and security technologies.

Preferred Qualifications

  • Experience supporting AI, scientific computing, research, higher education, government, or regulated environments.
  • Experience designing multi-cloud architectures or enabling applications that can use models and services across multiple cloud providers.
  • Experience with Kubernetes and container platforms such as Amazon EKS, Google Kubernetes Engine, Docker, or related technologies.
  • Experience with enterprise AI gateways, model-routing platforms, API management, vector databases, data catalogs, and observability platforms.
  • Experience evaluating and integrating commercial and open-source foundation models.
  • Familiarity with high-performance computing, GPU-based workloads, distributed model training, or large-scale scientific datasets.
  • Familiarity with frameworks and standards such as the NIST AI Risk Management Framework, NIST security controls, or comparable responsible-AI and cybersecurity practices.
  • Relevant AWS, Google Cloud, machine learning, data engineering, security, or Kubernetes certifications.

Skills

  • Effective Decisions: Uses job knowledge and sound judgment to make quality decisions in a timely manner.
  • Self-Development: Pursues a variety of opportunities to continue learning and developing.
  • Dependability: Can be counted on to deliver results and accepts personal responsibility for expected outcomes.
  • Initiative: Pursues work and interactions proactively, with optimism, positive energy, and motivation to move initiatives forward.
  • Adaptability: Responds constructively to change and maintains an open outlook while adjusting to evolving needs.
  • Communication: Ensures effective information flow across audiences and creates a shared understanding.
  • A collaborative and service-oriented approach, with an interest in understanding the needs of researchers, engineers, business teams, and operational staff before proposing a solution.
  • The ability to balance rapid experimentation with the security, reliability, governance, and long-term support requirements of a national laboratory.
  • Comfort working in an evolving environment where requirements, platforms, and organizational priorities may not yet be fully defined.
  • A practical, platform-oriented mindset that favors reusable capabilities, open standards, interoperability, and shared solutions over isolated or vendor-specific implementations.
  • The judgment to determine when a solution should use AWS, GCP, on-premises infrastructure, open-source technologies, or a combination of platforms.
  • An understanding that scientific and accelerator workloads may have requirements that differ from traditional enterprise IT, including large datasets, specialized computing, low-latency operations, and long-lived research workflows.
  • The ability to work effectively with teams at different levels of cloud and AI maturity, including mentoring colleagues and enabling others rather than becoming a single point of dependency.
  • A willingness to be hands-on—building, testing, troubleshooting, documenting, and operationalizing solutions in addition to providing architectural guidance.
  • Intellectual curiosity and a willingness to ask questions, challenge assumptions constructively, and evaluate technologies based on evidence.
  • Strong ownership and follow-through, including the ability to move an initiative from early discovery through production readiness and operational handoff.
  • An appreciation for knowledge sharing, transparency, and clear documentation so that services can be operated and improved by the broader team.
  • The ability to communicate limitations and risks honestly while remaining focused on helping teams identify a workable path forward.

Similar jobs

AI Specialist

Bealls, Inc.Bradenton Beach, FL· 2 mo ago
OTHRapply on beallsinc.taleo.net

AI Specialist

LaddersUnited States· 1 mo ago
RemoteOTHR$183k–$251k/yrapply on theladders.com

AI Specialist

Stanford UniversityUnited States· 1 mo ago
RemoteOTHRapply on careersearch.stanford.edu

AI Specialist

RotheraChicago, IL· 1 mo ago
OTHR$125k–$150k/yrapply on grnh.se

AI Specialist

Intellibee IncPontiac, MI· 1 mo ago
OTHRapply on intellibee.my.salesforce-sites.com

AI Specialist

LaunchDarklyUnited States· 1 mo ago
RemoteOTHR$215k–$295k/yrapply on job-boards.greenhouse.io