AI Operations Architect & AWS Engineer, Life Sciences
About the role
We are seeking an LLM Operations Architect & AWS Engineer to design, deploy, and operate production-grade LLM systems on AWS infrastructure at Norstella. You will build and maintain the agentic AI pipelines and cloud architectures that power large-scale clinical data extraction from real-world clinical notes, directly accelerating RWD-driven drug development and treatment access for patients with unmet need. This role sits at the intersection of ML infrastructure, cloud engineering, and applied NLP — you will own the operational backbone that turns LLM capabilities into reliable, scalable data products.
Responsibilities
- Design and maintain AWS cloud architectures optimized for LLM inference and fine-tuning workloads, including GPU instance management, auto-scaling, and cost optimization across spot and on-demand capacity.
- Build and operate agentic AI pipelines — multi-step LLM orchestration workflows with tool use, structured output extraction, retrieval-augmented generation (RAG), and automated quality control loops.
- Deploy and manage open-weight LLMs (Llama, Gemma, Qwen, and successors) in production, including model serving infrastructure, batching strategies, and latency/throughput optimization.
- Implement robust evaluation and monitoring frameworks for LLM-based extraction systems: automated accuracy measurement, drift detection, prompt regression testing, and human-in-the-loop QC integration.
- Develop and maintain CI/CD pipelines for prompt versioning, model updates, and configuration management across development, staging, and production environments.
- Collaborate with NLP scientists, data engineers, and software developers to translate prototype extraction logic into production-hardened, recurrent processing feeds.
- Ensure compliance with data security requirements (HIPAA, PHI handling) in all LLM deployment and data processing workflows, working closely with IT security and compliance teams.
- Analyze and optimize computational resource utilization — GPU hours, storage, network throughput — balancing cost efficiency against processing SLAs.
- Evaluate and integrate emerging LLM serving technologies, orchestration frameworks, and inference optimization techniques (quantization, speculative decoding, structured generation) to maintain operational edge.
Requirements
- Master's in Computer Science, Data Science, or related field.
- 5+ years in production NLP/ML systems development and operations on AWS.
- Hands-on experience deploying and serving open-weight LLMs (Llama, Gemma, Qwen families) at scale, using frameworks such as vLLM, TGI, or equivalent serving infrastructure.
- Expert Python proficiency and working knowledge of LLM orchestration tooling (e.g., LangChain/LangGraph, custom agent frameworks).
- Experience with AWS ML/compute services (SageMaker, EC2 GPU instances, S3, Step Functions, Lambda) and infrastructure-as-code (Terraform, CloudFormation, or CDK).
- Proven track record deploying large-scale NLP/LLM solutions in healthcare or life sciences.
- Strong analytical and communication skills; comfortable operating across multiple concurrent projects.
Preferred Qualifications
- AWS certifications (Solutions Architect, Machine Learning Specialty, or DevOps Engineer).
- Direct experience with PHI handling and HIPAA-compliant ML infrastructure.
- Experience with AWS Redshift, OpenSearch/Elasticsearch, or similar large-scale data stores.
- Familiarity with ML experiment tracking and lifecycle management (MLflow, Weights & Biases, or equivalent).
- Experience building or operating RAG systems, structured extraction pipelines, or clinical NLP applications.
Benefits
At Norstella, we offer a competitive compensation package, comprehensive benefits, and a supportive work environment. We also provide opportunities for professional growth and development, as well as a culture that values diversity, equity, and inclusion.
Pay
The salary range for this position is $100,000 - $150,000 annually, depending on experience and qualifications.
Schedule
This role is typically full-time, with standard business hours. Remote work is possible for qualified candidates.