Staff Software Development Engineer - Enterprise AI Infrastructure - #4898
Our mission is to detect cancer early, when it can be cured. We are working to change the trajectory of cancer mortality and bring stakeholders together to adopt innovative, safe, and effective technologies that can transform cancer care. We are a healthcare company, pioneering new technologies to advance early cancer detection. We have built a multi-disciplinary organization of scientists, engineers, and physicians and we are using the power of next-generation sequencing (NGS), population-scale clinical studies, and state-of-the-art computer science and data science to overcome one of medicine’s greatest challenges.
This role is based in Sunnyvale, California (new headquarters opening September) or Durham, NC. We offer a flexible hybrid schedule requiring a minimum of 60% (24 hours) on-site, with Tuesdays and Thursdays as key collaboration days.
About the role
The Staff Software Development Engineer - Enterprise AI Infrastructure is a senior technical role responsible for leading the design, development, and scaling of a centralized, highly governed enterprise AI platform. This position serves as a technical expert focused on AWS and Kubernetes-based (EKS) AI infrastructure, agentic development, and multi-agent orchestration operating within a regulated environment. The Staff Engineer partners closely with cross-functional stakeholders across Software Engineering, Data Science, Security, Regulatory Affairs, and Product to build a unified control plane that securely connects large language models with enterprise tools and company knowledge. This role drives technical excellence in cloud infrastructure, container orchestration, AI governance, identity-scoped integrations, and agentic workflows while mentoring engineering teams and advancing the organization's enterprise AI strategy.
Responsibilities
- Lead the end-to-end design, development, deployment, and monitoring of a scalable, governed enterprise AI platform leveraging Amazon EKS and AWS native services (e.g., Bedrock, OpenSearch Serverless, KMS, VPC).
- Design and implement agentic AI workflows, specialized autonomous agents, and multi-agent systems using advanced LLM orchestration techniques and agent frameworks.
- Architect and manage secure integrations using the Model Context Protocol (MCP) to connect the AI platform with internal systems, vector databases, and third-party SaaS applications (e.g., Google Workspace, Slack).
- Build and enforce strict identity, authorization, and zero-trust token brokering flows leveraging Okta, Auth0, and custom JWT authorizers to ensure secure, least-privilege tool execution.
- Implement deterministic policy controls (e.g., Cedar policy engine) to enforce role-based access, approval gates, and human-in-the-loop checks at the API gateway level.
- Develop and maintain highly isolated, scalable containerized runtime environments (e.g., Kubernetes pods on Amazon EKS) for secure AI model execution, tool usage, and knowledge retrieval.
- Establish and maintain comprehensive audit trails and observability for all AI interactions, utilizing AWS CloudTrail and GenAI observability tools (e.g., OpenTelemetry) to track cost, latency, and tool calls.
- Collaborate with Product Management, Security, Regulatory, and business stakeholders to translate enterprise requirements into scalable, compliant AI infrastructure solutions.
- Troubleshoot and resolve complex technical issues involving cloud infrastructure, Kubernetes networking, network isolation (PrivateLink), and agentic workflows.
- Contribute to technology roadmaps, AI infrastructure strategy, and long-term platform evolution initiatives.
- Mentor engineers, software developers, and technical teams while promoting engineering excellence, infrastructure-as-code (IaC) best practices, and continuous improvement.
- Partner with Quality, Regulatory, Privacy, Security, and Compliance functions to ensure software and AI systems operate in accordance with applicable regulatory requirements and company policies.
Adaptability and Growth Expectations
- Flexibility in responsibilities and duties as the organization evolves, including taking on additional responsibilities, participating in cross-functional projects, or adapting to new technologies and methodologies.
- Supporting other departments or teams during periods of high demand or contributing to special projects as needed.
Requirements
- Bachelor's degree or equivalent in Computer Science, Software Engineering, Artificial Intelligence, Cloud Computing, or related field; Master's or PhD preferred.
- 8-12 years of relevant software development and cloud infrastructure experience with demonstrated technical leadership.
- Deep expertise in AWS cloud architecture and container orchestration, specifically with Amazon EKS, Kubernetes networking, network isolation (VPC, PrivateLink), IAM, KMS, and GenAI services (e.g., AWS Bedrock).
- Proven experience in agentic AI development, building autonomous agents, and orchestrating LLM tool-calling workflows using frameworks like LangChain, LangGraph, AutoGen, or Claude Agent SDK.
- Hands-on experience implementing the Model Context Protocol (MCP) or building robust, governed API/tool integrations for LLMs.
- Strong background in identity and access management (IAM), OAuth, JWT, and integrating with enterprise IdPs (Okta, Auth0) for scoped, token-based authorization.
- Advanced proficiency in programming languages such as Python, TypeScript, or Go, and infrastructure-as-code tools (Terraform, AWS CDK).
- Experience with vector databases, RAG (Retrieval-Augmented Generation) architectures, and row-level access controls (e.g., OpenSearch, FAISS, pgvector).
- Proficiency with CI/CD pipelines, MLOps practices, Kubernetes ecosystem tools (e.g., Helm), containerization, and modern observability stacks.
- Demonstrated knowledge of applicable regulatory standards, including:
- Cybersecurity principles, tools, and control frameworks (e.g., ISO 27001, NIST, SOC 2, HIPAA).
- Operations within the regulated medical device environment (e.g., IVDD, IVDR, FDA 21 CFR 800 series, FDA 21 CFR Part 11).
- AI governance, software validation, data integrity, and risk management principles applicable to regulated environments.
- Exceptional problem-solving and analytical skills with the ability to address ambiguous, high-impact technical challenges in AI orchestration and Kubernetes scaling.
- Strong leadership and influence skills, capable of driving alignment across engineering, security, regulatory, and business stakeholders.
- Excellent communication skills with the ability to explain complex LLM behaviors, infrastructure architectures, and security boundaries to technical and non-technical audiences.
- Proven mentoring and coaching capabilities that elevate cloud engineering and AI talent.
- Strong understanding of AI safety, prompt injection defenses, secure tool execution, and deterministic policy enforcement.
- Strategic thinking with the ability to balance long-term enterprise AI platform vision with near-term business delivery.
- High adaptability and intellectual curiosity regarding emerging agentic AI frameworks, MCP specifications, and cloud computing trends.
Physical Working Conditions
- Standard office or hybrid work environment depending on company policy.
- Frequent use of software development tools, AI/ML platforms, cloud infrastructure, data engineering tools, and collaboration systems.
- May require extended hours during major project deadlines, AI model deployments, production incidents, regulatory audits, or strategic initiatives.
- Operates with significant independence and responsibility and is expected to provide leadership on complex software, AI, technical, and organizational decisions.
Pay
The expected, full-time, annual base pay scale for this position is $169,000–$224,000 in Sunnyvale, CA and $147,000–$195,000 in Durham, NC. This role may be eligible for other forms of compensation, including an annual bonus and/or incentives, subject to the terms of the applicable plans and company discretion.
Benefits
- Flexible time-off or vacation.
- 401(k) retirement plan with employer match.
- Medical, dental, and vision coverage.
- Mindfulness programs.