Principal Software Engineer, AI Platform
Kai · San Jose, CA · 4 wk ago
Engineering$125/hrFull-time
About the role
We are building an AI-powered cybersecurity platform that helps enterprises manage vulnerabilities at scale. Our AI team has delivered real, working products — services that use LLMs for natural language filtering, container image analysis, package standardization, and maintenance assessment. The AI science is strong. Now, the AI team needs dedicated engineering leadership to match.
Responsibilities
- Audit and roadmap. Assess the current services, architecture, and deployment practices. Produce an engineering roadmap that prioritizes the highest-impact improvements (we'll give you a head start — we know where the gaps are).
- Establish engineering foundations. Design and implement the shared library architecture — common patterns for configuration, DB access, error handling, structured logging, and health checks that all services adopt. Define the coding standards, PR review process, and quality gates.
- Build the test infrastructure. Not just write tests but design the testing strategy: how we mock LLM APIs, how we run integration tests against staging, how we measure coverage, and how it all plugs into CI. Set the standard that the team follows.
- Consolidate and fix CI/CD. You'll design a parameterized pipeline template, add test and lint stages, and establish the deployment strategy (staging, canary, rollback).
- Fix critical production issues. You'll prioritize and address the highest-risk issues while establishing patterns that prevent new ones.
- Owarchitecture decisions across the AI platform — API contracts, caching strategies, data flow, service boundaries, and technology choices.
- Mentor AI scientists on engineering practices — testing, code structure, version control workflows — in a way that improves their code without slowing their research velocity.
- Design evaluation and reliability frameworks for LLM-powered features — how we measure accuracy, detect regressions, and monitor production behavior.
- Partner with backend engineering to define how AI services integrate with the core cybersecurity platform — API versioning, contract testing, SLAs, and data flow.
- Translate business problems into technical architecture — work with product and AI leadership to turn ambiguous requirements into well-scoped engineering work.
Requirements
- 8+ years of professional software engineering experience, with at least 2 years in a technical leadership or staff+ role where you set standards for a team or organization.
- Deep software architecture expertise. You've designed shared libraries, defined API contracts, established coding standards, and made technology decisions that a team lived with for years. You think in terms of patterns, not just solutions.
- Production systems ownership at scale. You've been paged at 2am, you've run incident postmortems, you've built the monitoring that catches problems before customers do. You understand what production-grade actually means.
- Strong Python expertise with emphasis on clean architecture — dependency injection, proper module boundaries, testable design, async patterns done correctly.
- Testing leadership. You haven't just written tests — you've established testing culture. You've designed test strategies for systems with complex external dependencies.
- Cloud platform expertise. Deep experience with Azure (preferred) or AWS/GCP. You've designed and operated containerized microservice architectures in production.
- Ci/CD and DevOps maturity. You've consolidated messy build pipelines, added quality gates, implemented deployment strategies (blue/green, canary), and established release processes.
- Mentorship and technical leadership. You've raised the engineering bar on a team. You've done code reviews that teach, not just gatekeep. You've helped junior and mid-level engineers grow. You can work with AI scientists who are domain experts but still developing engineering practices — and you can do it with empathy and patience.
- Excellent communication skills. You'll interface with AI scientists, backend engineers, product managers, and executive leadership. You need to translate between these audiences fluently.
Preferred Qualifications
- Experience working with or alongside AI/ML teams — you understand the workflow of AI scientists and how to build infrastructure that serves their needs without constraining their exploration.
- Familiarity with LLM provider APIs (Anthropic, OpenAI, Azure OpenAI) and the engineering challenges of LLM integration (prompt management, output parsing, cost optimization, latency).
- Experience with Azure specifically — Cosmos DB, Container Apps, Azure Identity, Azure DevOps.
- Background in cybersecurity, vulnerability management, or compliance-sensitive environments (SOC2, data privacy).
- Track record of building engineering teams from early stage — you've been the first senior engineering hire and built the team around you.
- Experience with RAG systems, embedding pipelines, vector databases, or LLM evaluation frameworks.
- Exposure to MLOps practices — model versioning, experiment tracking, evaluation pipelines (this becomes increasingly relevant as we mature).