Senior Cloud Infrastructure Engineer, AI Platform
Procore Technologies · Austin, TX · 1 mo ago
On-siteInformation TechnologyFull-time
About the role
We are hiring a Senior Cloud Infrastructure Engineer to build the core infrastructure that powers the next generation of AI agents. Our agents must ingest, enrich, and vectorize massive, multi-modal datasets from over 100 customer sources. The core challenge is twofold: how do we do this in a way that is radically cost-efficient, while still allowing agents to deliver thorough responses in seconds? This position reports to a Senior Manager, Software Engineering and will be 2 days per week hybrid role in our Austin office.
Responsibilities
- Build & Optimize for Scale & Cost: Implement and optimize our highly scalable, multi-tenant data ingestion and vectorization pipeline, with a relentless focus on improving cost and performance.
- Implement Robust Monitoring: Create and maintain dashboards, alerts, and logging to ensure system health, identify performance bottlenecks, and provide immediate visibility into production issues.
- Contribute to Millisecond Latency: Be a key contributor in performance tuning across the stack. You will help identify and eliminate bottlenecks in data retrieval, model inference, and agent response times to ensure a snappy, real-time user experience.
- Engineer Multi-Tenant Architectures: Design and implement scalable, secure, and cost-effective multi-tenant infrastructures for our SaaS-based AI products, ensuring strict tenant isolation and fair resource allocation.
Requirements
- 5+ years of hands-on experience with cloud platforms
- Hands-on experience with cloud platforms, with strong expertise in Google Cloud Platform (GCP) and Infrastructure as Code (e.g., Terraform)
- A demonstrated history of performance analysis, tuning, and infrastructure cost optimization, with the ability to speak about trade-offs and quantified impact
- Experience building or working on multi-tenant SaaS platforms
- Experience setting up end-to-end observability, including logging, metrics, and alerting using tools like Prometheus, Grafana, Datadog, or GCP Operations suite
Preferred Qualifications
- A fundamental understanding of the challenges in training and serving large machine learning models (e.g., memory constraints, computational complexity)
- Strong understanding of VectorDBs, LLMs, and Agentic Observability tools (Datagrid uses Milvus for our VectorDB and Arize for agent tracing)
- Experience with the Gemini API, specifically managing LLM quotas and load balancing
- Hands-on experience with LLM serving frameworks and optimization techniques (quantization, tensor parallelism, FlashAttention)
- Experience designing multi-tenant SaaS architectures and implementing resource quotas and cost allocation
- High-level knowledge of agentic systems and best practices
Pay
Base Pay Range: $140,960.00 - $193,820.00 USD Annual. This role may also be eligible for Equity Compensation and/or Bonus Incentive Compensation.