Staff Engineer, Data and AI
This is a senior technical leadership role focused on building and scaling the data and AI infrastructure that powers high-impact products, experimentation, personalization, and trusted decision-making. You will help translate strategic business and product opportunities into scalable data and AI solutions, from early exploration through production. The role combines hands-on engineering with architecture, technical leadership, and cross-functional collaboration.
About the role
You’ll design reliable batch and real-time data systems while improving platform scalability, observability, security, and cost efficiency. You’ll also help advance modern AI capabilities, including LLM infrastructure, vector retrieval, evaluation, and agentic workflows. As a Staff-level engineer, you’ll raise engineering standards through mentorship, design leadership, and influence across teams.
Responsibilities
- Partner with product managers and cross-functional teams to translate business and product opportunities into scalable data and AI solutions, from initial exploration through production deployment.
- Design, build, and operate reliable batch and real-time data pipelines supporting reporting, experimentation, machine learning, feature generation, model evaluation, personalization, and production inference.
- Evolve cloud-native data and ML platform architecture across orchestration, storage, compute, streaming, and data access, using technologies such as Airflow, Kafka, Redshift, Glue, vector databases, and related services.
- Establish and improve data quality and trust through robust data modeling, schema evolution, data contracts, automated testing, lineage, privacy controls, freshness monitoring, and recoverability practices.
- Develop reusable tools, engineering standards, and self-service workflows that enable teams to discover data and build dependable data and AI solutions with greater independence.
- Maintain platform resilience and operational excellence through observability, alerting, runbooks, incident response, on-call participation, performance optimization, reliability improvements, and cost management.
- Lead architectural decisions and technical initiatives across data and AI infrastructure while reducing technical debt and establishing scalable engineering patterns.
- Contribute to the development and adoption of modern AI and agentic workflows, including embeddings, vector retrieval, LLM evaluation and observability, and production AI systems.
- Mentor engineers through design reviews, code reviews, technical guidance, and hands-on coaching while helping raise the overall technical bar.
- Participate in an on-call rotation and take ownership of the reliability and operational health of critical data and AI systems.
Requirements
- 7+ years of software engineering experience, with substantial experience building distributed systems, data platforms, ML platforms, or comparable production infrastructure.
- Proven hands-on experience designing, building, and operating large-scale batch and/or streaming data systems, ideally with technologies such as Kafka, Spark, and workflow orchestration platforms.
- Strong software engineering skills with the ability to build maintainable, testable, production-grade systems using Python and/or comparable programming languages.
- Advanced data modeling and SQL expertise, with experience creating scalable, reliable data models and data products.
- Experience designing and operating cloud-native infrastructure using AWS, GCP, or comparable cloud platforms, including infrastructure as code, containers, orchestration, and managed data services.
- Demonstrated experience taking data, ML, or AI systems into production, including deployment, reliability, observability, evaluation, and ongoing operational ownership.
- Strong understanding of distributed data architecture and the ability to make sound trade-offs across compute, storage, orchestration, streaming, infrastructure, scalability, and cost.
- Practical experience with modern AI infrastructure, such as embeddings and vector retrieval, LLM evaluation and observability, or agentic workflows.
- Strong operational mindset with experience in monitoring, incident response, on-call practices, runbooks, performance tuning, and building resilient systems.
- Demonstrated technical leadership, including architectural decision-making, mentoring, code and design reviews, and the ability to influence engineering direction across teams.
- Strong collaboration and communication skills, with the ability to work effectively with technical and non-technical stakeholders.
- Professional-level English proficiency; resumes and application responses should be submitted in English.
- Ability to work full-time from an eligible location in the United States, Canada, or Mexico.
Benefits
- Remote-first flexibility: Full-time opportunity with eligible employees based in the United States, Canada, or Mexico, subject to approved locations.
- Professional growth: Opportunities to mentor engineers, influence technical strategy, and contribute to advanced data and AI initiatives.
- Inclusive workplace: A collaborative environment that values diverse experiences, skills, and perspectives.
- Accessibility support: Reasonable accommodations are available throughout the recruitment process for candidates who need them.
Pay
- U.S. annual base salary: $236,000 in San Francisco and New York; $224,000 in Austin, Boston, Chicago, Washington, DC, Los Angeles, and Seattle; and $200,500 in other eligible U.S. locations.
- Canadian annual base salary: $221,500 CAD in Vancouver and Toronto; $205,500 CAD in Victoria and Calgary; and $202,000 CAD in other eligible Canadian locations.
- Mexico annual base salary: $1,727,000 MXN across eligible locations in Mexico.