Jobs · Engineering · California

Sr. Machine Learning Engineer

Illumio · San Jose, CA · 1 wk ago
On-siteEngineeringFull-time

About the role

The Senior Software Engineer will architect and optimize high-throughput, event-driven systems using Apache Kafka to handle real-time data flows. They will also build and maintain large-scale data pipelines using Apache Spark or Flink to provide high-volume analytics that power our AI. Additionally, they will design sophisticated AI agents capable of autonomous planning, memory management, and high-reliability tool-use across distributed environments. The role involves leading the architectural design of containerized services on Kubernetes, ensuring high availability and scalability across cloud infrastructure (AWS/Azure/GCP).

Responsibilities

  • Architect and optimize high-throughput, event-driven systems using Apache Kafka to handle real-time data flows.
  • Build and maintain large-scale data pipelines using Apache Spark or Flink to provide high-volume analytics that power our AI.
  • Design sophisticated AI Agents capable of autonomous planning, memory management, and high-reliability tool-use across distributed environments.
  • Lead the architectural design of containerized services on Kubernetes, ensuring high availability and scalability across cloud infrastructure (AWS/Azure/GCP).

Requirements

  • 5–8 years of experience in backend engineering using Java, Python, or Go.
  • Expertise in distributed systems, asynchronous architectures (Kafka), and large-scale data processing (Spark/Flink).
  • Hands-on experience with agentic frameworks (e.g., AutoGen, CrewAI, or custom orchestration layers), RAG, MCP, fine tuning models and prompt engineering.
  • Agentic observability using Langfuse, Evals frameworks for Testing/Resilience.

Qualifications

  • Advanced IaC: Expertise in building reusable Terraform modules and managing complex multi-region cloud deployments.
  • Vector DB Optimization: Deep experience in indexing strategies (HNSW vs IVF) and performance tuning for high-concurrency vector databases at scale.
  • AI Ops: Experience with LLM deployment optimization (e.g., vLLM, TensorRT-LLM) or managing proprietary model inference endpoints.

Skills

  • Experience with cloud-native technologies.
  • Strong understanding of machine learning and AI concepts.
  • Ability to work independently and in a team environment.
  • Excellent problem-solving skills and attention to detail.

Benefits

Our benefits package includes:

  • Competitive compensation and equity options.
  • Flexible working hours and remote work options.
  • Professional development opportunities and training programs.
  • Health, dental, and vision insurance.
  • Retirement savings plan with employer match.
  • Employee assistance program.

Pay

Salary range: $120,000 - $160,000 annually.

Schedule

Full-time, Monday through Friday, 9 AM to 5 PM.

Company Culture

We foster a culture of innovation, ownership, and fun. We believe in enabling ownership at all levels of the organization and empowering teams. Join us to shape the future of cybersecurity!

Similar jobs