Sr. Machine Learning Engineer
About the role
The Senior Software Engineer will architect and optimize high-throughput, event-driven systems using Apache Kafka to handle real-time data flows. They will also build and maintain large-scale data pipelines using Apache Spark or Flink to provide high-volume analytics that power our AI. Additionally, they will design sophisticated AI agents capable of autonomous planning, memory management, and high-reliability tool-use across distributed environments. The role involves leading the architectural design of containerized services on Kubernetes, ensuring high availability and scalability across cloud infrastructure (AWS/Azure/GCP).
Responsibilities
- Architect and optimize high-throughput, event-driven systems using Apache Kafka to handle real-time data flows.
- Build and maintain large-scale data pipelines using Apache Spark or Flink to provide high-volume analytics that power our AI.
- Design sophisticated AI Agents capable of autonomous planning, memory management, and high-reliability tool-use across distributed environments.
- Lead the architectural design of containerized services on Kubernetes, ensuring high availability and scalability across cloud infrastructure (AWS/Azure/GCP).
Requirements
- 5–8 years of experience in backend engineering using Java, Python, or Go.
- Expertise in distributed systems, asynchronous architectures (Kafka), and large-scale data processing (Spark/Flink).
- Hands-on experience with agentic frameworks (e.g., AutoGen, CrewAI, or custom orchestration layers), RAG, MCP, fine tuning models and prompt engineering.
- Agentic observability using Langfuse, Evals frameworks for Testing/Resilience.
Qualifications
- Advanced IaC: Expertise in building reusable Terraform modules and managing complex multi-region cloud deployments.
- Vector DB Optimization: Deep experience in indexing strategies (HNSW vs IVF) and performance tuning for high-concurrency vector databases at scale.
- AI Ops: Experience with LLM deployment optimization (e.g., vLLM, TensorRT-LLM) or managing proprietary model inference endpoints.
Skills
- Experience with cloud-native technologies.
- Strong understanding of machine learning and AI concepts.
- Ability to work independently and in a team environment.
- Excellent problem-solving skills and attention to detail.
Benefits
Our benefits package includes:
- Competitive compensation and equity options.
- Flexible working hours and remote work options.
- Professional development opportunities and training programs.
- Health, dental, and vision insurance.
- Retirement savings plan with employer match.
- Employee assistance program.
Pay
Salary range: $120,000 - $160,000 annually.
Schedule
Full-time, Monday through Friday, 9 AM to 5 PM.
Company Culture
We foster a culture of innovation, ownership, and fun. We believe in enabling ownership at all levels of the organization and empowering teams. Join us to shape the future of cybersecurity!