Director, AI Platform Reliability
About Us
We love going to work and think you should too. Our team is dedicated to trust, customer obsession, agility, and striving to be better every day. These values serve as the foundation of our culture, guiding our actions and driving us towards excellence. We foster a culture of performance and recognition, allowing us to transform growth as we enable our employees to do the best work of their careers. This role is open to candidates based in or near San Francisco, CA. At LogicMonitor, we hire within our Centers of Energy—vibrant locations where our teams connect, collaborate, and innovate.
What You'll Do
LogicMonitor® is the AI-first hybrid observability platform powering the next generation of digital infrastructure. LogicMonitor delivers complete visibility and actionable intelligence across on-premises, cloud, and edge environments. By anticipating issues before they strike, optimizing resources in real time, and enabling faster, smarter decisions, LogicMonitor helps IT and business leaders protect margins, accelerate innovation, and deliver exceptional digital experiences without compromise.
- Lead and scale multiple engineering teams responsible for high-volume, business-critical distributed systems and data platforms.
- Define the technical strategy and architecture for platforms processing hundreds of millions of transactions and terabytes of data.
- Guide the development of Java-based microservices, APIs, Kafka streaming pipelines, batch-processing workflows, and cloud-native services.
- Build and evolve scalable data lake and Data Lakehouse platforms supporting real-time, near-real-time, and batch analytics workloads.
- Establish reliable data ingestion, transformation, storage, governance, lineage, retention, and data-quality practices across streaming and batch pipelines.
- Build low-latency, highly available, fault-tolerant systems with strong scalability, resiliency, and disaster-recovery capabilities.
- Define and own operational SLAs, SLOs, availability targets, recovery objectives, and performance metrics for critical services and data pipelines.
- Drive capacity planning, load testing, throughput optimization, and improvements to p95 and p99 latency.
- Ensure effective Kafka design, including partitioning, consumer groups, ordering, schema evolution, replay, and lag management.
- Establish engineering standards for architecture, coding, testing, security, observability, and production readiness.
- Partner with Product, Architecture, SRE, Security, Data, and Infrastructure teams to deliver strategic platform initiatives.
- Strengthen operational excellence through monitoring, incident management, on-call practices, root-cause analysis, and continuous reliability improvements.
- Recruit, mentor, and develop engineering managers, architects, and senior technical leaders.
- Improve developer productivity, CI/CD automation, deployment safety, and release predictability.
- Manage technical debt, platform modernization, cloud costs, and long-term scalability investments.
Requirements
- 10+ years of professional software-engineering experience, including significant experience building large-scale distributed systems.
- Experience leading engineering teams, architects, and staff engineers.
- Demonstrated success delivering and operating platforms that process hundreds of millions of transactions, requests, or events.
- Deep technical expertise in Java, JVM performance, concurrency, multithreading, memory management, and application profiling.
- Strong experience designing and operating microservice-based and event-driven architectures.
- Extensive production experience with Apache Kafka or a comparable distributed streaming platform.
- Strong understanding of Kafka partitioning, replication, consumer groups, offset management, ordering, delivery semantics, schema evolution, and reprocessing.
- Experience designing low-latency, highly available APIs and backend services.
- Experience managing terabyte- or petabyte-scale datasets across relational, NoSQL, streaming, and object-storage technologies.
- Strong understanding of distributed-systems concepts, including consensus, replication, partitioning, consistency models, idempotency, backpressure, and fault tolerance.
- Experience operating cloud-native applications using Kubernetes, containers, infrastructure as code, and automated CI/CD pipelines.
- Experience with at least one major cloud platform, such as AWS, Google Cloud, or Microsoft Azure.
- Strong knowledge of observability practices involving metrics, logs, traces, profiling, alerting, dashboards, and service-level objectives.
- Demonstrated experience improving system reliability, scalability, latency, cost efficiency, and engineering productivity.
- Strong written and verbal communication skills, including the ability to explain complex technical decisions to engineering teams, executives, and business stakeholders.
- Proven ability to build inclusive, accountable, and high-performing engineering organizations.
Benefits
- Comprehensive health, dental, and vision coverage.
- Generous parental leave policies.
- Access to our Employee Assistance Program and various Wellness programs.
- 401K with company matching.
- Lifestyle Spending Account.
- Unlimited vacation policy.
Pay
The Base Salary range for this role is $247,500 USD - $275,000 USD. Compensation packages at LogicMonitor for eligible roles include base salary, a variable plan depending on role, along with comprehensive benefits.