Senior Kafka SRE Engineer
About the role
We are seeking a Kafka Site Reliability Engineer to help build, operate, and continuously improve Schwab's enterprise streaming platform ecosystem. This role combines deep expertise in Confluent Kafka technologies with modern Site Reliability Engineering practices to deliver highly available, secure, and resilient streaming services that support critical business capabilities across the firm.
As a member of the team, you will drive operational excellence through automation, observability, Infrastructure as Code, and AIOps-driven capabilities. You will play a key role in designing, deploying, and supporting Confluent Kafka environments across on-premises and cloud platforms while helping engineering teams deliver reliable real-time data solutions. Success in this role requires the ability to proactively identify risks, solve complex technical challenges, improve platform performance, and enhance system reliability through data-driven decision-making and continuous improvement.
You will contribute to the evolution of Kafka infrastructure supporting capabilities such as Schema Registry, Kafka Connect, ksqlDB, Cluster Linking, and cloud-native deployments. Through the application of automation, predictive analytics, and AI-driven operational insights, you will help reduce operational toil, strengthen platform resiliency, improve incident response, and accelerate issue resolution.
The ideal candidate thrives in highly distributed environments and enjoys partnering with engineers, architects, and infrastructure teams to improve platform health, streamline deployment processes, increase observability, and establish best practices across the streaming ecosystem. This role offers the opportunity to influence the future of event streaming at Schwab while developing expertise in emerging AIOps, cloud, automation, and reliability engineering capabilities.
Responsibilities
- Support and administer enterprise-scale production Confluent Kafka platforms, including Confluent Platform and Confluent Cloud.
- Develop Python automation, operational tooling, observability dashboards, and alerting solutions.
- Apply AIOps concepts, including anomaly detection, event correlation, alert reduction, predictive analytics, automated remediation, and AI-driven operational insights within production environments.
- Leverage predictive monitoring and intelligent automation to improve platform reliability, incident management, and operational efficiency.
- Deploy and operate Kafka platforms on Kubernetes, preferably Google Kubernetes Engine (GKE), using Helm charts and cloud-native operational practices.
- Use Terraform and Ansible to automate infrastructure provisioning, configuration management, and platform lifecycle activities.
- Work with the Confluent Kafka ecosystem, including ZooKeeper, KRaft, Schema Registry, Kafka Connect, ksqlDB, Cluster Linking, MirrorMaker, and role-based access controls (RBAC).
- Support large-scale distributed systems, highly available platforms, fault-tolerant architectures, and production-critical workloads.
- Perform incident response, root-cause analysis, problem resolution, and post-incident continuous improvement activities.
- Implement and maintain observability solutions, operational metrics, service-level objectives (SLOs), dashboards, and actionable monitoring controls.
Requirements
- 5-7 years of experience supporting and administering enterprise-scale production Confluent Kafka platforms.
- 5-7 years of experience developing Python automation, operational tooling, observability dashboards, and alerting solutions.
- Hands-on experience applying AIOps concepts in production environments.
- Experience deploying and operating Kafka platforms on Kubernetes, preferably Google Kubernetes Engine (GKE).
- Strong experience using Terraform and Ansible for infrastructure automation.
- Deep expertise with the Confluent Kafka ecosystem, including ZooKeeper, KRaft, Schema Registry, Kafka Connect, ksqlDB, Cluster Linking, MirrorMaker, and RBAC.
- 3+ years of experience working with public cloud technologies, with Google Cloud Platform (GCP) preferred.
- Strong understanding of Kafka architecture and client internals, including producers, consumers, partitions, replication, serialization, consumer groups, performance tuning, and exactly-once processing.
- Strong Linux administration, troubleshooting, performance tuning, and networking experience supporting high-throughput distributed systems.
- Experience supporting large-scale distributed systems and highly available platforms.
- Experience performing incident response, root-cause analysis, and problem resolution.
- Experience implementing and maintaining observability solutions, SLOs, and monitoring controls.
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.
Preferred Qualifications
- Confluent certifications such as CCDAK or CCAAK.
- Google Cloud certifications such as Professional Cloud Architect or Professional Cloud DevOps Engineer.
- Experience operating Confluent for Kubernetes (CFK) and Confluent Cloud APIs.
- Experience implementing or supporting AIOps platforms for predictive incident detection and self-healing infrastructure.
- Experience with observability and monitoring technologies such as Grafana, InfluxDB, BigQuery, Prometheus, Splunk, or Datadog.
- Experience with CI/CD technologies such as GitHub Actions or Cloud Build.
- Experience supporting additional messaging and streaming technologies such as RabbitMQ, IBM MQ, Solace, or Google Pub/Sub.
- Understanding of modern Site Reliability Engineering practices, including SLIs, SLOs, error budgets, and reliability engineering.
- Strong communication, collaboration, and relationship-building skills.
- Demonstrated ability to adapt to changing priorities and drive initiatives independently.
Benefits
- 401(k) with company match and Employee Stock Purchase Plan.
- Paid time for vacation, volunteering, and a 28-day sabbatical after every 5 years for eligible positions.
- Paid parental leave and family building benefits.
- Tuition reimbursement.
- Health, dental, and vision insurance.
In addition to the salary range, this role is eligible for bonus or incentive opportunities.