Sr. Platform Engineer, Kubernetes
Comcast · West Chester, PA · 1 mo ago
On-siteEngineeringFull-time
Responsibilities
- Architecting and managing the platforms where Spark runs, such as Kubernetes clusters, or cloud services like AWS (EKS).
- Packaging Spark workloads (often via Docker/Kubernetes) and integrating them with orchestration systems like Apache Flyte.
- Deploying Infrastructure via Terraform/Ansible.
- Troubleshooting and resolving job failures, memory/resource issues, and execution anomalies. This includes optimizing Spark configurations to reduce cloud compute and storage costs.
- Building automation and tools in languages like Python, Java, or Scala, Linux Scripting (Bash) to increase the productivity of development teams.
- Writing medium to complex SQL Queries as needed.
- Implementing and maintaining systems for monitoring, logging, and alerting (e.g., Prometheus, Grafana) to ensure platform stability and reliability.
- Create and maintain comprehensive documentation for Kubernetes infrastructure, processes, and procedures.
- Provide training and support to team members as needed.
Qualifications
- Bachelor's degree in computer science or a related field, or equivalent experience, typically 7 years in a DevOps or Systems Engineering role.
- Expertise in Apache Spark: Deep understanding of Spark architecture, including RDDs, DataFrames, execution hierarchy, lazy evaluation, shuffling, and fault tolerance.
- Proficiency in languages used for Spark development and automation, such as Python, Pyspark and Scala/Java.
- Proficient in Linux Scripting (Bash).
- Proficient in writing SQL.
- Experience in CI/CD tools, Github.
- Experience in setting up and using observability tools like Prometheus, Grafana etc.
- Strong knowledge on Networking Protocols (TCP/IP, DNS, Load Balancer etc.,) and hardware components.
- Automation via Terraform/Ansible.
- Hands-on experience with on-prem and major cloud providers (AWS, Azure, GCP) and container orchestration tools like Docker and Kubernetes.
- Familiarity with related technologies and formats like Delta Lake, Apache Iceberg, Apache Kafka, Hadoop, and various data storage systems (S3, HDFS, etc.).
- Hands-on experience with Databricks, Snowflake, Apache Iceberg, Unity Catalog, or similar tools.
- Solid understanding of data lakes and governance.
- Experience setting up, maintaining caching layers like Alluxio.
- Strong analytical skills for debugging complex distributed systems issues.
- Strong communication and collaboration abilities.