Jobs · Engineering · Pennsylvania

Sr. Platform Engineer, Kubernetes

Comcast · West Chester, PA · 1 mo ago
On-siteEngineeringFull-time

Responsibilities

  • Architecting and managing the platforms where Spark runs, such as Kubernetes clusters, or cloud services like AWS (EKS).
  • Packaging Spark workloads (often via Docker/Kubernetes) and integrating them with orchestration systems like Apache Flyte.
  • Deploying Infrastructure via Terraform/Ansible.
  • Troubleshooting and resolving job failures, memory/resource issues, and execution anomalies. This includes optimizing Spark configurations to reduce cloud compute and storage costs.
  • Building automation and tools in languages like Python, Java, or Scala, Linux Scripting (Bash) to increase the productivity of development teams.
  • Writing medium to complex SQL Queries as needed.
  • Implementing and maintaining systems for monitoring, logging, and alerting (e.g., Prometheus, Grafana) to ensure platform stability and reliability.
  • Create and maintain comprehensive documentation for Kubernetes infrastructure, processes, and procedures.
  • Provide training and support to team members as needed.

Qualifications

  • Bachelor's degree in computer science or a related field, or equivalent experience, typically 7 years in a DevOps or Systems Engineering role.
  • Expertise in Apache Spark: Deep understanding of Spark architecture, including RDDs, DataFrames, execution hierarchy, lazy evaluation, shuffling, and fault tolerance.
  • Proficiency in languages used for Spark development and automation, such as Python, Pyspark and Scala/Java.
  • Proficient in Linux Scripting (Bash).
  • Proficient in writing SQL.
  • Experience in CI/CD tools, Github.
  • Experience in setting up and using observability tools like Prometheus, Grafana etc.
  • Strong knowledge on Networking Protocols (TCP/IP, DNS, Load Balancer etc.,) and hardware components.
  • Automation via Terraform/Ansible.
  • Hands-on experience with on-prem and major cloud providers (AWS, Azure, GCP) and container orchestration tools like Docker and Kubernetes.
  • Familiarity with related technologies and formats like Delta Lake, Apache Iceberg, Apache Kafka, Hadoop, and various data storage systems (S3, HDFS, etc.).
  • Hands-on experience with Databricks, Snowflake, Apache Iceberg, Unity Catalog, or similar tools.
  • Solid understanding of data lakes and governance.
  • Experience setting up, maintaining caching layers like Alluxio.
  • Strong analytical skills for debugging complex distributed systems issues.
  • Strong communication and collaboration abilities.

Similar jobs