Job Summary
As a Sr. Platform Engineer, you will be responsible for building, managing, and optimizing the underlying infrastructure and tools that enable efficient, scalable, and reliable execution of large-scale data processing workloads. Designing systems for collecting metrics (Prometheus) and visualizing data (Grafana) to provide deep insights into application/infrastructure performance. This role is a specialized subset of data platform engineering, ensuring the environment where data engineers and data scientists run their Spark jobs is robust and cost-efficient.
Job Description
RESPONSIBILITIES AND DUTIES:
- Architecting and managing the platforms where Spark runs, such as Kubernetes clusters, or cloud services like AWS (EKS).
- Packaging Spark workloads (often via Docker/Kubernetes) and integrating them with orchestration systems like Apache Flyte.
- Deploying Infrastructure via Terraform/Ansible
- Troubleshooting and resolving job failures, memory/resource issues, and execution anomalies. This includes optimizing Spark configurations to reduce cloud compute and storage costs.
- Building automation and tools in languages like Python, Java, or Scala, Linux Scripting (Bash) to increase the productivity of development teams.
- Write medium to complex SQL Queries as needed.
- Implementing and maintaining systems for monitoring, logging, and alerting (e.g., Prometheus, Grafana) to ensure platform stability and reliability.
- Develop and optimize the data catalog platform (e.g., Apache Iceberg, Unity Catalog) for authorization, search, and lineage.
- Automate workflows, monitoring, and incident resolution.
- Collaborate with Data Stewards, Analysts, and Scientists to address data needs and issues.
- Promote best practices and assess emerging technologies.
- Working closely with data engineers, data scientists, and other engineering teams to define requirements, advise on best practices, and ensure successful delivery of data objectives.
- Engaging with open-source communities (like Apache Spark, Delta Lake, or Apache Iceberg) to discuss technical challenges and contribute improvements.
- Create and maintain comprehensive documentation for Kubernetes infrastructure, processes, and procedures. Provide training and support to team members as needed.
QUALIFICATIONS:
- Bachelor's degree in computer science or a related field, or equivalent experience, typically 7 years in a DevOps or Systems Engineering role.
- Expertise in Apache Spark: Deep understanding of Spark architecture, including RDDs, DataFrames, execution hierarchy, lazy evaluation, shuffling, and fault tolerance.
- Proficiency in languages used for Spark development and automation, such as Python, Pyspark and Scala/Java.
- Proficient in Linux Scripting (Bash).
- Proficient in writing SQL.
- Experience in CI/CD tools, Github.
- Experience in setting up and using observability tools like Prometheus, Grafana etc.,
- Strong knowledge on Networking Protocols (TCP/IP, DNS, Load Balancer etc.,) and hardware components,
- Automation via Terraform/Ansible
- Hands-on experience with on-prem and major cloud providers (AWS, Azure, GCP) and container orchestration tools like Docker and Kubernetes.
- Handson experience setting up IAM, VPC, EC2 etc.,
- Familiarity with related technologies and formats like Delta Lake, Apache Iceberg, Apache Kafka, Hadoop, and various data storage systems (S3, HDFS, etc.).
- Hands-on experience with Databricks, Snowflake, Apache Iceberg, Unity Catalog, or similar tools.
- Solid understanding of data lakes and governance.
- Experience setting up, maintaining caching layers like Alluxio.
- Strong analytical skills for debugging complex distributed systems issues.
- Strong communication and collaboration abilities.
Disclaimer: This information has been designed to indicate the general nature and level of work performed by employees in this role. It is not designed to contain or be interpreted as a comprehensive inventory of all duties, responsibilities and qualifications.
Comcast is an equal opportunity workplace. We will consider all qualified applicants for employment without regard to race, color, religion, age, sex, sexual orientation, gender identity, national origin, disability, veteran status, genetic information, or any other basis protected by applicable law.
Skills:
Kubernetes; Terraform (Software); Amazon S3; Databricks Platform; Linux; Adobe Spark; Go Programming Language
Base pay is one part of the Total Rewards that Comcast provides to compensate and recognize employees for their work. Most sales positions are eligible for a Commission under the terms of an applicable plan, while most non-sales positions are eligible for a Bonus. Additionally, Comcast provides best-in-class Benefits to eligible employees. We believe that benefits should connect you to the support you need when it matters most, and should help you care for those who matter most. That's why we provide an array of options, expert guidance and always-on tools, that are personalized to meet the needs of your reality - to help support you physically, financially and emotionally through the big milestones and in your everyday life. Please visit the compensation and benefits summary on our careers site for more details.
Education
Bachelor's Degree
While possessing the stated degree is preferred, Comcast also may consider applicants who hold some combination of coursework and experience, or who have extensive related professional experience.
Relevant Work Experience
7-10 Years
Skills Required
- Bachelor's degree in computer science or related field, or equivalent experience
- 7+ years in DevOps or Systems Engineering roles
- Expertise in Apache Spark (architecture, RDDs, DataFrames, shuffling, fault tolerance)
- Proficiency in Python and PySpark
- Experience with Scala or Java for Spark development
- Proficient in Linux scripting (Bash)
- Proficient in writing SQL (medium to complex queries)
- Experience with CI/CD and GitHub
- Experience deploying infrastructure with Terraform and/or Ansible
- Hands-on experience with Kubernetes and Docker
- Hands-on experience with cloud providers and services (AWS, Azure, GCP) including EKS, EC2, VPC, IAM
- Experience implementing observability/monitoring (Prometheus, Grafana) and logging/alerting
- Familiarity with data lake technologies and formats (Delta Lake, Apache Iceberg, Unity Catalog)
- Hands-on experience with Databricks, Snowflake, or similar platforms
- Familiarity with streaming and storage systems (Apache Kafka, Hadoop, S3, HDFS)
- Experience with caching layers such as Alluxio
- Strong networking knowledge (TCP/IP, DNS, load balancers) and hardware components
- Experience troubleshooting distributed systems, job failures, memory and resource issues
- Strong communication and collaboration skills
What We Do
Welcome to Comcast. From the connectivity and platforms we provide to the content and experiences we create, we bring people together, globally. Our people think the world of our work, and that’s why our work is the best in the world.
Why Work With Us
We believe you can achieve extraordinary things when you feel connected - to the work you do and who you do it with. From the platforms we provide to millions of people, to the content and experiences we create - we bring our customers, viewers and teammates closer together across the globe.
Gallery
Comcast Offices
Hybrid Workspace
Employees engage in a combination of remote and on-site work.

.png)




.png)


