Apache Spark Developer

Posted 4 Days Ago
Be an Early Applicant
Flower Mound, TX, USA
In-Office
125K-185K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
The Role
Design, develop, and optimize large-scale Apache Spark applications and scalable ETL/ELT pipelines for batch and streaming data. Collaborate with architects, cloud and ML teams to deploy Spark workloads on Databricks/EMR/Synapse, ensure data quality, monitor performance, troubleshoot production issues, and modernize legacy ETL into cloud-native Spark architectures.
Summary Generated by Built In
Apache Spark Developer – Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Apache Spark Developer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $125,000–$185,000 Annually
Experience Required: 6+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are seeking an experienced Apache Spark Developer to design, develop, and optimize large-scale distributed data processing applications supporting enterprise analytics, machine learning, real-time reporting, and cloud-based data platforms. This role focuses on building high-performance Spark applications capable of processing billions of records across structured and semi-structured data sources while delivering scalable, reliable, and cost-efficient data pipelines.
You will work closely with data architects, data engineers, cloud platform teams, machine learning engineers, and business intelligence developers to build modern data processing solutions leveraging Apache Spark, cloud-native technologies, and distributed computing frameworks. The ideal candidate possesses deep expertise in Spark architecture, distributed systems, performance optimization, and cloud-based big data ecosystems.
Key Responsibilities
  • Design, develop, and maintain high-performance distributed data processing applications using Apache Spark.
  • Build scalable batch and real-time ETL/ELT pipelines processing large volumes of enterprise data.
  • Develop Spark applications using PySpark, Scala, or Spark SQL for data transformation, aggregation, and analytics.
  • Optimize Spark jobs for memory utilization, partitioning strategies, shuffle performance, and execution efficiency.
  • Process structured, semi-structured, and streaming data from enterprise databases, APIs, Kafka, cloud storage, and data lakes.
  • Develop reusable Spark libraries, data processing frameworks, and metadata-driven ingestion pipelines.
  • Collaborate with cloud engineering teams to deploy Spark workloads on Databricks, EMR, Azure Synapse, or Kubernetes.
  • Implement data quality validation, reconciliation, monitoring, and automated error handling across distributed pipelines.
  • Integrate Spark applications with enterprise data warehouses, lakehouses, and reporting platforms.
  • Participate in architecture reviews, code reviews, technical design discussions, and Agile development activities.
  • Troubleshoot production issues involving distributed processing, cluster performance, resource utilization, and data quality.
  • Support cloud migration initiatives by modernizing legacy ETL workloads into Spark-based architectures.

Required Skills
  • Six or more years of professional software or data engineering experience.
  • Four or more years of hands-on Apache Spark development experience in enterprise production environments.
  • Strong proficiency in PySpark, Scala, or Spark SQL for distributed data processing.
  • Deep understanding of Apache Spark architecture including RDDs, DataFrames, Datasets, Catalyst Optimizer, DAG execution, and Tungsten engine.
  • Strong experience with distributed computing concepts including partitioning, shuffling, caching, broadcast joins, and fault tolerance.
  • Advanced SQL skills with databases such as SQL Server, Oracle, PostgreSQL, Snowflake, or Teradata.
  • Experience working with Hadoop ecosystem technologies including Hive, HDFS, YARN, and Parquet.
  • Experience processing streaming data using Spark Structured Streaming, Apache Kafka, or Event Hubs.
  • Hands-on experience with cloud platforms including Azure Databricks, AWS EMR, AWS Glue, Azure Synapse Analytics, or Google Dataproc.
  • Experience integrating Spark applications with Delta Lake, Apache Iceberg, or Apache Hudi.
  • Strong understanding of data warehousing concepts, dimensional modeling, and data lake architecture.
  • Experience using Git, CI/CD pipelines, Azure DevOps, GitHub Actions, or Jenkins.
  • Strong debugging, troubleshooting, and Spark performance tuning skills.
  • Experience working in Agile Scrum development environments.

Preferred Qualifications
  • Experience building enterprise Lakehouse architectures using Databricks or Delta Lake.
  • Familiarity with Apache Airflow, Azure Data Factory, AWS Step Functions, or Control-M for workflow orchestration.
  • Experience with machine learning workflows using Spark MLlib, MLflow, or feature engineering pipelines.
  • Knowledge of Kubernetes, Docker, and containerized Spark deployments.
  • Experience implementing Data Quality frameworks using Great Expectations or Deequ.
  • Familiarity with Apache NiFi, Apache Flink, Trino, or Presto.
  • Experience working with cloud object storage including Amazon S3, Azure Data Lake Storage (ADLS Gen2), or Google Cloud Storage.
  • Knowledge of Infrastructure as Code using Terraform or ARM templates.
  • Experience with enterprise monitoring tools including Prometheus, Grafana, Datadog, or OpenTelemetry.
  • Cloud certifications in Azure, AWS, Databricks, or Apache Spark-related technologies are highly desirable.

Project Environment
You will be joining a modern data engineering team responsible for building cloud-native big data platforms supporting enterprise analytics, AI, and business intelligence initiatives. Current projects include:
  • Enterprise data lakehouse implementation using Databricks and Delta Lake
  • Real-time streaming analytics processing billions of daily events
  • Large-scale customer analytics and behavioral data platforms
  • Financial risk modeling and fraud detection pipelines
  • Healthcare clinical and operational analytics solutions
  • Cloud migration of legacy Hadoop and ETL workloads
  • Machine learning feature engineering and model training pipelines
  • Enterprise reporting platforms supporting executive dashboards and self-service analytics
  • Distributed data processing infrastructure deployed on Azure and AWS
This is a hands-on engineering role where you will contribute to distributed system architecture, Spark application development, cloud migration, performance optimization, production support, and continuous improvement of enterprise-scale data processing platforms.
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 676-4399. Learn more about Bright Vision Technologies at www.bvteck.com.
 
Bright Vision Technologies is an Equal Opportunity Employer.
 

Skills Required

  • Six or more years of professional software or data engineering experience
  • Four or more years of hands-on Apache Spark development experience in enterprise production environments
  • Proficiency in PySpark, Scala, or Spark SQL for distributed data processing
  • Deep understanding of Apache Spark architecture (RDDs, DataFrames, Datasets, Catalyst Optimizer, DAG execution, Tungsten)
  • Experience with distributed computing concepts (partitioning, shuffling, caching, broadcast joins, fault tolerance)
  • Advanced SQL skills with SQL Server, Oracle, PostgreSQL, Snowflake, or Teradata
  • Experience with Hadoop ecosystem technologies including Hive, HDFS, YARN, and Parquet
  • Experience processing streaming data using Spark Structured Streaming, Apache Kafka, or Event Hubs
  • Hands-on experience with cloud platforms including Azure Databricks, AWS EMR, AWS Glue, Azure Synapse Analytics, or Google Dataproc
  • Experience integrating Spark applications with Delta Lake, Apache Iceberg, or Apache Hudi
  • Understanding of data warehousing concepts, dimensional modeling, and data lake architecture
  • Experience using Git and CI/CD tools such as Azure DevOps, GitHub Actions, or Jenkins
  • Strong debugging, troubleshooting, and Spark performance tuning skills
  • Experience working in Agile Scrum development environments
  • Experience building enterprise Lakehouse architectures using Databricks or Delta Lake
  • Familiarity with Apache Airflow, Azure Data Factory, AWS Step Functions, or Control-M for workflow orchestration
  • Experience with Spark MLlib, MLflow, or machine learning feature engineering pipelines
  • Knowledge of Kubernetes, Docker, and containerized Spark deployments
  • Experience implementing data quality frameworks using Great Expectations or Deequ
  • Familiarity with Apache NiFi, Apache Flink, Trino, or Presto
  • Experience with cloud object storage such as Amazon S3, ADLS Gen2, or Google Cloud Storage
  • Knowledge of Infrastructure as Code using Terraform or ARM templates
  • Experience with enterprise monitoring tools like Prometheus, Grafana, Datadog, or OpenTelemetry
  • Cloud certifications in Azure, AWS, Databricks, or Apache Spark-related technologies
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
53 Employees
Year Founded: 2020

What We Do

Bright Vision Technologies is a minority-owned organization founded in July 2020 and based in New Jersey, USA. The company specializes in delivering top-tier staffing and IT consulting services, including custom computer programming and systems design. Additionally, they are a product engineering firm with a flagship AI-powered talent intelligence and enterprise automation platform called Lumina, which helps transform IT into a strategic asset for their valued partners.

Similar Jobs

Cloudera Logo Cloudera

Staff Software Engineer

Artificial Intelligence • Cloud • Software • Big Data Analytics
In-Office or Remote
7 Locations
3092 Employees
165K-230K Annually

BlackLine Logo BlackLine

Artificial Intelligence Engineer

Cloud • Fintech • Information Technology • Machine Learning • Software • App development • Generative AI
Remote or Hybrid
USA
1810 Employees
128K-160K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Technical Director, Commercial Auto

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
8 Locations
40000 Employees
106K-197K Annually

Enverus Logo Enverus

Owner Relations Agent - 25270

Big Data • Information Technology • Software • Analytics • Energy
In-Office or Remote
3 Locations
1800 Employees
43K-58K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account