Sr. Data Integration Engineer

Posted 2 Days Ago
Hiring Remotely in Maryland, USA
Remote
Senior level
Information Technology • Software
The Role
Designs and operates resilient batch and event-driven data pipelines using Python and AWS services. Responsibilities include ingestion, transformation, schema validation, data quality, monitoring, security, testing, infrastructure automation, performance optimization, incident resolution, documentation, and customer requirements management. The role also supports Agile story development, acceptance criteria, UAT, risk analysis, stakeholder communication, and production maintenance.
Summary Generated by Built In

Join us at Sparksoft, where we're not just another tech company—we're a catalyst for change. Our mission isn't just to offer IT solutions; it's to revolutionize the way you work. Here, passion isn't just a buzzword; it's the fuel behind groundbreaking ideas and transformative technologies. We serve a wide range of government clients, delivering impact that's felt across the nation.

Our true strength lies in our people. They're the problem-solvers and innovators consistently delivering extraordinary outcomes. With Sparksoft, you're not stepping into a routine job; you're joining a team committed to innovation and excellence. Our innovation extends beyond just delivering projects. Through our specialized Innovation Centers, we continuously refine our methods, ensuring we remain industry leaders.

We are Sparksoft!

ROLE AND RESPONSIBILITIES:

  • Design and develop resilient batch and event-driven data pipelines using Python, AWS Glue, and appropriate AWS managed services.
  • Build AWS Glue jobs, crawlers, workflows, triggers, and Data Catalog integrations to discover, transform, govern, and publish datasets.
  • Ingest and process structured and semi-structured data from files, APIs, databases, and streaming or messaging sources, with particular expertise in JSON and NDJSON formats.
  • Develop efficient Python components for parsing, schema validation, normalization, enrichment, deduplication, aggregation, and data quality checks.
  • Use AWS services such as Amazon S3, AWS Lambda, Amazon EventBridge, AWS Step Functions, Amazon SQS, Amazon SNS, Amazon Kinesis, Amazon Athena, Amazon Redshift, AWS Lake Formation, AWS Secrets Manager, AWS KMS, Amazon CloudWatch, and AWS IAM as solution needs dictate.
  • Create automated unit, integration, regression, and data reconciliation tests; embed data quality controls throughout the pipeline lifecycle.
  • Implement operational monitoring, logging, alerting, traceability, restartability, error handling, and recovery mechanisms for production pipelines.
  • Apply security and privacy requirements through least-privilege access, encryption, secure secret management, audit logging, and appropriate handling of sensitive data.
  • Automate infrastructure and deployment processes using infrastructure as code and CI/CD practices.
  • Optimize pipeline performance, reliability, scalability, and cost through profiling, tuning, service selection, and ongoing operational analysis.
  • Investigate and resolve performance issues, failed jobs, and production incidents; document root causes and preventive actions.
  • Create and maintain technical documentation, including source-to-target mappings, pipeline designs, data contracts, runbooks, lineage, test evidence, and operational procedures.
  • Analyze and organize requirements into stories under epics, generate acceptance criteria, lead refinement sessions and work with Product Owners on prioritization.
  • Work with external teams on timelines and raise risks appropriately.
  • Maintain continuous communication with the customer, project SMEs, and key stakeholders to collect and document business requirements in support of their vision.
  • Review test scenarios and work with team to include any missed impact points.
  • Understand project delivery mechanisms and ensure owned stories/epics are tracked to closure.
  • Track customer requirements from inception through delivery.
  • Assist with user acceptance testing (UAT).
  • Analyze data to understand business problems and opportunities.
  • Identify and evaluate potential risks and impacts of proposed solutions.
  • Provide ongoing support and maintenance for implemented solutions.
  • Assist the product owner and development team to achieve customer satisfaction.

REQUIRED EXPERIENCE:

  • 7+ years of relevant experience
  • Strong Python development skills, including modular design, testing, debugging, packaging, dependency management, and performance optimization.
  • Hands-on experience developing data pipeline solutions with AWS Glue and integrating Glue with Amazon S3 and the AWS Glue Data Catalog.
  • Practical experience with multiple AWS data, integration, security, and monitoring services used to deliver end-to-end data pipelines.
  • Demonstrated experience ingesting, parsing, validating, transforming, and troubleshooting JSON and NDJSON, including nested structures, malformed records, schema drift, and large-file processing.
  • Experience with data modeling, schema design, data partitioning, metadata, lineage, and data quality practices.
  • Experience using Git-based version control, automated testing, CI/CD pipelines, and infrastructure-as-code approaches.
  • Excellent analytical, problem-solving, documentation, and communication skills, with a proactive and customer-focused approach.
  • Exhibit strong verbal and written communication skills, attention to detail, and the ability to follow up in a timely manner.
  • Have experience creating detailed reports and presenting information to both technical and non-technical audiences.
  • Possess expertise in using JIRA and Confluence for managing requirements.
  • Must be able to obtain and maintain a Public Trust clearance.
  • Must have lived in the United States 3 out of the past 5 years.

PREFERRED EXPERIENCE:

  • SAFe Agile Certification
  • AWS certification relevant to data engineering, architecture, or development.
  • Experience in healthcare IT and understanding of regulatory requirements such as HIPAA.
  • Experience designing source-to-target mappings, canonical data models, and data integration patterns across heterogeneous data providers.
  • Experience partnering with Data Quality and Data Governance teams to establish data quality metrics, validation rules, profiling processes, and remediation workflows.
  • Experience managing data delivery requirements, service level agreements (SLAs), and operational readiness processes.
  • Experience and/or knowledge of CMS programs, processes, and standards.
  • Experience with Apache Spark or PySpark, Parquet, Avro, Iceberg, or other distributed processing and open table or columnar storage technologies.

EDUCATION AND CERTIFICATIONS:

  • Bachelor's degree

WHAT WE OFFER: 

At Sparksoft, we know that people do their best work when they feel supported, inspired, and connected. That’s why we’ve built a workplace that balances comprehensive benefits with a culture of collaboration and innovation. From flexible time off to professional growth opportunities, we’re committed to helping you thrive both inside and outside of work. When you join Sparksoft, you’ll enjoy:

•    Competitive compensation and a 401(k) with employer contributions to help you plan for the future
•    Flexible paid time off and hybrid ways of working that support true work-life balance
•    Comprehensive health coverage—including medical, dental, vision, life, and disability insurance
•    A curated in-office experience designed to foster community, team connections, and innovation
•    Opportunities to give back through Sparksoft Cares, including annual company-wide fundraising events
•    Training and development programs that build new skills and prepare you for leadership roles
•    A collaborative, transparent, and fun culture—recognized as a Great Place to Work®

Accessibility and Accommodations: Sparksoft Corporation is committed to providing equal employment opportunities to all individuals. If you require accommodations during the application or interview process, please contact us at [email protected] or call 410-424-7700. Requests are reviewed and fulfilled on a case-by-case basis.

Security Notice: Your privacy and data security are important to us. Sparksoft Corporation will never request sensitive personal information via email. If you receive any suspicious communication claiming to be from Sparksoft, please report it immediately to our security team at [email protected].

Artificial Intelligence (AI) Policy: While Sparksoft values the appropriate use of artificial intelligence in the workplace, candidates must complete interviews and independent assessments using their own knowledge and abilities. Unless expressly permitted by Sparksoft, the use of AI during interviews and assessments is strictly prohibited. Violations may result in disqualification. Please contact your recruiter regarding accommodation needs.


Skills Required

  • 7+ years of relevant experience
  • Strong Python development skills, including modular design, testing, debugging, packaging, dependency management, and performance optimization
  • Hands-on experience developing data pipelines with AWS Glue, Amazon S3, and AWS Glue Data Catalog
  • Experience with multiple AWS data, integration, security, and monitoring services
  • Experience ingesting, parsing, validating, transforming, and troubleshooting JSON and NDJSON data
  • Experience with data modeling, schema design, data partitioning, metadata, lineage, and data quality
  • Experience with Git-based version control, automated testing, CI/CD, and infrastructure as code
  • Excellent analytical, problem-solving, documentation, and communication skills
  • Strong verbal and written communication skills, attention to detail, and timely follow-up
  • Experience creating detailed reports and presenting to technical and non-technical audiences
  • Expertise using JIRA and Confluence for requirements management
  • Ability to obtain and maintain a Public Trust clearance
  • Must have lived in the United States for 3 of the past 5 years
  • SAFe Agile Certification
  • AWS certification relevant to data engineering, architecture, or development
  • Experience in healthcare IT and knowledge of HIPAA requirements
  • Experience designing source-to-target mappings, canonical data models, and data integration patterns
  • Experience partnering with Data Quality and Data Governance teams
  • Experience managing data delivery requirements, SLAs, and operational readiness
  • Experience or knowledge of CMS programs, processes, and standards
  • Experience with Apache Spark or PySpark, Parquet, Avro, Iceberg, or similar distributed processing and storage technologies
  • Bachelor's degree
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Columbia, MD
162 Employees
Year Founded: 2004

What We Do

Sparksoft helps the clients achieve their business objectives by providing Innovative, best-of-breed software products and technology solutions at substantial cost savings. Sparksoft Team has considerable industry experience with wide range of leading companies.

Similar Jobs

Zapier Logo Zapier

Legal Counsel

Artificial Intelligence • Productivity • Software • Automation
Remote
2 Locations
800 Employees
158K-238K Annually

General Motors Logo General Motors

District Manager, OnStar & Loyalty - Alabama

Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Remote or Hybrid
United States
165000 Employees

DFIN Logo DFIN

Director Of Sales

Fintech • Software
Remote or Hybrid
United States
1750 Employees

Engine Logo Engine

Senior Software Engineer

Consumer Web • Software • Travel
Easy Apply
Remote
United States
1000 Employees
135K-187K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account