Data Engineer

Posted 5 Days Ago
Be an Early Applicant
Kondapur, Sangareddy, Telangana, IND
In-Office
Senior level
Cloud • Information Technology • Consulting
The Role
Build and operate reliable batch and streaming data pipelines, curated analytical data models, and data quality controls. Own production monitoring, incident response, performance and cost optimization, security, governance, lineage, and documentation. Partner with analysts, data scientists, product engineers, and customers to define durable data contracts and deliver Azure-based data platforms. Apply Python, PySpark, SQL, CI/CD, infrastructure as code, and modern lakehouse technologies.
Summary Generated by Built In

Coretek is looking for a Data Engineer to build and operate the pipelines and data models that the rest of the business runs on. You'll own ingestion from source systems through to curated, well-documented datasets that analysts, data scientists, and application teams depend on. This is a hands-on engineering role: you'll write production code, design schemas, and be accountable for the reliability and cost of what you ship.

Responsibilities

  • Design, build, and maintain batch and streaming data pipelines that are idempotent, observable, and recoverable.
  • Model data for analytics (dimensional models, semantic layers, and curated marts), balancing query performance against maintainability.
  • Integrate data from operational databases, SaaS APIs, files, and event streams, including handling schema drift and late-arriving data.
  • Build data quality checks (freshness, volume, uniqueness, referential integrity) into pipelines rather than bolting them on afterward, and define how failures alert and escalate.
  • Own pipelines in production: monitoring, on-call rotation for data incidents, root-cause analysis, and backfills.
  • Tune performance and cost (partitioning, clustering, file sizing, warehouse and cluster sizing) and make the tradeoffs explicit.
  • Apply engineering discipline to data: version control, code review, CI/CD, automated testing, and infrastructure as code.
  • Implement access controls, PII handling, retention, and lineage and audit requirements in partnership with security and compliance.
  • Partner with analysts, data scientists, and product engineers to turn ambiguous requirements into durable data contracts.
  • Maintain data dictionaries, lineage, and pipeline runbooks so consumers can find a dataset, understand what each field means and how current it is, and use it correctly without having to ask the team that built it.

Requirements
  • 5+ years building production data pipelines.
  • Strong hands-on Python development for data engineering, with real testing, packaging, and code review practice, not scripting alone.
  • Working knowledge of PySpark: DataFrame and SQL APIs, joins and aggregations at scale, partitioning and shuffle behavior, and the ability to read a Spark UI to diagnose a slow or failing job.
  • Strong SQL: window functions, query plans, and performance tuning, not just SELECTs.
  • Hands-on experience with the Azure data platform: Data Factory, Databricks, Synapse/Fabric, and ADLS.
  • Solid data modeling fundamentals: normalization, star schemas, slowly changing dimensions.
  • Git-based workflow and experience shipping through CI/CD.
  • Excellent communication skills, with the ability to debug a failing pipeline end to end and articulate the impact to diverse audiences, including non-technical stakeholders.
  • Exceptional analytical and problem-solving skills, with the judgment to find the root cause of a data issue rather than patching the symptom.
  • Strong knowledge and experience in working with customers in a consultative approach in a technical environment.

Additional Qualifications

  • Streaming experience (Kafka, Event Hubs).
  • Lakehouse formats: Delta Lake, Iceberg.
  • Infrastructure as code (Terraform, Bicep) and containerization (Docker, Kubernetes).
  • Experience in a regulated environment (HIPAA, SOC 2, PCI, GDPR): auditability, encryption, data residency.
  • Experience building data platforms for ML or supporting feature pipelines.
  • Proven ability to manage multiple client projects and deliver high-quality results on time.
  • Experience in Azure DevOps or GitHub for source control and pipelines.

Skills Required

  • 5+ years building production data pipelines
  • Strong hands-on Python development for data engineering, including testing, packaging, and code review
  • Working knowledge of PySpark DataFrame and SQL APIs, joins, aggregations, partitioning, shuffle behavior, and Spark UI diagnosis
  • Strong SQL skills, including window functions, query plans, and performance tuning
  • Hands-on experience with Azure Data Factory, Databricks, Synapse or Fabric, and ADLS
  • Data modeling fundamentals, including normalization, star schemas, and slowly changing dimensions
  • Git-based workflow and CI/CD delivery experience
  • Excellent communication and end-to-end pipeline troubleshooting skills
  • Strong analytical and root-cause problem-solving skills
  • Experience working with customers in a consultative technical environment
  • Streaming experience with Kafka or Event Hubs
  • Experience with Delta Lake or Iceberg
  • Infrastructure as code using Terraform or Bicep
  • Containerization experience with Docker or Kubernetes
  • Experience in regulated environments such as HIPAA, SOC 2, PCI, or GDPR
  • Experience building ML data platforms or feature pipelines
  • Ability to manage multiple client projects and deliver results on time
  • Experience with Azure DevOps or GitHub for source control and pipelines
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Farmington Hills, MI
147 Employees
Year Founded: 2005

What We Do

Coretek is the #1 Microsoft Azure Partner in the U.S. and an Azure Expert Managed Service Provider. Coretek consults, builds, manages, and maintains IT infrastructure, enabling business leaders to spend less time thinking about technology and more time focused on their customers, culture, and communities. Coretek solves the world's most complex business challenges with the cloud.

Similar Jobs

Optum Logo Optum

Data Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

CrowdStrike Logo CrowdStrike

Data Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
India
11000 Employees

Optum Logo Optum

Data Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees
Hybrid
Hyderabad, Telangana, IND
289097 Employees

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account