Data Engineer

Posted Yesterday
Be an Early Applicant
Hyderabad, Telangana, IND
In-Office
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
The Role
Build and operate scalable data pipelines, warehouses, models, quality controls, and curated analytics datasets. Manage cloud data infrastructure and troubleshoot data reliability issues. Own the production lifecycle of a computer-vision cart classifier, including data workflows, feature engineering, evaluation, deployment, monitoring, and rollback. Collaborate with BI, product, software engineering, operations, and machine-learning teams.
Summary Generated by Built In
About Beyond Key:

We are a Microsoft Gold Partner and a Great Place to Work-certified company. "Happy Team Members, Happy Clients" is a principle we hold dear. We are an international IT consulting and software services firm committed to providing. Cutting-edge services and products that satisfy our clients' global needs. Our company was established in 2005, and since then we've expanded our team by including more than 350+ Talented skilled software professionals. Our clients come from the United States, Canada, Europe, Australia, the Middle East, and India, and we create and design IT solutions for them. If you need any more details, you can get them at https://www.beyondkey.com/about.

Role Summary:
This role bridges the gap between core data engineering and practical machine learning applications. Primarily, you will be a data platform engineer responsible for owning core data pipelines, data models, and quality controls that power client analytics and future data products.
Secondarily, you will drive the production lifecycle of our shopping cart computer-vision feature. You will orchestrate the data workflows that interface with our Machine Learned models to ensure accurate cart classification, while leveraging the FaceFirst ML team for deeper capacity. You will collaborate with BI Analysts, software engineers, and product teams to transform raw data into actionable insights.

Key Responsibilities
Data Platform, Pipelines, & Quality (Primary Focus)
  • Pipeline Design & Operation: Design, build, and operate scalable ELT/ETL pipelines that ingest data from IoT/smart-cart telemetry, video events, operational systems, and external partners into our cloud data lake/warehouse.
  • Infrastructure Management: Build and maintain robust data infrastructure, including databases (SQL and NoSQL), data warehouses, and data integration solutions.
  • Data Modeling: Establish canonical data models and definitions (schemas, event taxonomy, metrics) so teams can trust and reuse the same data across products, BI, and analytics.
  • Data Quality Assurance: Own data quality end-to-end by implementing validation rules, automated tests, anomaly detection, and monitoring/alerting to prevent and quickly detect regressions.
  • Consistency & Governance: Drive data consistency improvements across systems (naming, identifiers, timestamps, joins, deduplication) and document data contract expectations with producing teams.
  • Root Cause Analysis: Troubleshoot pipeline and data issues, perform root-cause analysis, and implement durable fixes that improve reliability and reduce operational load.
  • Collaboration & Analytics: Partner with BI Analysts and Product teams to create curated datasets and self-serve analytics foundations (e.g., marts/semantic layer), as well as support internally facing dashboards to communicate system health.
Applied ML Ownership - Smart Exit Cart-Empty Classifier (Secondary Focus)
  • Lifecycle Management: Own the production lifecycle for the cart classification capability, including data collection/labeling workflows, evaluation, threshold tuning, and safe release/rollback processes.
  • Pipeline Implementation: Implement and optimize machine learning pipelines, from feature engineering and model training to deployment and monitoring in production.
  • Evaluation & Monitoring: Build and maintain an evaluation harness (offline metrics + repeatable test sets) and ongoing monitoring (accuracy drift, data drift, false positive/negative analysis).
  • Cross-Team Collaboration: Collaborate with the FaceFirst ML team to incorporate improvements (model updates, feature changes) while keeping client’s production integration stable.
  • Integration: Work with software engineers to ensure the classifier integrates cleanly into the product workflow with robust telemetry, logging, and operational runbooks.
Skills & Experience Required:
Work Experience: 5+ years

  • Core Engineering: Strong experience building and operating production ELT/ETL pipelines and data warehouses.
  • Programming: Fluency in SQL and Python (or similar) for data transformation, validation, and automation.
  • Cloud Platforms: Experience with cloud data platforms (Azure and/or GCP), including object storage, security/access controls, and cost-aware design.
  • Tooling: Hands-on experience with orchestration and transformation tooling (e.g., Airflow/Prefect) and batch processing frameworks (e.g., Spark/Databricks).
  • Quality Practices: Practical experience implementing data quality practices (tests, monitoring/alerting, lineage/documentation) and improving data consistency across systems.
  • Operations: Collaborate with operational teams to identify, diagnose, and remediate in-field system issues.
  • Bachelor’s degree in computer science, Software Engineering, Information Systems, Mathematics, Statistics, or a related technical field.
Nice to Have
  • Rule Engines: Working knowledge and understanding of Rule Engines and the integration of those engines to process complex data relationships.
  • ML & Vision: Experience with computer vision/video analytics concepts or deploying ML models to production.
  • Streaming & IoT: Experience with streaming/near real-time data (e.g., Kafka, Pub/Sub), IoT telemetry pipelines, or edge computing.
  • MLOps: Familiarity with MLOps practices (model/version tracking, reproducible training, monitoring).

Skills Required

  • 5+ years of professional work experience
  • Experience building and operating production ELT/ETL pipelines and data warehouses
  • Fluency in SQL and Python or a similar programming language
  • Experience with Azure and/or GCP cloud data platforms, including object storage, security/access controls, and cost-aware design
  • Hands-on experience with orchestration and transformation tools such as Airflow or Prefect
  • Experience with batch processing frameworks such as Spark or Databricks
  • Experience implementing data quality tests, monitoring, alerting, lineage, documentation, and data consistency improvements
  • Ability to collaborate with operational teams to diagnose and remediate in-field system issues
  • Bachelor's degree in computer science, software engineering, information systems, mathematics, statistics, or a related technical field
  • Working knowledge of rule engines and their integration for processing complex data relationships
  • Experience with computer vision, video analytics, or deploying machine-learning models to production
  • Experience with streaming or near-real-time data, Kafka, Pub/Sub, IoT telemetry pipelines, or edge computing
  • Familiarity with MLOps practices, including model/version tracking, reproducible training, and monitoring
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
379 Employees
Year Founded: 2005

What We Do

Beyond Key is a global IT services and consulting company that helps organizations modernize and transform through Microsoft 365, Dynamics 365, cloud, business intelligence, data, artificial intelligence, automation, and custom software development. It delivers enterprise web, mobile, and cloud applications, integrations, and consulting, serving clients across the United States, Canada, Europe, Australia, and other markets as a long-term technology partner.

Similar Jobs

Optum Logo Optum

Data Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

MetLife Logo MetLife

Data Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Hyderabad, Telangana, IND
43000 Employees
Hybrid
Hyderabad, Telangana, IND
289097 Employees

Capco Logo Capco

Data Engineer

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account