Advanced Data Scientist

Posted 21 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Aerospace
The Role
Develops advanced multimodal machine learning solutions spanning text, audio, and vision. Responsibilities include mathematically formulating business problems, engineering production-grade code, building scalable big-data pipelines, training deep learning models, implementing NLP and computer vision systems, developing OCR and document AI solutions, and deploying machine learning applications on cloud platforms. Requires strong foundations in mathematics, machine learning, deep learning, distributed computing, and 4–6 years of industry experience.
Summary Generated by Built In

Key Responsibilities

  • Mathematical Formulation: Translate ambiguous business problems into mathematically sound framework objectives and optimisation targets.
  • Production-Grade Engineering: Write clean, modular, and maintainable code using production-level design patterns to scale mathematical models.
  • Big Data Processing: Design and manage scalable data pipelines to process massive datasets efficiently for model training and inference.
  • Deep Learning & Vision Development: Build, train, and fine-tune complex neural networks across text, audio, and visual modalities.
  • Cloud Deployment: Architect and deploy models to cloud environments, leveraging distributed computing and robust cloud infrastructure.

Required Technical Skills & Competencies


1. Tooling, Libraries & Software Engineering

  • Core Language: Advanced proficiency in Python with a strict adherence to Object-Oriented Programming (OOP) principles, clean coding standards, and design patterns.
  • Machine Learning Libraries: Advanced proficiency in scikit-learn (sklearn) for data preprocessing, feature engineering, and baseline modelling.
  • Deep Learning Frameworks: Core expertise in PyTorch (preferred) or TensorFlow for building, customizing, and training deep neural networks from scratch.
  • Big Data Ecosystem: Experience with Apache Spark (PySpark) and the Hadoop Ecosystem (HDFS, Hive, MapReduce) for handling, transforming, and querying large-scale distributed datasets.
  • Cloud Architecture: Experience building and deploying scalable machine learning applications on major cloud platforms (AWS, Azure, or GCP).

2. Core Mathematics & First-Principles ML

  • Foundational Math: Solid foundation in Linear Algebra (eigenvalues, SVD, matrix decompositions), Multivariable Calculus (partial derivatives, gradients, Jacobians), and Probability Theory (Bayesian inference, probability distributions, expectation maximization).
  • Machine Learning: In-depth understanding of standard Machine Learning algorithms (Trees, Boosting, SVMs, GMMs) with the ability to explain the underlying loss functions and optimizations mathematically.
  • Deep Foundations: Thorough understanding of Multi-Layer Perceptrons (MLPs), mathematical derivation of backpropagation, hyperparameter initialization strategies (Xavier, He), optimization variants (Adam, RMSProp), and advanced regularization techniques (L1/L2, Dropout, Batch Normalization). 

3. Advanced Natural Language Processing (NLP)

  • Sequential Networks: Hands-on experience with sequence modeling, including Word Embeddings (Word2Vec, FastText), RNNs, LSTMs, and GRUs.
  • Transformer Ecosystem: Deep structural knowledge of the Transformer architecture (Self-Attention math, Multi-Head mechanisms).
  • Pre-trained NLP Models: Experience implementing and fine-tuning encoder-only (BERT, RoBERTa) and decoder-only (GPT series) architectures.

4. Computer Vision (CV) & Document AI


  • Spatial Networks: Deep understanding of Convolutional Neural Networks (CNNs), feature map mathematics, pooling operations, and advanced CV backbones.
  • OCR & Document Processing: Proven track record building or customizing Optical Character Recognition (OCR) systems for complex text extraction pipelines.
  • Vision Transformers: Familiarity with the adaptation of attention mechanics to visual tasks (ViTs, Swin Transformers).

Education & Experience

Qualifications
  • Education: Bachelor’s, Master's, or Ph.D. in a highly quantitative field (Mathematics, Statistics, Econometrics, Computer Science, Physics, or Operations Research).
  • Experience: 4 to 6 years of industry experience working as a Data Scientist or Machine Learning Engineer with a portfolio of complex multimodal projects.
About UsHoneywell Technologies is a global, pure-play automation company with a legacy of innovating to help solve the world’s most mission-critical challenges, enhancing the quality of life for people and communities around the world. We serve the building, industrial and process sectors with a broad portfolio of services, solutions and products, underpinned by our Honeywell Technologies Accelerator operating system and Honeywell Technologies Forge intelligence layer. By combining the deep domain expertise of our more than 50,000 employees with decades of data from our global installed base, we are uniquely positioned to lead the industrial sector’s transition from automation to autonomy.

Skills Required

  • Advanced proficiency in Python, object-oriented programming, clean coding standards, and design patterns
  • Advanced proficiency with scikit-learn for preprocessing, feature engineering, and baseline modeling
  • Expertise with PyTorch or TensorFlow for developing and training deep neural networks
  • Experience with Apache Spark or PySpark and the Hadoop ecosystem, including HDFS, Hive, and MapReduce
  • Experience deploying scalable machine learning applications on AWS, Azure, or GCP
  • Strong foundation in linear algebra, multivariable calculus, and probability theory
  • In-depth understanding of standard machine learning algorithms, loss functions, and optimization
  • Thorough understanding of deep learning architectures, backpropagation, initialization, optimization, and regularization
  • Hands-on experience with Word2Vec, FastText, RNNs, LSTMs, and GRUs
  • Deep knowledge of Transformer architecture, self-attention, and multi-head attention
  • Experience implementing and fine-tuning BERT, RoBERTa, and GPT-series models
  • Deep understanding of CNNs, feature maps, pooling, and advanced computer vision backbones
  • Proven experience building or customizing OCR systems and complex text extraction pipelines
  • Familiarity with Vision Transformers and Swin Transformers
  • Bachelor's, Master's, or Ph.D. in Mathematics, Statistics, Econometrics, Computer Science, Physics, Operations Research, or another highly quantitative field
  • 4 to 6 years of industry experience as a Data Scientist or Machine Learning Engineer
  • Portfolio of complex multimodal machine learning projects
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Mississauga, Ontario
10,000 Employees
Year Founded: 1914

Similar Jobs

Honeywell Logo Honeywell

Data Scientist

Aerospace • Security • Energy • Industrial
In-Office
2 Locations
110269 Employees
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
10000 Employees

Honeywell Logo Honeywell

Data Scientist

Aerospace • Security • Energy • Industrial
In-Office
2 Locations
110269 Employees

Honeywell Logo Honeywell

Data Scientist

Aerospace • Security • Energy • Industrial
Hybrid
4 Locations
110269 Employees

Similar Companies Hiring

Rangeview Thumbnail
Manufacturing • Defense • Aerospace
Berkeley, CA
25 Employees
Outpost Space Thumbnail
Aerospace • Defense
US
24 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account