The Role
Build predictive models, end-to-end data science pipelines, and cloud-deployed AI/ML APIs. Develop image classification, NLP, chatbot, RAG, and LLM-agent solutions while monitoring data quality, drift, governance, performance, latency, and scalability. Collaborate with cross-functional teams to deliver production machine-learning projects and translate business requirements into technical solutions.
Summary Generated by Built In
Data Scientist
Pune
About Us
Coditation Systems was founded by a serial entrepreneur, and a team of young talented technologists, some of who have grown to spearhead the organization. With its inception in 2016, we became a boutique technology services and solutions firm specializing in Machine Learning & AI, Data Engineering, and Cloud. We have a team of ninja architects, data scientists, data engineers, and software engineers having decades of collective experience of applying emerging technologies to build cutting edge software products.
What are we looking for?
We are looking for a skilled Databricks Developer with hands-on experience in building and optimizing data pipelines on the Databricks platform. The ideal candidate should have strong expertise in big data processing, Spark, and cloud data engineering, with the ability to work in a fast-paced, data-driven environment.
A Day in the Life
- Build and optimize predictive models for churn forecasting, revenue contraction, and risk analysis at customer-product levels.
- Develop end-to-end data science libraries and orchestrate pipelines using Kedro, managing data loading, feature engineering, model training, inference, and explainability.
- Deploy AI/ML solutions on cloud platforms (GCP / AWS) and expose models as APIs.
- Collaborate with cross-functional teams to translate business requirements into technical solutions.
- Monitor data quality, detect drifts, and implement data governance frameworks.
- Perform image classification, text moderation, and other AI/ML-based solutions for client products.
- Design, develop, and deploy LLM-based AI agents for business intelligence, automated workflows, and decision support.
What you will need
- 7+ years of experience in Data data science
- Programming Languages: Python, SQL
- Frameworks & Libraries: scikit-learn, PyTorch, XGBoost, LightGBM, pandas, NumPy, Polars, LangChain, LlamaIndex, Crew.ai
- Databases: Postgres (pgvector, tsvector), Redshift, Snowflake (and/or any other vector db)
- Strong understanding of data science workflows, including data cleaning, feature engineering, model selection, validation, and explainability (e.g., SHAP).
- Experience in developing chatbots or automated query systems.
- Experience in image classification and transformer-based NLP models.
- Good to have: Cloud & Platforms: GCP, API deployment using Flask / FastAPI
- Good to have: Experience with RAG, LLM prompt optimisation, and GraphLLMs to reduce hallucinations.
Competencies
- Proven track record of end-to-end ML project delivery, from raw data to production.
- Strong analytical skills with experience in clustering, forecasting, and predictive modeling.
- Experience working with subscription and contractual business models.
- Ability to optimize LLMs and ML models for performance, latency, and scalability.
- Excellent problem-solving and communication skills.
Skills Required
- 7+ years of experience in data science
- Python programming
- SQL programming
- Experience with Databricks, big data processing, Spark, and cloud data engineering
- Experience with scikit-learn, PyTorch, XGBoost, LightGBM, pandas, NumPy, Polars, LangChain, LlamaIndex, and CrewAI
- Experience with Postgres, Redshift, Snowflake, and/or vector databases
- Understanding of data cleaning, feature engineering, model selection, validation, and explainability such as SHAP
- Experience developing chatbots or automated query systems
- Experience with image classification and transformer-based NLP models
- Strong analytical skills in clustering, forecasting, and predictive modeling
- Experience with end-to-end ML project delivery from raw data to production
- Experience with subscription and contractual business models
- Ability to optimize LLMs and ML models for performance, latency, and scalability
- Excellent problem-solving and communication skills
- Experience with GCP
- API deployment using Flask or FastAPI
- Experience with RAG, LLM prompt optimization, and GraphLLMs
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
is a boutique technology services and solutions firm specializing in Machine Learning & AI, Data Engineering, and Cloud. It has a team of ninja architects, data scientists, data engineers, and software engineers having a decades of collective experience of applying emerging technologies to build cutting edge software products









