Are you passionate about building production-grade AI systems that continuously learn and improve from real-world feedback? We are looking for a Senior ML Engineer / Data Scientist to help develop intelligent recognition and entity-matching solutions for a large-scale media data platform.
In this fully remote role across Europe, you will work with Large Language Models, evaluation frameworks, and cloud-based ML pipelines to improve automation quality and reduce manual processing efforts. You will collaborate closely with Data Engineering teams and Customer stakeholders while owning the ML lifecycle end-to-end.
We at Sigma Software create impactful technology solutions for global customers and provide engineers with opportunities to work on meaningful, high-scale products using modern AI technologies. This role offers significant ownership, challenging engineering tasks, and the ability to influence production AI systems at scale.
CUSTOMER
Our Customer operates a large-scale platform focused on processing and structuring advertising and media operational data. The company is actively investing in intelligent automation and machine learning solutions to improve recognition accuracy across multiple station and network layouts while minimizing manual intervention in data processing workflows.
PROJECT
The project focuses on building a self-learning Postlog and Prelog recognition system capable of automatically understanding new layouts, extracting structured data, and improving from production feedback. The solution leverages Large Language Models and modern ML practices to optimize recognition quality, entity matching, and confidence-based automation.
You will contribute to the development of scalable AI-driven workflows designed to achieve high automation accuracy, observability, and operational efficiency in production environments.
Job Description- Design and develop a self-learning Postlog and Prelog recognition system using modern ML and LLM techniques
- Build and maintain versioned prompts, evaluation datasets, and few-shot exemplars
- Apply production-grade LLM practices including schema-constrained extraction, grounding strategies, and low-confidence fallback handling
- Improve recognition quality and optimize layout and header mapping performance
- Analyze production failures and enhance prompts, retrieval pipelines, and model behavior
- Run evaluation pipelines and shadow-mode comparisons against legacy systems and gold datasets
- Monitor confidence scores, latency, operational quality, and infrastructure costs
- Develop entity-matching systems for Station, Advertiser, and CreativeID master data
- Implement confidence scoring, thresholding, and auditability mechanisms
- Transform human and machine corrections into labeled signals for continuous model improvement
- Monitor prompt and model drift in production environments
- Collaborate with Data Engineering teams on ML integration and operationalization
- Communicate technical findings and recommendations to engineering teams and Customer stakeholders
- At least 5 years of experience in Machine Learning, Data Science, or ML Engineering
- Proven experience delivering ML models or LLM-powered systems into production
- Strong hands-on experience with Large Language Models in real products or pipelines
- Deep understanding of prompt engineering, prompt versioning, evaluation methodologies, and grounding strategies
- Experience handling low-confidence scenarios and optimizing cost and latency for LLM systems
- Strong Python and SQL skills
- Solid knowledge of statistics, confidence estimation, sampling, hypothesis testing, and threshold optimization
- Experience with classification, ranking, matching, or recommendation-related problems
- Understanding of offline evaluation metrics, holdout validation, and production monitoring
- Hands-on experience with AWS cloud services including S3, IAM, CloudWatch, and orchestration services
- Strong communication and collaboration skills
- Upper-Intermediate English level or higher
WILL BE A PLUS
- LLM-related certifications
- Experience with Amazon Bedrock or equivalent enterprise LLM platforms
- Production experience with Claude/Sonnet-class models
- Experience with Excel or layout extraction systems
- Knowledge of confidence calibration, active learning, or weak supervision techniques
- Experience with cost-aware LLM operations including caching, routing, and fallback models
- Advertising or media domain knowledge
- Familiarity with Glue, Airflow, or similar orchestration and data pipeline tools
PERSONAL PROFILE
- Strong ownership mindset
- Analytical and data-driven thinking
- Ability to work independently in ambiguous environments
- Continuous improvement approach
- Attention to quality and operational excellence
- Effective collaboration and communication skills
Skills Required
- At least 5 years of experience in machine learning, data science, or ML engineering
- Experience delivering machine learning models or LLM-powered systems into production
- Hands-on experience with Large Language Models in real products or pipelines
- Knowledge of prompt engineering, prompt versioning, evaluation methodologies, and grounding strategies
- Experience handling low-confidence scenarios and optimizing cost and latency for LLM systems
- Strong Python skills
- Strong SQL skills
- Knowledge of statistics, confidence estimation, sampling, hypothesis testing, and threshold optimization
- Experience with classification, ranking, matching, or recommendation problems
- Understanding of offline evaluation metrics, holdout validation, and production monitoring
- Hands-on experience with AWS services including S3, IAM, CloudWatch, and orchestration services
- Strong communication and collaboration skills
- Upper-Intermediate English level or higher
- LLM-related certifications
- Experience with Amazon Bedrock or equivalent enterprise LLM platforms
- Production experience with Claude or Sonnet-class models
- Experience with Excel or layout extraction systems
- Knowledge of confidence calibration, active learning, or weak supervision techniques
- Experience with cost-aware LLM operations, including caching, routing, and fallback models
- Advertising or media domain knowledge
- Familiarity with AWS Glue, Airflow, or similar orchestration and data pipeline tools
What We Do
Sigma Software Group, an award-winning and trusted IT partner, has been serving customers for over 21 years, providing comprehensive IT solutions to various businesses, ranging from startups to established software product houses. As one of Europe's substantial IT consultancies, it brings together a dedicated workforce of over 2,100 professionals in 40 offices across 19 countries. With a diverse client base, including more than 300 enterprises, including Fortune 500 stalwarts, Sigma Software Group is a preferred choice for developing solutions that help businesses create cutting-edge products while meeting their unique needs. Sigma Software Group operates as a dynamic ecosystem of tech companies, offering 25 ready-to-implement innovative products and 40+ value-added services. Furthermore, Sigma Software Group is committed to fostering innovation through initiatives such as the Sigma Software Labs business incubator, Sigma Software University, the SID Venture Partners VC Fund, UA Tech Network, Techosystem, the European Business Association, and other collaborative efforts. Since 2015, Sigma Software Group has consistently earned recognition on the IAOP's prestigious World's Top 100 Outsourcing list. The company's accomplishments have also been acknowledged by prominent global media outlets such as Forbes, CNBC, The Times, and Reuters







