Application Development & Support Specialist

Position Overview

The AI/ML Engineer is a hands-on technical specialist responsible for designing, building, and operationalizing AI and machine learning capabilities within the MDM platform and its surrounding ecosystem. This role focuses on applying AI/ML to core MDM problems — entity matching, deduplication, data quality scoring, anomaly detection, and intelligent automation of data stewardship workflows.

This is not a research role — the AI/ML Engineer builds production-grade models and integrates them into the MDM platform's operational pipelines. The role operates under the "you build it, you support it" model, owning AI/ML features from experimentation through production deployment and ongoing monitoring.

The AI/ML Engineer works closely with the MDM Architect (who drives overall technical direction), the Engineering Manager (who drives delivery), and MDM/Integration developers (who build the platform and pipelines that AI/ML models plug into).

Key Responsibilities AI/ML Model Development & Integration

  • Design, build, and train ML models for entity matching, deduplication, and record linkage across customer, supplier, contact, and item domains
  • Develop probabilistic and deterministic matching algorithms that improve match accuracy over traditional rule-based approaches
  • Build data quality scoring models that assess completeness, accuracy, consistency, and timeliness of master data records
  • Develop anomaly detection models to identify data quality issues, unusual patterns, and potential duplicates in real-time data flows
  • Design and implement intelligent survivorship logic using ML to determine optimal golden record attribute values from multiple sources
  • Build NLP/text processing capabilities for entity name standardization, address parsing, and fuzzy matching
  • Integrate AI/ML models into MDM platform workflows — matching, merging, stewardship routing, and exception handling
  • Develop automated data classification and categorization models for incoming records
  • Build recommendation engines for data stewards — suggest merge candidates, flag potential false positives, prioritize review queues
  • Experiment with graph-based approaches for relationship discovery and network analysis across MDM entities

MLOps & Production Operations

  • Deploy ML models to production environments with proper versioning, monitoring, and rollback capabilities
  • Build and maintain ML pipelines for model training, validation, and deployment (e.g., MLflow, Kubeflow, SageMaker, Azure ML)
  • Implement model monitoring — track prediction accuracy, data drift, concept drift, and model degradation over time
  • Design A/B testing frameworks to compare model performance against rule-based baselines
  • Build automated retraining pipelines triggered by performance degradation or data distribution changes
  • Manage feature stores and feature engineering pipelines for MDM-specific features
  • Optimize model inference performance for real-time matching scenarios (latency, throughput)
  • Maintain model documentation including training data, hyperparameters, performance metrics, and decision thresholds

AI Platform & Agentic Workflows

  • Design and build agentic AI workflows integrated with MDM processes (e.g., using AWS Bedrock, LangChain, or similar frameworks)
  • Develop LLM-powered capabilities for data enrichment, entity extraction, and intelligent data validation
  • Build AI-assisted stewardship tools that reduce manual review effort through intelligent automation
  • Implement RAG (Retrieval-Augmented Generation) patterns for contextual data quality recommendations
  • Evaluate and integrate foundation models and LLMs for MDM-specific use cases
  • Design prompt engineering strategies and guardrails for production LLM integrations
  • Build conversational interfaces for data stewards to query and interact with MDM data using natural language

Data Engineering for AI/ML

  • Design and build feature engineering pipelines that extract ML-ready features from MDM, ERP, CRM, and data lake sources
  • Build training data pipelines — extract, label, and version training datasets from production MDM data
  • Implement data preprocessing, cleansing, and normalization pipelines specific to ML model inputs
  • Collaborate with data lake and integration teams to ensure AI/ML pipelines have access to required data sources
  • Design and maintain data schemas for ML feature stores and model input/output contracts

Production Support & Troubleshooting

  • Own production support for AI/ML features — monitor model performance, investigate prediction failures, and resolve issues within SLAs
  • Debug model prediction errors — trace through feature extraction, model inference, and post-processing to isolate root cause
  • Analyze model logs and prediction outputs to identify systematic errors or bias
  • Collaborate with MDM developers when AI/ML model outputs cause downstream data quality issues
  • Perform root cause analysis when match/merge accuracy degrades and implement corrective actions
  • Maintain runbooks for AI/ML model operations, retraining procedures, and incident response

Testing & Quality

  • Design and execute model evaluation frameworks — precision, recall, F1, AUC for matching models
  • Build automated test suites for model validation including edge cases, boundary conditions, and adversarial inputs
  • Conduct A/B testing and champion/challenger experiments to validate model improvements
  • Perform bias and fairness testing across different data segments and domains
  • Participate in code reviews for ML code and provide feedback on data pipeline quality
  • Maintain regression test datasets to ensure model updates don't degrade performance on known scenarios

Security & Compliance

  • Working knowledge of data security, information security practices, and SOX compliance for AI/ML model changes in production
  • Ensure PII/sensitive data handling compliance in model training and inference pipelines
  • Implement model explainability and audit trails for compliance-sensitive matching decisions

Required Qualifications AI/ML Expertise

  • 10 - 12 years hands-on experience building and deploying ML models to production (not just research/experimentation)
  • 4+ years experience with entity matching, record linkage, or deduplication using ML approaches (e.g., probabilistic matching, deep learning for entity resolution)
  • 4+ years experience with Python ML ecosystem — scikit-learn, pandas, NumPy, TensorFlow or PyTorch
  • 3+ years experience with NLP/text processing for entity name matching, address parsing, fuzzy matching (e.g., spaCy, NLTK, Hugging Face Transformers)
  • 3+ years experience with MLOps tools and practices — model versioning, deployment, monitoring (e.g., MLflow, Kubeflow, SageMaker, Azure ML, Vertex AI)
  • 2+ years experience with LLMs and generative AI — prompt engineering, RAG patterns, agentic workflows (e.g., AWS Bedrock, OpenAI API, LangChain, Claude)
  • Experience with graph-based ML or network analysis techniques (e.g., Neo4j, GraphSAGE, node2vec)
  • Experience with feature engineering and feature store management (e.g., Feast, Tecton, SageMaker Feature Store)

Data Engineering & Platform Skills

  • 5+ years advanced SQL experience — complex queries, performance tuning, data analysis (e.g., Oracle, SQL Server, PostgreSQL, Snowflake)
  • 4+ years hands-on coding in Python — production-quality code, not just notebooks (including testing, error handling, logging)
  • 3+ years experience building data pipelines for ML — feature extraction, training data preparation, model serving (e.g., Apache Spark, Airflow, dbt)
  • 2+ years experience with cloud ML platforms and services (e.g., AWS SageMaker, Azure ML, GCP Vertex AI)
  • 2+ years experience with containerization for model deployment (e.g., Docker, Kubernetes)
  • Experience with CI/CD for ML models — automated testing, deployment, and rollback (e.g., Jenkins, GitLab CI, GitHub Actions)
  • Experience with log analysis and monitoring tools for model observability (e.g., Splunk, ELK, CloudWatch, Prometheus/Grafana)

MDM & Data Quality Domain Knowledge

  • 3+ years experience working with MDM platforms or data quality systems (e.g., Informatica MDM, Reltio, Informatica DQ, Ataccama)
  • Understanding of MDM concepts: match/merge, survivorship, golden record, hierarchy management, stewardship workflows
  • Experience with data quality dimensions — completeness, accuracy, consistency, timeliness, uniqueness
  • Experience with incident management using ticketing systems (e.g., ServiceNow, Jira)
  • Working knowledge of SDLC: development, testing, CI/CD, change management, release management

Preferred Qualifications

  • Manufacturing industry experience — customer, supplier, item master data
  • Experience applying ML to multi-domain MDM (customer, supplier, contact, item)
  • Experience with Informatica MDM or Reltio platform internals and extensibility
  • Experience with graph databases for relationship modeling (e.g., Neo4j, Amazon Neptune)
  • Experience with real-time ML inference at scale (low-latency model serving)
  • Experience with data labeling and annotation workflows for training data creation
  • Experience with model explainability frameworks (e.g., SHAP, LIME)
  • Publications or patents in entity resolution, record linkage, or data quality
  • Familiarity with event-driven architecture and streaming ML (e.g., Kafka + ML inference)
  • Agile/Scrum delivery experience

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available