- Analyze production issues and perform root cause analysis
- Define and implement data quality frameworks metrics and controls
- Develop and validate data pipelines with PySpark and Databricks
- Execute source to target reconciliation
- Identify investigate and resolve data anomalies
- Implement proactive monitoring and automated quality checks
- Monitor data across Databricks Lakehouse
- Optimize data processing jobs for performance and scalability
- Perform end to end data pipeline validation
- Run data quality checks and monitoring
- Track and manage data defects through resolution
- Validate business rules and transformation logic against source systems
- Write advanced SQL for data profiling and root cause analysis