Data Engineer

Role summary

We are looking for a hands-on data engineer with strong ETL and PySpark expertise to design, build, and support data pipelines and data marts within a banking environment. The ideal candidate will own the full SDLC lifecycle — from build through UAT, production deployment, and post-production support — while working across structured, semi-structured, and unstructured data.

Key responsibilities

  • Design, develop, and maintain ETL pipelines and data marts using PySpark and Python
  • Write clean, maintainable, and production-grade Python code following software engineering best practices
  • Own end-to-end SDLC activities: build, UAT support, UAT bug fixes, production deployment, and post-production support
  • Perform data analysis and debugging using Oracle SQL and PySpark
  • Work across structured, semi-structured, and unstructured data sources
  • Build and maintain data warehousing solutions supporting banking/financial reporting needs
  • Debug and optimize PySpark jobs for performance and reliability
  • Collaborate with cross-functional teams (QA, DBAs, business analysts) through the release cycle
  • Participate in CI/CD pipeline processes, including testing and validation of data pipelines
  • Ensure data pipeline reliability, scalability, and adherence to banking data governance/compliance standards

Required skills & experience

  • 5+ years of commercial experience in a data-driven engineering role
  • Hands-on experience building data marts and ETL pipelines
  • Expert-level PySpark and Python for ETL scripting
  • Strong command of Oracle SQL for data analysis and debugging
  • Proven experience across the full SDLC — build, UAT, bug fixing, deployment, post-prod support
  • Strong understanding of software engineering concepts and best practices for production pipelines
  • Experience working with structured, semi-structured, and unstructured data
  • Prior experience with banking clients or strong banking domain knowledge
  • Strong data warehousing fundamentals

Tech stack (daily use)

  • Languages: Python
  • Big Data: Spark / PySpark, Hadoop, MapReduce, Hive
  • Data libraries: Pandas
  • Databases: SQL and NoSQL DBMS
  • Tools: Jupyter
  • Practices: CI/CD, data testing & validation

Nice to have (optional — add if applicable)

  • Cloud experience (AWS/Azure/GCP) — not mentioned in your input, confirm with client
  • Airflow or other orchestration tools
  • Experience with regulatory/compliance reporting in banking

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available