Data Engineer- Senior Associate
Industry/Sector
Not ApplicableSpecialism
Software EngineeringManagement Level
Senior AssociateJob Description & Summary
The Opportunity
Join our Acceleration Center India and help shape the future of business for our diverse client portfolio across geographies and jurisdictions. You’ll work at the heart of global teams across Advisory, Assurance, Tax and Business Services—solving real client challenges through connected collaboration. We’ll help you grow your skills so you can go further. With hands-on learning, cutting-edge tools and an inclusive culture, this is your opportunity to do inspiring work that makes a difference—every day.
As a Data Engineer- Senior Associate, you will focus on designing and building data infrastructure and systems to enable efficient data processing and analysis. Within our Internal Firm Services practice, you will transform raw data into actionable insights, enabling informed decision-making and driving business growth. As a Senior Associate, you will build meaningful client connections and learn how to manage and inspire others. You will navigate increasingly complex situations, growing your personal brand and technical skills. You are expected to anticipate the needs of your teams and clients, delivering quality work. Embracing increased ambiguity, you will be comfortable when the path forward isn’t clear, using these moments as opportunities to grow.
In this role at PwC Acceleration Center India, you will develop and implement data pipelines, data integration, and data transformation solutions. You will use a broad range of tools and methodologies to generate new ideas and solve problems. By interpreting data, you will inform insights and recommendations, upholding professional and technical standards. This position offers a chance to deepen your understanding of the business context and how it is evolving, while building relationships and enhancing your skills.
Responsibilities
- Designing and developing data infrastructure and systems to facilitate efficient data processing and analysis
- Implementing data pipelines, integration, and transformation solutions to support client needs
- Leveraging advanced technologies and techniques to convert raw data into actionable insights
- Collaborating with internal stakeholders to understand project objectives and align data solutions with business strategies
Design and develop reliable ETL/ELT pipelines using SQL, Python, PySpark, and Databricks.
- Build and maintain data models across Oracle, relational databases, MongoDB, and lakehouse platforms.
- Develop and optimize complex SQL queries, stored procedures, views, indexes, and database objects.
- Administer databases, including user access, security, backup and recovery, monitoring, patching, and performance tuning.
- Build Databricks workflows using Spark SQL, Delta Lake, and medallion data layers.
- Prepare structured and unstructured data for analytics, machine learning, and Generative AI applications.
- Support AI use cases involving embeddings, vectorization, semantic search, vector databases, and Retrieval-Augmented Generation (RAG).
- Implement data-quality checks, metadata management, security controls, and data-governance standards.
- Troubleshoot production database and data-pipeline issues and perform root-cause analysis.
- Collaborate with data architects, application teams, analysts, data scientists, and AI engineers.
- Mentor junior team members and contribute to data engineering and database best practices.
What You Must Have
- Minimum bachelor degree
- At least 4-9 years of experience
- Oral and written proficiency in English required
What Sets You Apart
-Strong hands-on experience with Oracle Database, SQL, and PL/SQL.
- Experience with relational database design, query optimization, indexing, partitioning, and performance tuning.
- Working knowledge of MongoDB, including document modeling, aggregation, indexing, and administration.
- Experience with Databricks, Apache Spark, Spark SQL, PySpark, and Delta Lake concepts.
- Experience building batch or near-real-time data pipelines.
-Knowledge of database administration, including backup, recovery, security, monitoring, migration, and high availability.
- Understanding of AI-ready data pipelines, embeddings, vector search, similarity matching, and vector databases.
- Proficiency in Python or another data-engineering scripting language.
- Strong analytical, troubleshooting, communication, and documentation skills.
Strong hands-on experience with Oracle Database, SQL, and PL/SQL.
- Experience with relational database design, query optimization, indexing, partitioning, and performance tuning.
- Working knowledge of MongoDB, including document modeling, aggregation, indexing, and administration.
- Experience with Databricks, Apache Spark, Spark SQL, PySpark, and Delta Lake concepts.
- Experience building batch or near-real-time data pipelines.
- Knowledge of database administration, including backup, recovery, security, monitoring, migration, and high availability.
- Understanding of AI-ready data pipelines, embeddings, vector search, similarity matching, and vector databases.
- Proficiency in Python or another data-engineering scripting language.
- Strong analytical, troubleshooting, communication, and documentation skills.
Travel Requirements
Not SpecifiedJob Posting End Date