Data Engineer
Tebra only initiates contact with candidates via email from an official Tebra email address (@, @, or @) or through our applicant tracking system, Greenhouse. We will only ask you to provide sensitive personal information through our official application portal — not via social media or text message. We do not conduct interviews via instant messaging.
About the RoleAs a Data Engineer focused on AI/ML, you’ll build, maintain, and optimize the data infrastructure that powers Tebra’s intelligent features. You’ll partner closely with Machine Learning Engineers, Data Scientists, and Software Engineers to transform complex healthcare data into high-quality datasets and real-time features that enable machine learning models.
This is a hands-on engineering role where you’ll contribute to scalable data pipelines, improve data quality, and help ensure our AI systems are powered by reliable, performant, and well-governed data. You’ll work on modern data platforms and gain experience building solutions that support both model training and production inference.
Your Area of Focus- Design, build, and maintain scalable data pipelines for feature extraction, training data generation, and model monitoring.
- Develop and enhance data systems that support analytics and machine learning workloads, including data lakehouse and feature store technologies.
- Monitor production data pipelines, identify data quality issues or pipeline failures, and implement improvements to ensure reliability and freshness.
- Participate in engineering design discussions and contribute to technical decisions around data architecture and pipeline implementation.
- Build reusable data engineering components, including automated data quality checks, schema validation, and testing frameworks.
- Translate business requirements into scalable data solutions that enable analytics and machine learning use cases.
- Optimize SQL queries, Spark workloads, and data processing pipelines to improve performance and scalability.
- Collaborate with ML Engineers and cross-functional partners to support MLOps best practices, including data versioning, lineage, and reproducibility.
- Break down technical work into manageable tasks and deliver high-quality solutions within an agile team.
- 3+ years of professional experience in Data Engineering, Software Engineering, or a related field.
- 2+ years of hands-on experience building and maintaining production data pipelines supporting analytics, reporting, or machine learning workloads.
- Strong proficiency in Python and SQL with experience developing production-quality data pipelines.
- Experience with modern data processing technologies such as Spark, Airflow, Kafka, or similar distributed data platforms.
- Experience working with cloud-based data platforms such as Databricks, Snowflake, Delta Lake, or equivalent lakehouse technologies.
- Understanding of data modeling, data warehousing, and data governance best practices.
- Familiarity with machine learning data workflows, including training datasets, feature engineering, and data quality concepts.
- Experience deploying and supporting production data pipelines with monitoring, testing, and CI/CD practices.
- Strong problem-solving skills, attention to detail, and the ability to collaborate effectively across engineering and product teams.
- Excellent communication skills and a desire to continuously learn new technologies and engineering practices.
#LI-SS1 #LI-Remote