Data Engineer
Whisker is redefining what it means to live with cats—designing intelligent systems that remove friction, elevate the everyday, and celebrate the quiet brilliance of feline companionship. Today, Litter-Robot leads the category. Tomorrow, an entire ecosystem that expands what’s possible for cats and the people who love them. We believe the future is feline. And we’re imagining that future today.
We work onsite 4+ days a week, with our team based in Auburn Hills, Michigan, and Juneau, Wisconsin. Our team of 700+ passionate pet people thrives on collaboration, innovation, and the occasional office cameo from a four-legged friend.
What You’ll Do:
We are seeking a Senior Data Engineer / Data Engineer with strong hands-on Databricks experience to design, build, and optimize scalable data pipelines and platforms that power analytics and decision-making across the organization. This role partners closely with data science, analytics, and engineering teams to deliver reliable, high-performance data solutions in a modern cloud lakehouse environment.
Summary:
The Senior Data Engineer / Data Engineer is responsible for designing, developing, and maintaining scalable data pipelines and infrastructure using Databricks and related cloud technologies. This role ensures data is reliable, secure, and readily accessible to support analytics, reporting, and machine learning initiatives.
Essential Duties and Responsibilities
- Designs, builds and maintains scalable ETL/ELT pipelines using Databricks, Apache Spark, and Delta Lake
- Develops and optimizes data workflows for batch and streaming data processing
- Architects and implements data lakehouse solutions following medallion architecture (bronze/silver/gold layers using DBT)
- Collaborates with data scientists, analysts, and business stakeholders to understand data requirements and deliver solutions
- Writes efficient, well-documented PySpark/SQL code for data transformation and processing
- Develops and supports MLOps workflows, including model tracking, versioning, and lifecycle management using MLflow
- Builds and maintains CI/CD pipelines for both data engineering and machine learning model deployment
- Implements monitoring and alerting for data pipelines and ML models to proactively identify and troubleshoot performance issues, data drift, and failures
- Monitors, troubleshoots, and optimizes pipeline and model performance, resource utilization, and cost efficiency within Databricks
- Ensures data quality, integrity, and governance across all pipelines and datasets
- Manages and optimizes Databricks clusters, jobs, and workflows
- Integrates Databricks with cloud platforms (Azure, AWS, or GCP) and other enterprise systems
- Implements data security best practices, including access controls and data masking
- Mentors junior data engineers and provide technical guidance on best practices, including MLOps standards
- Participates in architecture reviews and contribute to the evolution of the data and ML platform strategy
- Documents technical designs, data models, pipeline processes, and ML deployment workflows
- Stays current with emerging tools, technologies, and best practices in data engineering and MLOps
- Performs additional responsibilities when required