Data Engineer AI/ML & Advanced Analytics
Remote(Minnetonka Mills, MN)
Data Engineer AI/ML & Advanced Analytics Position Summary
We are seeking a highly skilled Data Engineer to support enterprise AI/ML and Advanced Analytics initiatives by building scalable, reliable, and high-performance data platforms and pipelines.
This role will focus on designing and engineering modern data solutions that enable machine learning, advanced analytics, and AI-driven decision making. The ideal candidate has strong expertise in data architecture, cloud platforms, large-scale data processing, and data engineering best practices.
The successful candidate will build and optimize data ecosystems that empower data scientists, analysts, and business teams to generate actionable insights and operational value.
Key Responsibilities
Design, develop, and maintain scalable data pipelines supporting AI/ML and analytics workloads.
Build and optimize batch and real-time data ingestion frameworks.
Develop data integration solutions across multiple internal and external data sources.
Engineer reliable datasets and feature stores that support machine learning model development.
Implement data transformation, cleansing, enrichment, and validation processes.
Design and maintain modern data lake, lakehouse, and data warehouse architectures.
Ensure data quality, integrity, security, and governance standards are met.
Optimize data processing performance, scalability, and operational efficiency.
Collaborate with Data Scientists and Analysts to support feature engineering and model deployment requirements.
Enable MLOps and ML platform capabilities to support model operationalization.
Implement monitoring, observability, and operational support processes for data platforms.
Maintain documentation, data lineage, and metadata management standards.
Required Qualifications
Bachelor's degree in Computer Science, Engineering, Information Systems, Data Engineering, or related field.
4+ years of experience in Data Engineering, Data Platforms, or Analytics Engineering.
Strong proficiency in Python, SQL, Spark, Scala, or equivalent technologies.
Experience building cloud-native data solutions using Azure, AWS, or GCP.
Experience with ETL/ELT frameworks and large-scale data processing.
Knowledge of distributed data processing technologies and modern data architectures.
Experience with data lake, warehouse, and lakehouse platforms.
Understanding of data governance, security, and data quality practices.
Strong collaboration and problem-solving skills.
Role Descriptions: Design and develop reusable data engineering components| templates| and standards for enterprise-wide use.Build and optimize data processing pipelines using Python| PySpark| and Databricks (Delta Lake| workflows| jobs| notebooks| Unity Catalog).Strong AWS experience across compute| storage| networking| and orchestration services.Triage| diagnose| and resolve complex data processing and pipeline failures with strong analytical skills.Mentor and guide developers across multiple Scrum teams| providing technical leadership| code reviews| and architectural direction.Produce clear| thorough documentation of code| patterns| and solution designs for enterprise reuse.Model and transform complex healthcare datasets with strong understanding of claims| eligibility| provider| and other payer data.Experience with CICD| automated testing| and job promotion using GitHub Actions.Ensure data security| governance| lineage| and HIPAAPHI compliance in all solutions.Participate in AgileScrum ceremonies| lead design discussions and support story refinement.
Essential Skills: Data Engineering| Databricks| PySpark| SQLmust have experience on agentic AI
Desirable Skills:
Keyword:
Skills: Digital : Databricks~Digital : PySpark~AI Agents~Azure Data Factory
Experience Required: 4-6
Data Engineer AI/ML & Advanced Analytics Position Summary
We are seeking a highly skilled Data Engineer to support enterprise AI/ML and Advanced Analytics initiatives by building scalable, reliable, and high-performance data platforms and pipelines.
This role will focus on designing and engineering modern data solutions that enable machine learning, advanced analytics, and AI-driven decision making. The ideal candidate has strong expertise in data architecture, cloud platforms, large-scale data processing, and data engineering best practices.
The successful candidate will build and optimize data ecosystems that empower data scientists, analysts, and business teams to generate actionable insights and operational value.
Key Responsibilities
Design, develop, and maintain scalable data pipelines supporting AI/ML and analytics workloads.
Build and optimize batch and real-time data ingestion frameworks.
Develop data integration solutions across multiple internal and external data sources.
Engineer reliable datasets and feature stores that support machine learning model development.
Implement data transformation, cleansing, enrichment, and validation processes.
Design and maintain modern data lake, lakehouse, and data warehouse architectures.
Ensure data quality, integrity, security, and governance standards are met.
Optimize data processing performance, scalability, and operational efficiency.
Collaborate with Data Scientists and Analysts to support feature engineering and model deployment requirements.
Enable MLOps and ML platform capabilities to support model operationalization.
Implement monitoring, observability, and operational support processes for data platforms.
Maintain documentation, data lineage, and metadata management standards.
Required Qualifications
Bachelor's degree in Computer Science, Engineering, Information Systems, Data Engineering, or related field.
4+ years of experience in Data Engineering, Data Platforms, or Analytics Engineering.
Strong proficiency in Python, SQL, Spark, Scala, or equivalent technologies.
Experience building cloud-native data solutions using Azure, AWS, or GCP.
Experience with ETL/ELT frameworks and large-scale data processing.
Knowledge of distributed data processing technologies and modern data architectures.
Experience with data lake, warehouse, and lakehouse platforms.
Understanding of data governance, security, and data quality practices.
Strong collaboration and problem-solving skills.
Role Descriptions: Design and develop reusable data engineering components| templates| and standards for enterprise-wide use.Build and optimize data processing pipelines using Python| PySpark| and Databricks (Delta Lake| workflows| jobs| notebooks| Unity Catalog).Strong AWS experience across compute| storage| networking| and orchestration services.Triage| diagnose| and resolve complex data processing and pipeline failures with strong analytical skills.Mentor and guide developers across multiple Scrum teams| providing technical leadership| code reviews| and architectural direction.Produce clear| thorough documentation of code| patterns| and solution designs for enterprise reuse.Model and transform complex healthcare datasets with strong understanding of claims| eligibility| provider| and other payer data.Experience with CICD| automated testing| and job promotion using GitHub Actions.Ensure data security| governance| lineage| and HIPAAPHI compliance in all solutions.Participate in AgileScrum ceremonies| lead design discussions and support story refinement.
Essential Skills: Data Engineering| Databricks| PySpark| SQLmust have experience on agentic AI
Desirable Skills:
Keyword:
Skills: Digital : Databricks~Digital : PySpark~AI Agents~Azure Data Factory
Experience Required: 4-6