Job ID: 108766-1
Title: Data and AI Engineer Automotive Engineering Analytics
Location: Warren, OH (ONSITE)
Duration: 12+ Months
Pay Range: $45 - $48 an hour on W2/ C2C (All Inclusive)
Role: Data and AI Engineer Automotive Engineering Analytics
Must Have Skills:
Must Have:
- Strong experience in Python, SQL, Data Engineering, and Analytics.
- Experience with Automotive, Manufacturing, Engineering, Quality, Reliability, Warranty, or Telemetry data.
- Hands-on experience with Machine Learning, Statistical Analysis, and AI/GenAI technologies (LLMs, RAG, Embeddings, Vector Search, Knowledge Graphs).
- Experience building ETL/Data Pipelines on Cloud platforms.
- Experience with dashboards, web applications, APIs, and data visualization.
- Knowledge of CI/CD, testing frameworks, and data quality monitoring.
- Ability to process structured and unstructured data and collaborate with business stakeholders.
Responsibilities:
- Build and maintain scalable data pipelines and analytical solutions.
- Clean, transform, validate, and analyze data from multiple sources.
- Develop ML models and AI-powered applications.
- Implement LLM/RAG-based solutions and ensure accuracy, traceability, and data security.
- Create dashboards, reports, and actionable insights for engineering teams.
- Monitor model performance, data quality, and system reliability.
Pre-Screening Questionnaire
Describe your experience using Python and SQL to clean, transform, join, validate, and analyze data from multiple sources.
Describe one machine-learning or statistical-analysis project you delivered. What methods, evaluation metrics, validation approach, and business or engineering outcome were involved?
What hands-on experience do you have with AI-enabled applications such as LLMs, retrieval-augmented generation, embeddings, vector search, or knowledge graphs, and how did you address accuracy, traceability, human review, and data protection?
Describe a production-grade data pipeline you designed or supported. How did you handle ingestion, transformation, missing or duplicate records, validation, data lineage, and scalability?
How would you monitor a deployed machine-learning model for drift, changing data quality, false positives, false negatives, and performance regression?