Job Title: AI Data Engineer (SP)
Job Location: Princeton, NJ & NYC, NY
Job Type: Full-Time
Job Description:
- Design and implement specialized data pipelines for AI model metadata, training data lineage, and model performance metrics tracking.
- Build data infrastructure on Databricks leveraging Spark for large-scale distributed dataset processing.
- Develop MCP servers and enable AI data distribution via MCP.
- Develop feature engineering pipelines and data preprocessing workflows for AI model training and inference.
- Implement model versioning, experiment tracking, and model registry integration using MLflow or similar tools.
- Create automated workflows for AI agent discovery, classification, and inventory management across the enterprise.
- Design and maintain knowledge graph structures for representing AI model relationships, dependencies, and data lineage.
- Build real-time data pipelines for AI model monitoring, drift detection, and performance tracking.
- Develop data quality frameworks specific to AI training datasets and validation data.
- Collaborate with data scientists to optimize data access patterns and feature store implementations.
- Implement security and compliance controls for sensitive AI training data and model artifacts.
- Create comprehensive documentation for AI data architectures, schemas, and integration patterns.
Required Skills and Qualifications
- Bachelor's or Master's degree in Computer Science, Data Science, Machine Learning, or related field.
- 5-7 years of hands-on experience in data engineering, with at least 2 years focused on AI/ML workloads.
- Expert proficiency in Python and experience with ML frameworks such as PyTorch, TensorFlow, or scikit-learn.
- Strong experience with Databricks, Apache Spark, and distributed computing for ML workflows.
- Deep understanding of the machine learning lifecycle, including model training, deployment, and monitoring processes.
- Experience with feature engineering, data preprocessing techniques, and ML data pipelines.
- Knowledge of vector databases, embeddings, and similarity search for AI applications.
- Proficiency in SQL for structured and unstructured data management.
- Understanding of data governance, model governance, and AI ethics principles.
- Strong analytical and problem-solving capabilities with attention to data quality.
- Excellent collaboration skills for working with data scientists, ML engineers, and architects.
Preferred/Nice-to-Have Skills
- Experience with generative AI applications, including RAG (Retrieval-Augmented Generation) and fine-tuning.
- Knowledge of LangChain, HuggingFace, or other GenAI frameworks.
- Familiarity with Azure ML, AWS SageMaker, or Google Vertex AI platforms.
- Experience with graph databases (Neo4j, Amazon Neptune) for knowledge graph implementation.
- Understanding of AI model explainability and interpretability techniques.
- Experience with A/B testing frameworks for ML model evaluation.
- Certification in Databricks, AWS, Azure, or GCP AI/ML services.
- Publications or contributions to open-source ML projects