Description:
Qualifications for Data Scientist Strong problem solving skills with an emphasis on product development. Experience using statistical computer languages (R, Python, SQL, etc.) to manipulate data and draw insights from large data sets. Experience working with and creating data architectures. Knowledge of a variety of machine learning techniques (clustering, decision tree learning, artificial neural networks, etc.) and their real-world advantages drawbacks. Knowledge of advanced statistical techniques and concepts (regression, properties of distributions, statistical tests and proper usage, etc.) and experience with applications. Excellent written and verbal communication skills for coordinating across teams. A drive to learn and master new technologies and techniques. Experience manipulating data sets and building statistical models, has a Master s or PHD in Statistics, Mathematics, Computer Science or another quantitative field, and is familiar with the following software tools: Coding knowledge and experience with several languages: C, C++, Java, JavaScript, etc. Knowledge and experience in statistical and data mining techniques: GLM Regression, Random Forest, Boosting, Trees, text mining, social network analysis, etc. Experience querying databases and using statistical computer languages: R, Python, SLQ, etc. Experience using web services: Redshift, S3, Spark, , etc. Experience creating and using advanced machine learning algorithms and statistics: regression, simulation, scenario analysis, modeling, clustering, decision trees, neural networks, etc. Experience analyzing data from 3rd party providers: Client Analytics, Site Catalyst, Coremetrics, Adwords, Crimson Hexagon, Client Insights, etc. Experience with distributed data/computing tools: Map/Reduce, Hadoop, Hive, Spark, Gurobi, MySQL, etc. Experience visualizing/presenting data for stakeholders using: Periscope, Business Objects, D3, ggplot, etc. Top Daily Responsibilities: 1. Support Data-Science and other analytics as needed. 2. Develop SQL queries and data sets 3. Develop business and client facing reports Skills a Top Candidate Should Have: 1. Knowledge and experience in statistical and data mining techniques: GLM/Regression, Random Forest, Boosting, Trees, text mining, social network analysis, etc. 2. Experience querying databases and using statistical computer languages: R, Python, SLQ, etc. 3. Experience using web services: Redshift, S3, Spark, DigitalOcean, etc. 4. Experience creating and using advanced machine learning algorithms and statistics: regression, simulation, scenario analysis, modeling, clustering, decision trees, neural networks, etc. 5. Experience analyzing data from 3rd party providers: Client Analytics, Site Catalyst, Coremetrics, Adwords, Crimson Hexagon, Client Insights, etc. 6. Experience with distributed data/computing tools: Map/Reduce, Hadoop, Hive, Spark, Gurobi, MySQL, etc. 7. Experience visualizing/presenting data for stakeholders using: Periscope, Business Objects, D3, ggplot, etc. Desired Skills: 8. Strong problem solving skills with an emphasis on product development. 9. Experience using statistical computer languages (R, Python, SQL, etc.) to manipulate data and draw insights from large data sets. 10. Experience working with and creating data architectures. 11. Knowledge of a variety of machine learning techniques (clustering, decision tree learning, artificial neural networks, etc.) and their real-world advantages/drawbacks. 12. Knowledge of advanced statistical techniques and concepts (regression, properties of distributions, statistical tests and proper usage, etc.) and experience with applications. 13. Excellent written and verbal communication skills for coordinating across teams. 14. A drive to learn and master new technologies and techniques. 15. We re looking for someone with experience manipulating data sets and building statistical models, has a Master s or PHD in Statistics, Mathematics, Computer Science or another quantitative field, and is familiar with software. Soft Skills: 1. Excellent Communication Skills. 2. Ability to work with business to gather report requirements 3. Team player. Custom Job Description: If you have a custom job description that you would like to use. Please paste it here: Knowledge and experience with large data sets, event streams and distributed computing (Hive,Impala,Hadoop etc.) Ability to gather requirements and develop reports in tool selected by business and KPIT. ENTER YEARS OF EXPERIENCE REQUIRED.
Enable Skills-Based Hiring
No
Additional Job Details
Data Scientist LLM role
The ideal candidate will have a strong background in natural language processing (NLP), deep learning, and experience with large-scale model training. This role requires expertise in model architecture, optimization techniques, and the ability to work with cross-functional teams to integrate AI solutions into products.
**Key Responsibilities:
- Design and develop LLMs for specific use cases, including but not limited to text generation, summarization, and conversational agents.
- Prompt engineer pre-trained LLMs on domain-specific data to optimize performance for target applications.
-Develop RAG based applications.
- Collaborate with data scientists, engineers, and product teams to integrate LLM capabilities into existing systems and platforms.
- Optimize model performance in terms of accuracy, speed, and resource utilization.
- Ensure the ethical use of AI models and implement fairness, accountability, and transparency measures.
**Required Qualifications:
- Strong grasp of NLP concepts such as tokenization, embeddings, attention mechanisms, and sequence modeling.
- Experience in handling domain adaptation, multi-task learning, and transfer learning.
- Experience with popular deep learning frameworks (e.g., TensorFlow, PyTorch).
- Proven track record of promt engineering, RAG system development and fine-tuning large pre-trained models (e.g., GPT, BERT, T5, etc.).