Shibbir Ahmed Arif
Data Scientist @ DetaPent
About
Data Scientist with 3+ years of experience designing and deploying AI/ML solutions across healthcare, finance, and nonprofit sectors. Currently leading intelligent system development at DetaPent Inc., and previously led Generative AI research at Montclair State University’s Data Science Lab, where I built LLM-based RAG pipelines for clinical NLP. Skilled in predictive modeling, natural language processing, and end-to-end design of real-time machine learning systems. Experienced in A/B testing, data augmentation, annotation labeling, response generation, similarity scoring, relevance evaluation, and ranking. Deeply proficient in Python, PyTorch, LangChain, and cloud platforms, with a proven track record of transforming unstructured data into scalable, high-impact solutions aligned with strategic goals. Committed to building AI systems that drive automation, improve decision-making at scale, and deliver meaningful business and societal outcomes. TECHNICAL SKILLS • Programming Languages: Python, SQL, R • Machine Learning Frameworks: PyTorch, TensorFlow, Scikit-learn • Natural Language Processing (NLP): Hugging Face Transformers, Sentence Transformers, spaCy, NLTK • Large Language Models (LLMs): DeepSeek R1, LLaMA , OpenAI APIs, Groq, LangChain, RAG Systems, Prompt Engineering, Fine-tuning • Data Processing & Visualization: Pandas, PySpark, Dataiku, Plotly, Looker, Power BI, Tableau • Data Engineering Tools: ETL Pipelines, Apache Spark, Google BigQuery, Azure Synapse, Azure Data Factory • Databases: MS SQL Server, PostgreSQL, MySQL, OracleDB, MongoDB • Cloud Platforms: AWS (S3, SageMaker), Azure, Google Cloud (GCP) • Version Control: Git, GitHub • Containerization & Deployment: Docker, Flask, Streamlit, Hugging Face Spaces • Scientific Writing & Research Tools: LaTeX, Overleaf • Other Skills: Linux/Bash Scripting
United States
Montclair
Information Technology & Services
R Programming, Image Classification, Hugging Face Transformers, Machine Learning & Predictive Modeling, Cloud & Retrieval Systems: AWS, FAISS, ChromaDB, Generative AI & LLMs (LangChain, Prompt Engineering), Python, PyTorch & Deployment (Streamlit, Flask), PyTorch, Statistics, Retrieval-Augmented Generation (RAG), Application Programming Interfaces (API), Continuous Integration and Continuous Delivery (CI/CD), AWS Lambda, AWS S3, Big Data Analytics, LangChain, Chatbot Development, Vector Databases, Bootstrap (Framework), Sentiment Analysis
Experience

Data Scientist
New Jersey, United States
• Uncovered critical business and clinical insights, as measured by strategic decisions influenced across teams, by analyzing large-scale datasets using Python, SQL, and AWS. • Improved predictive accuracy for patient outcomes and financial planning, as measured by successful deployment of models in production, by building machine learning models (e.g., XGBoost, Random Forest) to forecast readmissions, disease progression, claims costs, and payment delays. • Boosted performance of GenAI and NLP systems, as measured by increased model accuracy and reduced preprocessing errors, by engineering scalable data pipelines to handle structured and unstructured multimodal data. • Enabled data-driven decision-making across departments, as measured by increased adoption of dashboards among technical and non-technical stakeholders, by designing interactive reports using Power BI and Tableau. • Delivered scalable, production-grade AI solutions, as measured by successful deployment of GenAI chatbots and clinical trial matching tools, by leading end-to-end development using PyTorch, LangChain, FAISS, and AWS SageMaker.

Adjunct Faculty | Data Science
Montclair, NJ
• Mentored 100+ students in Data Science and Statistics through hands-on labs and interactive modules using R and JMP, enhancing engagement and practical skills. • Assessed learning outcomes by grading assignments, labs, and exams, ensuring consistent performance improvement and high course satisfaction.

Associate Data Scientist
Data Science Lab at Montclair State University
Montclair, NJ
• Improved patient screening efficiency by 20%, as measured by faster eligibility assessments, by developing an LLM-powered NLP pipeline using Python, applying deep learning and machine learning techniques to process over 10,000 MIMIC-IV clinical records. • Increased model precision and recall by 35%, as measured by evaluation metrics on test datasets, by engineering a robust LLaMA 3.2-based system with predictive modeling and bias evaluation to assess patient eligibility across diverse demographics. • Reduced real-time query latency by 50%, as measured by response time benchmarks, by designing and deploying a FAISS vector store with optimized preprocessing for efficient patient data retrieval.

Graduate Student Worker | AI & Big Data
New Jersey, United States
• Improved student engagement and comprehension, demonstrated through increased participation and positive feedback, by leading weekly discussion sessions on Python programming. • Ensured consistent and fair evaluation, reflected in timely grading and strong student outcomes, by designing and assessing assignments and labs via Canvas and zyBooks. • Assessed advanced coding proficiency, evidenced by midterm and final exam performance, by creating original problem-solving assessments focused on real-world Python applications. • Enabled successful completion of 8 AI-driven IoT projects (47 students), validated by on-time delivery and high-quality submissions, by supervising and mentoring teams using Python and Raspberry Pi 5, and managing project timelines. • Streamlined course coordination and student support, confirmed by smooth class operations and effective communication, by collaborating with the faculty and 6 Teaching Assistants to support instruction.

Graduate Teaching Assistant | Python & Database
New Jersey, United States
• Improved lab engagement and learning outcomes, by assisting faculty in delivering hands-on sessions for courses such as Python Programming I and Database Systems. • Maintained instructional quality, by creating lab assignments and delivering substitute lectures during faculty absences. • Enhanced student support, by holding regular office hours for both in-person and online learners, leading to increased course satisfaction. • Streamlined course delivery, by managing and updating Canvas platforms to support assignments, grades, and resources. • Ensured timely and consistent grading, by evaluating labs and assignments across core computing courses.

Data Science Analyst Intern
Hillside, NJ
• Optimized resource allocation by 20% by analyzing over 100GB of structured and unstructured data using Python, SQL, and Power BI, driving smarter funding and logistics decisions. • Increased operational efficiency by 25% by implementing data-driven strategies that supported the equitable distribution of a $4M grant, benefiting over 1,000 community members. • Boosted executive insight and visibility by building interactive Power BI dashboards and reports that translated complex analyses into clear, actionable decisions. • Developed a dynamic Maply heat map to identify underserved regions across New Jersey, guiding outreach strategies for food access and community support. • Analyzed food distribution and partner performance data, supporting the success of partner tiering initiatives to align resources with impact. • Collaborated cross-functionally with sourcing, programs, and partnerships teams to align data insights with frontline decision-making.

Senior Data Science Analyst
Narayanganj Sadar Upazila, Dhaka, Bangladesh
• Improved data integrity by 25% by optimizing ETL pipelines and unifying messy data from relational databases and APIs using Python and advanced data wrangling techniques. • Accelerated decision-making by deploying interactive Tableau dashboards, boosting cross-functional engagement by 30% and enabling real-time insights. • Increased SQL query efficiency by 15% by restructuring complex workflows, enhancing access to high-impact business intelligence. • Extracted actionable insights from large datasets using statistical analysis in Python, uncovering trends that informed strategic decisions. • Translated analytical findings into strategic recommendations by creating clear, data-driven reports for product, marketing, and operations teams. • Streamlined data ingestion pipelines by integrating datasets from spreadsheets, databases, and APIs into a centralized analytics workflow.

Research Analyst
Chattogram, Bangladesh
• Achieved 87% classification accuracy using SVM and TF-IDF to analyze tweet reviews, outperforming four other supervised machine learning models. • Enhanced text preprocessing by converting emojis and emoticons into English tokens, improving opinion signal clarity and model performance. • Processed and labeled 50GB+ customer tweets, categorizing sentiment into five classes to analyze perceptions of leading smartphone brands. • Automated the NLP pipeline, reducing preprocessing time by 40% and enabling scalable sentiment analysis across large datasets. • Presented findings at the EICT Conference, showcasing applied NLP techniques and insights to a professional and academic audience.
Education

Computer Science & Engineering
Key Notes: - Class Rank (2 out of 45 students) - PCIU Volunteers Team Leader - PCIU Math Olympiad Team Member - B.Sc. Thesis Title: Supervised Machine Learning Based Opinion Mining on Twitter Data - Coursework includes Artificial Intelligence, Algorithm, Data Structure, Database Systems, Compiler, Computer Graphics, Computer Network, Digital Logic Design, Computer Organisation & Architecture, etc.
Shibbir Ahmed Arif's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.


