Harold COMPAORE
Data Scientist @ CommonShare
About
Hi, I’m COMPAORE Yolemba Harold — a junior data scientist passionate about leveraging data insights to drive impactful decisions. I have hands-on experience in data science, machine learning, and deep learning, particularly in using statistical analysis techniques and models for regression and classification tasks. During my internships, I worked on projects involving time-series forecasting and startup status prediction, using tools such as Pandas, NumPy, Matplotlib, Seaborn, Scikit-learn, TensorFlow, Google Colab, Kaggle, Git, and GitHub. Currently, I’m pursuing engineering studies at ENSIAS, where I’m taking courses in Machine Learning, Data Visualization, Statistical Analysis, Data Warehousing, J2EE Web Development, and Industrial IoT. I also continue to grow my skills through platforms like Coursera and YouTube. Beyond my technical background, I bring strong soft skills such as communication, problem-solving, teamwork, and adaptability. I’m fluent in French (native) and have an intermediate level of English.
Morocco
Rabat
Information Technology & Services
Large Language Models (LLM), IBM Watson, AI Agents, Retrieval-Augmented Generation (RAG), LangChain, Transformers, Hugging Face Products, Generative AI, Deep Learning, PyTorch, Robots de trading, Talend Open Studio, PostgreSQL, Django, Amazon Web Services (AWS), Trading, Spring Boot, MariaDB, React.js, Credit Scoring
Experience

Data Scientist
Inde
During this internship, I worked on a project focused on predicting the current status of a startup. The main goal was to classify startups as either Operating, IPO, Acquired, or Closed. To achieve this, I performed a variety of tasks, including: -cleaning original dataset by deleting redundant and irrelevant information, handling missing data and outliers and date variables transformation. This changed dataset shape from (196 553, 44) to (57 660, 16) -applying univariate, bivariate and multivariate analysis in order to find out insights in data. -performing feature engineering by applying log transformation, standardisation, creating 4 new features and encoding variables; -oversampling to balance classes by creating over 100 000 new instances -modelling using algorithms like LogisticRegression, GaussianNB, AdaBoostClassifier, GradientBoostingClassifier to achieve accuracy of 0.99 in binary classification and 1.0 in multiclass classification Tools & Technologies: -Data Manipulation: Pandas, Numpy -Visualization: Matplotlib, Seaborn -Machine Learning: Scikit-learn, XGBoost -Oversampling: imbalanced-learn (Imblearn) -Collaboration & Version Control: Git, GitHub -Environment: Google Colab, Kaggle

Research Assistant
Rabat, Rabat-Salé-Kénitra, Maroc
During this internship, the project I worked on involved imputing values into datasets with different percentages of missingness (5%, 10%, 20%, 30%, 40%, 50%). The approach used here was training machine learning and deep learning models on these datasets. Main tasks: -Implementing an algorithm to convert a time series dataset into sequences of fixed size (10) to prepare it for supervised learning -Designing an algorithm to systematically scans the dataset to detect all occurrences of TS-MD (Time Series Missing Data) and train a model to impute this missing values -Training ML/DL models like SVR, RNN, and LSTM on the training set and comparing their performance on test set. SVR performed best, with a median RMSE (Root Mean Squared Error) value of 7.473013. Tools: Numpy, Pandas, Scikit learn, Tensorflow, Keras, metrics, Google Colab, Kaggle
Harold COMPAORE 's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.




