Guichong Li
Data Scientist/AI Engineer
About
SUMMARY • The main research focuses on Bayes learning, multiple layer Bayes learning; • Effectively designed and implemented machine learning algorithms for churn prediction and Add a Line analysis and machine learning pipelines for telecom business; • Developed advanced machine learning algorithms for text, POS outlet items, category hierarchy classification; multilabel and multitask classification algorithms; • Developed advanced machine learning regression/classification algorithms for food component analysis using chemometrics and spectroscopy; • Recent postdoctoral research on uniformly and unbiased sampling/crawling online social networks using advanced Markov Chain Monte Carlo techniques; developed an innovative and unique sampling algorithm, a new coupling technique, implemented by Ruby and Rails and Twitter API, DataMapper; Unix/Linux, Amazon EC2; social media analysis using Python, NLTK, SKLearn. • Previous postdoctoral research in DRDC, CORA, Canada, for Complex Dynamic Network Analysis; Simulation of Autonomous Underwater Vehicles; using MatLab and VB, etc. • Two year research contract in Health Canada for nuclear explosion and pollution monitoring, and environmental anomaly detection; using J# and Weka (Java) software package, Eclipse. • The main research interest focus on Machine Learning and Data Mining algorithms and technology; in particular, one-class learning using kernel methods for anomaly detection and its application on big data using MapReduce/Hadoop with Pig/Hive/HBase, and advanced Markov Chain Monte Carlo techniques for fast and unbiased sampling/crawling online social networks such as Twitter and Facebook. • Having both Mathematics and Computer Science education backgrounds; 10 year professional experience for software development, various artificial intelligence algorithm design, and leadership for transaction and database applications using SQL Server, ORACLE, C/C++, Java/J#, VB, JDBC, .Net., PowerBuilder, TCP/IP, OpenGL;
Canada
Markham
Government Administration
Keras, Statistics, TensorFlow, Data Science, Deep Learning, Computer Vision, Statistical Modeling, Python, Machine Learning, Algorithms, Data Mining, Java, Eclipse, MySQL, XML, Artificial Intelligence, Software Development, PostgreSQL, C++, C#
Experience

Data Scientist/AI Engineer
Transportation Canada
Ottawa, ON
Data Scientist/AI engineer, Aug- 2025, Feb, 2026, Transportation Canada • Design and implement the shortterm rail prediction system on azure databricks using pyspark. • railstate and weather data parsing and cleaning, transformation on azure databricks using pyspark. • building ETL pipeline to extract 7 days forecasting from open-meteo service. • create training sets for train carload prediction from historical data and conduct time series analysis • training LSTM time series forecast model and RandomForestRegressor model for 7 days railstate prediction based on 7 days open-meteo forecasting • MLFlow for model training, tracking, registration, and inference. • Implement and deploy a conversational AI agent for Geo.ca in National Natural Resource on AWS Sagemaker AI • Implement AI agent using langchain and bedrock, detect intent and tool routing; create knowbase for RAG pipeline; design step function for orchestration. • FastAPI for asynchronzation and pydantic for formatting; lambda function and deployment using sam and docker container

AI Engineer
Mississauga, Ontario, Canada
• Develop GenAI solutions for banking services: automatic validation of interest rate swaps; pipeline for management of trade/cashflow/payment; payment near missing. • Genai for interest rate forecasting using gemini 2 and gpt 4, meta llama 3. LLM models evaluation use Gval, etc; • Designing agent chains for transaction visualization analytics. • PyRIT development for AI Banking services: transformation of seedprompts from different sources, Gandalf prompt and evaluation, customization of targets; enhancement by parallel inference; • MCP for powered RAG, etc. • Designing and implementing multi-agent using LangChain and LangGraph to resolve payment near missing problem occurring in cashflow. Work with agentic AI systems that can plan, reason, and execute tasks autonomously to improve cashflow. Design, monitor, and optimize AI agents for task automation, decision-making, and workflow orchestration. • Fine-tune image recognition models for cheque segmentation; create training sets by annotating check images using SuperAnnotate, VoTT, LabelImg, and annotation format conversion between csv and pascalVoc, etc. • Image preprocessing such as dilation, erosion, Gaussian blur, contour/boxes, for OCR enhancement for check segmentation using tesseract and paddleOCR. Technology: GenAI business solution using Gemini 2, GPT 4, meta llama 3; Gval evaluation metrics; PyRIT development for financial security; MCP and RAG enhancement; agent chain and multi-agent orchestration using langGraph; Annotation using SuperAnnotate, Vott, LabelImg, PaddleOCR for check segmentation

AI/ML Engineer
TB, PSPC
Ottawa, Ontario, Canada
Data scientist, PSPC, MaxSys Collect and create data Catalog and dictionary in the organization; develop the data strategy by implementing the data flow of information in Translation Bureau which improve and business processes. Specially two AI platforms, azure synapse analytics and azure data factory, are widely investigated to improve efficiency and effectiveness for business analysis and applications. For example, a new data strategy using the dedicate pipeline in data factory implements millions of data query and transformation; automate Sigma analysis with GCLingua, and Sigma API for business financial analysis. Technology: azure synapse analytics, azure data factory, lakehouse, HDIinsight service manages OS systems, power query with incremental, directquery, query folding, computer engine; Sigma API

AI Engineer
Ottawa, Ontario, Canada
AI developer/Python programmer, IRCC, Altis Project I: IBM SPSS python conversion. Implemented end-to-end python pipeline with spark on AWS cloud for the original SPSS stream models built on large datasets. The main tasks include data collection, spss stream python conversion for spss type, filter, select, filler, merge nodes, aggregate nodes, flag nodes supernodes, cache, statistical outputs, with pyspark with spark on AWS cloud platform; unit tests and spark storage plan for optimization. Further, including techniques to transfer to Google Cloud Platform (GCP) services with BigQuery, DataFlow, Pub/Sub, BigTable, Data Fusion, DataProc, Cloud Composer, Cloud SQL, Compute Engine, Cloud Functions, and App Engine; BigQueryML, AutoML, Vertex AI; Graphic network analysis using Neo4j with community detection, key influencer, spread of information, bottleneck, connectivity Technology: Pandas, Pyspark(filter, count, columns, drop, withcolumrename), spark, spark, sql(to_date, col, lit, when, trim, substring, etc); spss (stream, filter node, filler node, type node, marge node, aggregate node, etc); GCP: BigQueryML, AutoML, Vertex AI. Project II: (computer vision, opencv, Pillow, deep learning, tensorflow, , keras, pytorch, Siamese networks, CNN, VGG16, Resnet50, Xception, etc) fraud detection for financial applications. Implement and develop python tools and platform to automatically extract handwriting and signature, signature verification, handwritten and printed text classification, from docs and check images, using deep learning models and image preprocessing methods, thresholding, connected component analysis; object detection and segmentation; remove noise from background; template matching, text extraction using keras ocr, tesseract, easyocr. LLM and ChatGPT for document verification

Data scientist and machine learning engineer
Montreal, Quebec, Canada
Data Scientist/Machine learning engineer, CN rail / SpruceInfotech Developed and implemented Python scripts for data parsing, data imputation, and data encoding using sklearn, pandas. Developed Python scripts to train and build models and to run tests to evaluate system performance of AI solutions, using sklearn, pandas; developed and implemented python scripts to AI solutions for forecast modeling and regression modeling and classification modeling; developed and implemented python scripts for log time series analysis; Developed MLflow for model tracking, training, logging, registration, inference, hyperopt/parameter sweep. Univariate/multivariate forecasting and regression Analyzed and validated business requirements and review of solutions with relevant stakeholders; the technical report for MLFlow project development and production solution for azure cloud AI solution; the research report for improvement of forecast models for rail transportation with Azure

Data Scientist
Ottawa, Ontario, Canada
Data scientist, Solana networks Project: Data Management, Labelling and Automation for Machine Learning; Activity Recognition and Hierarchical Labelling; OpenPose and HumanActionClassification platform Context: image/video datasets downloaded from publicly-available sources or created by using annotation tools, and parsing and importing the data items into the MongoDB database (both image and video content). A tool supports the creation of training data collections in the database, subject to different criteria such as ImageNet, COCO, StanfordDrone. Finally, implement training set creation; and active learning and experimentation; reporting the results of a human pose estimation and activity recognition prototype solution by building on existing open source packages. validation with TIDE and PyBrisque. Technology and tools: Python, Pytorch, Tensorflow, ImageNet, COCO, StanfordDrone, VoTT, LabelMe annotation, boundingbox, segmentation, active learning, OpenPose, HumanActionClassification, TIDE, PyBrisque

Data Scientist
Ottawa, Ontario, Canada
Data scientist, Solana networks Project: Anomaly Detection for In-Vehicle Networks Context: anomaly traffics (attacks) might be found in automotive bus system (CAN) which transmits signals between electronic control units (ECUs) as well as wireless interfaces such as GSM and Bluetooth. Attackers try to access the automotive network in order to inject messages, manipulate data or access confidential information. The task is to develop advanced ML algorithms to detect four attacks such as DoS, Fuzzy, Replay, Impersonation attacks in real-time with a low false positive rate and a high recall rate. Tools: python, Pandas, scikit-learn, statistical models Level of effort: proposed a new anomaly detection algorithm using statistical methods; super accuracy and fast detection; superior to previous methods such as random forest, SVM, Deep Learning, etc; the method has been tested on different CAN messages, the implementation has been delivered for production.

Senior Data Scientist
Ottawa, Ontario, Canada
Data scientist, Solana networks Project: ML design for threat & intent detection system Context: intrusion detection systems (IDSs) are used for network security by reporting network attacks and triggering alerts. The issue is that IDSs have been observed to trigger thousands of alerts per day, and thus high false positive rate. This makes it extremely difficult for the analyst to correctly identify the true positives. ML supervised learning methods have been successfully applied for network security by reducing false positives raised by IDSs. Further, The task also aims to intent attacks. For example, periodic network traffics might imply potential adversarial behaviors from botnets, which later launch network attacks. Therefore, Intent detection can be achieved by developing start of the art periodicity detection techniques. Technology: Kubernetes, Elasticsearch, Python, Pandas, scikit-learn, Naïve Bayes, PCA, randomforest, decision tree, clustering; cosine distance, segment distance, FFT.

AI Developer
Ottawa, Ontario, Canada
AI developer, CRA Canada Project: AWS architecture for QnABot development and deployment Context: create and deploy QnABot with AWS Lex and AWS Alexa and ElasticSearch and Kendra and CloudWatch. Using CloudFormation previsions and manage stacks and resources specified in template code. Building responseBots with intents and slottype and fulfilment Lambda functions. Using AWS CDK and CloudFormation and Cloud9 for infrastructure as code; CodeCommit for version control; CodePipeline for continuous delivery and CodeBuild to create AWS projects. Additional tasks using AWS SageMaker and Rekognition for image and video analysis Tools: AWS, Boto3, Python, Node.js, angularJS, javascript, typescript

Data Scientist
Subject: Improvement of forecasting models by machine learning algorithm design with R Context: the data scientist team in JM Smucker has made much effort to build forecasting models to predict price and unit sales and build a shiny analysis flatform to help sale management. The challenge is how to improve forecast models while training sample contains much missing values and noise. As a result, forecasting models are subject to overestimation for prediction of price and unit sale amount, and thus suffer poor performance for deployment. Level of effort: fully responsible and independent for prototype design and implementation Tasks and technical details: • Built forecasting models using Linear regression PCR and nonlinear regression models such as randomforest; combine with PCA for machine learning pipeline • Developed a new learning architecture based on Bayesian learning; combine Bayesian learning and PCA; • Developed a new algorithm for improvement of forecasting models by learning Bayesian variables using Bayesian methods; • Research on pricing prediction using forecasting models • Developed methods and experiments to create training sample for complex learning tasks including pricing predictions

Principal Machine Learning Scientist
Greater Boston Area
Principal Machine Learning scientist/Engineer • Develop advanced AI technology for predictive maintenance with Porsche of Volkswagen Group; • Machine learning pipeline: automate data imputation and data encoding; stability selection/recursive feature elimination; creating training data for predictive maintenance using XML parser • Recent research: mixture of Bayes with gaussian and kernel density • Free overfitting such that no difference between prediction accuracy and training accuracy; both up to 100; PROJECT SUMMARY • As a Principal Machine Learning scientist/Engineer, design and implement Predictive maintenance machine learning system with Porsche of Volkswagen Group, Python, Scikit-learn, Pandas, XML, Matlibplot; • Develop mixture of Bayes algorithm, Python, scikit-learn, Bayes Theorem

Senior Data Scientist
FreedomMobile Inc
Toronto, Canada Area
The main research focusing on Bayes learning algorithms; published the recent work on Data Science and Big Data Analytics(DSBDA) 2018; Design and implement Customer Churn Analysis using machine learning technology such as LogisticRegression, RidgeClassifier, RBF SVM, Neural Network, TensorDNN, PCA, Machine learning pipeline for data imputation and categorical encoding and algorithmic learning; Developed effective methods for data imputation and unseen missed values for categorical encoding; Design and implement Add a Line for Telecom business analysis using Zeppelin Notebook, Spark/pyspark, MLLib, Hadoop;
Education

Computer Science
PhD student, Computer Science in University of Ottawa, Analyzing and designing intelligent system for Health Canada for monitoring nuclear explosion and pollution. Using Machine Learning and Data Mining methods for designing and implementing this intelligent system; The research was done with two year research contract with Government Health Canada using C/C++, C#, J#, .Net, and Weka Machine Learning Package, and Eclipse, knowledge of EJB, WebSphere, Weblogic, AJAX, and SOAP.
Guichong Li's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.


