Bhavneet Singh
AI/ML Engineer – Lead Architect for LLM & Agentic Systems @ National Research Council Canada / Conseil national de recherches Canada
About
I am an AI Engineer specializing in building and operating production-scale Generative AI and machine learning systems, with expertise in LLMs, RAG, and distributed inference. My focus is on transforming advanced AI models into scalable, low-latency systems that deliver measurable real-world impact. At the National Research Council Canada (NRC), I architect and deploy production AI platforms: • Architected Graph-RAG-based multi-agent system enabling distributed inference over 150K+ documents, supporting 40+ stakeholders. • Developed distributed LLM microservices using FastAPI, vector databases, and cloud infrastructure (AWS, GCP), enabling scalable real-time inference • Designed multimodal AI pipelines achieving sub-300 ms real-time inference, optimized for production reliability and performance I specialize in designing end-to-end AI systems integrating model development, deployment, and infrastructure using PyTorch, HuggingFace, Docker, Kubernetes, MLflow, and cloud platforms (AWS, GCP, Azure), with a strong focus on scalability, observability, and production reliability. I am also the creator of RiskReg, an open-source rare-event regression framework that reduced prediction error by up to 3× across 34 real-world datasets. I focus on building production-ready AI systems that combine LLMs, MLOps, and scalable infrastructure to deliver reliable, high-performance real-world solutions.
Canada
Ottawa
Computer Software
Kubernetes, Generative AI • Large Language Models (LLMs) • Retrieval-Augmented Generation (RAG), Machine Learning • PyTorch, MLOps • LLMOps, AWS • GCP • Azure, Generative AI • LLMs • RAG • Machine Learning • PyTorch • Python • MLOps, LLMOps • AWS, GCP, Azure, TypeScript, MCP (Model Context Protocol), AzureML, Automated Machine Learning (AutoML), GPT-5, LangGraph, Data Engineering, Software Development Life Cycle (SDLC), Terraform, Large Language Models (LLMs), MLOps & MLflow, Cloud Computing (AWS | GCP | Azure), Full-Stack Development (React + Node.js), PyTorch & Deep Learning
Experience

AI/ML Engineer – Lead Architect for LLM & Agentic Systems
Ottawa, ON
• Architected a Graph-RAG multi-agent system enabling distributed inference over 150K+ documents, supporting 40+ stakeholders. • Built multi-agent orchestration (LangGraph + CrewAI + MCP + Ray) with policy routing, improving decision accuracy by 26%. • Implemented hybrid retrieval (MiniLM-L6-v2 + ColBERT) with feedback loops, improving retrieval precision@k by 31%. • Optimized FAISS, Milvus, and PostgreSQL vector stores to index 3M+ legal docs with sub-second retrieval latency. • Deployed LLM inference on vLLM + Triton, enabling multi-tenant serving with predictable latency/memory-efficient utilization. • Designed LLM microservices (FastAPI + ONNX + CUDA graphs) achieving <400ms inference under concurrent production load. • Applied VLMs (CLIP, BLIP-2) with hybrid retrieval and contextual reasoning, improving real-time inference accuracy by 18%. • Fine-tuned LLMs/vision models using LoRA-based tuning and quantization, boosting accuracy by 24% and reducing size by 48%. • Developed distributed ML pipelines (PySpark + Spark + PyTorch) over multi-node clusters, improving throughput by 65%. • Built real-time ingestion and embedding pipelines with PII filtering/schema validation, accelerating ETL throughput by 52%. • Integrated LangSmith + MLflow + W&B observability, reducing debugging cycles by 40% and standardizing RAG quality metrics. • Automated drift detection, retraining loops, and lineage tracking via Databricks, reducing performance degradation by 33%. • Implemented token tracking, prompt versioning, and cost monitoring to control LLM spend and predictable production scaling.

Teaching Assistant
Ottawa, ON
Teaching Assistant/Demonstrator/Lab Monitor -Intro to Software Engineering and Data Science • Supervising Labs on Software Engineering and Android concepts in Java, grading Assignments, Proctoring Exams, and Conducting Office Hours for doubt resolution. • Mentored labs in Data Science, guiding students in Python programming, and automated ML using pandas and scikit-learn.

Agentic AI Developer - IBM Cyber Range
Ottawa, Ontario, Canada
• Engineered an agentic-AI benchmarking framework that autonomously identifies and corrects regression imbalance across 34 real-world datasets spanning cybersecurity, climate, housing, and healthcare. • Integrated LangChain and AutoGen agents for adaptive data relevance estimation and self-driven model retraining workflows. • Developed autonomous analysis agents leveraging φ-weighted metrics (SERA, SERT, φ-RMSE/MAE) and SMOGN-based oversampling, reducing rare-event prediction error by up to 3×. • Designed CrewAI-based multi-agent pipelines coordinating preprocessing, hyperparameter tuning, and diagnostics, improving runtime efficiency by 45%. • Implemented AI reasoning modules using KDE, GMM, and spectral density estimators for self-assessment of model reliability in high-impact regions. • Created bias-detection and interpretability agents leveraging HuggingFace Transformers, uncovering 60% of previously missed rare outcomes. • Automated research reporting through an LLM-powered summarization agent (LangChain + LaTeX), cutting documentation time by 70%. • Released as a PyPI-versioned, CI/CD-enabled package with reproducible configs and agent orchestration templates, adopted by multiple AI research labs.

AI/ML Engineer (Research & Innovation Division)
New Delhi
• Engineered agentic AI pipelines for context-aware orchestration across varied multi-step workflows supporting 10+ clients. • Developed RAG workflows for document search using HNSW-based ANN indexing with Pinecone, improving retrieval precision by 22%. • Built end-to-end data pipelines using Python, AWS, and MLflow for transformer-based models, improving accuracy by 12%. • Implemented a batch preprocessing pipeline for large document inputs, reducing parsing time by 8% and improving backend performance. • Reduced inference time by 30% via a custom FFT-based model, optimized for real-time text categorization. • Processed 500K+ document requests with REST APIs on AWS EC2, sustaining <200ms latency. • Developed React.js dashboard for admins, simplifying monitoring & results visualization. • Integrated CloudWatch monitoring for ML endpoints, cutting debugging time by 60% and improving uptime. • Automated ML infrastructure CI/CD with GitHub Actions and Terraform, supporting scalable deployment across client systems.

Machine Learning Specialist
North Rhine-Westphalia, Germany
• Improved data retrieval speed by 35% via optimized SQL joins, indexed queries, and Redis caching in the Express.js service layer. • Optimized backend services with FastAPI and MongoDB, enabling fast retrieval of dynamic datasets and improving data access performance by 25%. • Reduced dashboard load latency by 20% using streaming visualizations with React.js, D3.js, and real-time backend event handling via WebSockets. • Automated ML training workflows on AWS SageMaker for 2.5M+ records, enabling smooth backend integration and scalability. • Performed statistical analysis and dimensionality reduction (PCA, Bayesian inference) to optimize ML model inputs. • Delivered containerized Python APIs for 3 production ML models via MongoDB and Docker, reducing deployment friction by 40%.

Machine Learning Engineer
Pune, Maharashtra, India
• Automated ingestion and preprocessing of large-scale geospatial climate datasets via Google Earth Engine, transforming and cleaning data using PostgreSQL and pandas for ML-ready pipelines. • Built and trained hybrid CNN–LSTM models in TensorFlow to classify crop outcomes from time-series climate data, achieving 91% precision and 90.1% recall. • Accelerated model convergence by 30% through hyperparameter optimization (batch size, learning rate) on AWS EC2 GPU instances. • Implemented robust model diagnostics and monitoring with TensorBoard, AUC–ROC, and confusion matrices to ensure sustained accuracy across training cycles. • Deployed trained deep learning models through Flask REST APIs with integrated logging and exception handling, enabling research teams to perform live forecasting and model validation.
Bhavneet Singh's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.



