Khaja Amir Kafeel Mohammed

Khaja Amir Kafeel Mohammed

Staff AI Engineer @ Thomson Reuters

About

As a data scientist at RBC Capital Markets, I apply my skills and knowledge in artificial intelligence, AWS, SQL, and Python to design, develop, test, and implement various data-driven solutions for the financial sector. I work with a diverse and collaborative team of analysts, engineers, and business stakeholders to deliver innovative and impactful projects that enhance the performance and efficiency of the organization. I have a strong educational background in artificial intelligence and business analytics, with a Master of Engineering from the University of Waterloo and an IPBA from the Indian Institute of Management, Indore. I have also acquired multiple certifications from Microsoft, Stanford University, and Intel Corporation that demonstrate my proficiency and interest in various domains and technologies. I am passionate about learning new skills and staying updated with the latest trends and developments in the field of data science. My goal is to leverage my expertise and experience to create value and make a difference in the world of data.

Country

Canada

City

Toronto

Industry

Information Technology & Services

Skill

Claude Agent SDK, claudecode, Agent Core, GitHub Copilot, MCP, aGEMT, Large Language Model Operations (LLMOps), Retrieval-Augmented Generation (RAG), Data Build Tool (DBT), Snowflake, Azure DevOps Services, Azure Data Studio, Data Warehousing, Azure Data Factory, Azure SQL, Azure Databricks, REST APIs, Kubernetes, Microsoft Azure, Path finder

Experience

Thomson Reuters

Staff AI Engineer

Thomson Reuters

LinkedIn
2025-12 - Present · 10 mos

Toronto, ON

Designed multi-agent orchestration systems using LangGraph (StateGraph, subgraphs, memory store) with MCP server integrations for tool/resource exposure to Claude and other LLM runtimes. Built Claude Code plugins and MCP-based skills servers for compliance modernization and tax workflows, enabling LLM-driven regulatory document analysis, risk flagging, and audit trail generation. Engineered end-to-end LLM interaction pipelines using the Anthropic Messages API with structured output patterns, tool-use chaining, and streaming for high-volume production deployments

RBC Capital Markets

Lead Machine Learning Engineer

RBC Capital Markets

LinkedIn
2024-1 - 2025-12 · 2 yrs

Toronto, ON

● Designed and deployed 15+ RAG-based AI systems integrating LangChain, LlamaIndex, FAISS, and Sentence-BERT (SBERT) for semantic search, retrieval optimization, and knowledge graph augmentation. ● Integrated MCP Server Agents and AutoGen frameworks for orchestrating multi-agent workflows across SAP, Pega, and ServiceNow, enabling intelligent process automation and real-time decisioning. ● Built browser automation agents and Python orchestration scripts to enable AI systems to interact dynamically with internal and external web interfaces. ● Developed a self-service analytics assistant capable of generating SQL queries, creating data visualizations, and summarizing insights automatically reducing manual reporting time by 50%. ● Created entity resolution models (98% accuracy) and PII detection pipelines for structured and unstructured data, reinforcing compliance with enterprise data governance and privacy policies. ● Innovated ETL scheduling and reliability using Apache Airflow, integrating custom sensors and Kubernetes operators for optimized data pipelines. ● Engineered high-availability FastAPI microservices for real-time inference and data retrieval from KDB, equipped with caching, rate limiting, and distributed tracing for observability. ● Transitioned authentication systems to Azure AD from LDAP, significantly bolstering security and streamlining user management. ● Developed in-house SNS and Lambda functionalities within Kubernetes, reducing dependency on AWS and enhancing system flexibility.

RBC Capital Markets

Senior Data Scientist

RBC Capital Markets

LinkedIn
2023-3 - 2023-12 · 10 mos

Toronto, ON

● Developed interactive dashboards any other using ECharts and Dash Mantine Components, focusing on credit usage, limit monitoring, and rating downgrades among counterparties, thus providing advanced visualization tools for informed credit management. ● Engineered high-availability FastAPI microservices for real-time inference and data retrieval from KDB, equipped with caching, rate limiting, and distributed tracing for observability. ● Built an anomaly detection model using Isolation Forest to identify potential data exfiltration via email, leveraging feature deviations rather than raw content to assign risk scores. Successfully flags 84% of synthetic high-risk emails during testing.

Quantiphi

Data Engineer

Quantiphi

LinkedIn
2021-10 - 2022-12 · 1 yr 3 mos

Ontario, Canada

● Proposed and implemented an ETL and Machine Learning boilerplates, resulting in a 40% reduction in overall project timelines. ● Designed, developed, and orchestrated data pipelines for real-time and batch data processing using AWS Step Functions & Airflow. ● Skilled in using Pyspark (Data Bricks, EMR) to load various data formats (JSON, XML, CSV, parquet) into a data lake and database tables. ● Experienced in writing SQL Queries, Stored procedures, functions, packages, tables, views, triggers using relational database like PostgreSQL, and MS SQL server. ● Migrated on-premises servers to the cloud, which involved examining, improving, and converting 3000 + on-premises server objects to a compatible cloud database. ● Worked with 10+ APIs (Rest, SOAP), including both consuming data from external APIs and creating APIs to expose data for use by other systems. ●Hands on Amazon Web Services (AWS) for creating and managing EC2, Elastic Map Reduce, Elastic Load-balancers, Elastic Container Service (Docker Containers), S3, Lambda, Elastic File system, RDS, Cloud Watch, Cloud Trail, IAM and Kinesis Streams. ● Experienced with real-time data sources such as Spark Streaming, Kafka, AWS Kinesis, Firehouse, and using asyncio to efficiently ingest data from various sources, resulting in reduced overall ingestion time 30%. ● Implemented data quality checks and validation processes to ensure the accuracy and completeness of data before loading it into the data warehouse which helped in improving overall decision making. ●Managed AWS resource limits by conducting regular cleanup operations, ensuring optimal application performance within accounts. ●Pioneered the use of AWS Bucket Policies to manage cross-account access, enhancing collaboration while ensuring strict adherence to data privacy standards

Accenture

Application Development Analyst

Accenture

LinkedIn
2020-6 - 2021-8 · 1 yr 3 mos

Hyderabad, Telangana, India

● Built 10+ data pipelines on different cloud platforms AWS, Azure, GCP for ingesting, translating, and analyzing massive amounts of data. ● Expertise in machine learning methodology including model selection, training, and evaluation, as well as data preparation techniques such as cross-validation, sampling, and imputation models. ● Experienced in MLOps and deployment using VertexAI, SageMaker, Kubflow and EC2, with a focus on automating and optimizing the end-to-end machine learning workflow, including model training, deployment, monitoring data drift and model drift in production machine learning models., and maintenance. ●Created a versatile and customizable Python package for efficient data management and manipulation. ●Created web-based tool using Python, Flask, HTML, CSS, and JavaScript for production support, job execution, and real-time issue troubleshooting, with customizable deployment options, and permission management. ●Played a pivotal role in integrating advanced normalized database solutions within the AWS-based quantitative platform, enhancing the research team's ability to perform flexible analysis and back-testing on a vast array of investment data. ●Experienced in implementing machine learning methodologies, including model selection, training, and evaluation, as well as utilizing data preparation techniques such as cross-validation, sampling, and imputation models. ●Orchestrated the seamless integration of Assume Role policies for cross-service functionality, facilitating secure and efficient inter-service communications and data sharing ● Proficient in implementing Python REST API using Flask for integrating machine learning models and exposing them through API endpoints for real-time predictions. ● Expertise in deep learning techniques including ANN, CNN, RNN (LSTM) using TensorFlow, with hands-on experience in applying them to natural language processing and image classification tasks.

Accenture

Associate Software Engineer

Accenture

LinkedIn
2018-8 - 2020-6 · 1 yr 11 mos

Hyderabad Area, India

● Developed 5+ Python & Pyspark packages utilizing best practices for naming conventions and coding techniques to guarantee the consistency of data models. ● Created different KPI using calculated key figures and parameters to make report dynamic and reduced no. of reports by 20%. ● Experience with experiment tracking using tools such as MLflow and DVC for keeping track of model performance and reproducibility, with a focus on efficient and organized data management. ● Excellent understanding of machine learning techniques and algorithms, such as Logistic Regression, SVM, Random Forests, Deep Learning, etc. ● Conducted sessions for up to 20-member professional team in upskilling on AWS, Python programming.

KPMG US

Data Analytics Consulting Virtual Internship

KPMG US

LinkedIn
2020-5 - 2020-5 · 1 mo

» Open access program designed to work with big data and learn the processes in producing business results » Validated data quality in preparation for analysis based on criteria of completeness, accuracy, and consistency » Utilized regression analysis to target high value customers based on customer demographics and attributes » Implemented visualizations to present data insights in Excel charts based on regression analysis

Education

University of Waterloo

University of Waterloo

LinkedIn

Artificial Intelligence

2021-9 - 2022-12 · 1 yr 4 mos

Masters in Artificial Intelligence & Machine Learning - Advance Topics on statistics and probability, including probability distributions, statistical inference, and hypothesis testing - Techniques for data exploration, visualization, and cleaning - Methods for modeling and analyzing data, including linear and logistic regression, decision trees, clustering and Deep learning (CNN ,RNN,LSTM) - Tools and technologies for working with data, such as SQL, Python, and R - Machine learning algorithms, such as supervised and unsupervised learning - Applications of data science, such as natural language processing, computer vision, and time series analysis - Communication skills and the ability to present findings and insights to a non-technical audience - Ethics and privacy considerations related to working with data.

Indian Institute of Management, Indore

Indian Institute of Management, Indore

LinkedIn

Business Analytics

2020 - 2021 · 1 yr

The Integrated Program in Business Analytics (IPBA) curriculum is designed to be a comprehensive program covering technical and business aspects of the application of analytics and data science. It starts by laying a strong foundation of essential tools and techniques including descriptive and inferential statistics, data extraction and manipulation with SQL, data manipulation and processing with Python and R and data visualization with Tableau. Students will also have access to Tableau's cutting-edge visual analytics software for free while enrolled in the program for study and practice purposes, through Tableau’s Global Academic Program. The bulk of the program builds on this foundation by extending the analysis capabilities to predictive models using statistical modelling and machine learning and covering data types that are both traditional and structured, to more unstructured types including text and images.

Osmania University

Osmania University

LinkedIn

Electronics and Communications Engineering

2014 - 2018 · 4 yrs

Khaja Amir Kafeel Mohammed's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.