Justin Tang
AI Engineer @ Manulife Wealth & Asset Management
About
Machine Learning Specialist with 20+ years of engineering experience. I ship production-grade AI systems and high-performance backends that solve critical business problems and deliver measurable results. High-Performance Backends: Architected SLO-driven services that reduced API tail latency by 80% (P99 from 1.0s to 200ms) while maintaining 99.9%+ availability. Production-Grade AI: Shipped a HIPAA/PIPEDA-compliant EHR CoPilot using RAG/NLP that increased clinician efficiency by 40%, providing safe and auditable responses from sensitive data. Robust MLOps & Governance: Established the complete ML lifecycle for model deployment, rigorous evaluation, and drift monitoring to ensure AI solutions are reliable, compliant, and have clear rollback paths.
Canada
Richmond Hill
Financial Services
Transformers, RNN, Fine Tuning, NLP, Multi-agent Systems, MCP, AWS Serverless (Lambda, API Gateway, DynamoDB), Amazon SageMaker (training/serving, Model Monitor), RAG (Retrieval-Augmented Generation), Apache Spark (ETL & feature pipelines), AWS SageMaker, Retrieval-Augmented Generation (RAG), Transformers (BERT/Bio_ClinicalBERT), MLOps, Explainable AI (SHAP), Java (Spring Boot), Data Governance (Apache Ranger / Apache Atlas), Chatbot Development, React, HBase
Experience

Senior Software & ML Engineer | Cloud • RAG/NLP | Clinical SaaS & Data Platforms
Toronto, Ontario, Canada
Architected highly available SaaS and production clinical AI on AWS. Grew from leading core product development to delivering RAG/NLP solutions over EHR data with strong SLOs, security, and measurable outcomes. Products: EHR CoPilot (RAG over EHR with citations), Risk Stratification Service (SageMaker XGBoost), Clinical SaaS platform (React + AWS), Glue/S3 data platform (millions/day ingestion). Key achievements Reduced tail latency by 80%: P99 improved from 1.0s to 200ms via cache-aside (Redis/ElastiCache), DynamoDB key/GSI redesign, and query refactors; typical P95 < 300ms. Delivered production RAG and clinical NLP: HIPAA/PIPEDA-compliant EHR CoPilot (Bio_ClinicalBERT + LLM, citations, safety filters) enabling ~40% faster clinician chart navigation. Shipped interpretable risk model: SageMaker XGBoost with Bayesian HPO achieving AUC 0.872 and precision@10% 0.713; pilot cohort saw a 12.7% reduction in ER visits. Implemented zero-downtime releases: Serverless backend on Lambda + API Gateway with blue-green and canary deployments, automated rollbacks, and 99.9% availability. Strengthened security and compliance: Cognito/Keycloak SSO, RBAC, KMS encryption, PHI de-identification, and audit trails; ~15% faster compliance reviews. Built data platform at scale: AWS Glue/S3/Athena pipelines ingesting millions of records per day; data contracts and lineage; BI queries P95 < 8s; internal freshness SLO T+2 hours. Established observability: Datadog APM/logs/traces with SLO burn-rate monitors; ~60% fewer noisy pages and 35–45% faster MTTR. Mentored team of three engineers: 18% reduction in critical bugs and 10% faster feature delivery through design reviews, testing strategy, and coding standards. Tech: AWS (Lambda, API Gateway, DynamoDB, MSK/Kafka, OpenSearch, S3, Glue, Athena, Step Functions), Node/TypeScript, Java, Python (ML), React/TypeScript, SageMaker, Vertex AI (TPU), Redis/ElastiCache, Datadog, OpenTelemetry, Terraform/SAM/CDK, Cognito/Keycloak.

Senior Software Engineer | Distributed Systems & Data Platforms
Beijing, China
Led a national-scale service desk for an energy major and an enterprise big-data platform for a top life insurer. Products: PetroChina Enterprise Service Desk (Spring/Oracle/Redis + CTI, 400k+ users, 99.95% uptime); New China Life Insurance Enterprise Data Platform (100-node Hadoop DW, 30M+ customers, 50+ TB). Built a 100-node Hadoop DW (HDFS/YARN) with Hive/Impala and Spark ETL via Oozie; standardized Oracle ingestion with Sqoop; served 50+ analysts. Cut reporting from 168h to 4h (~97.6%) and established a single source of truth. Improved ETL throughput ~60% via partitioning/bucketing, predicate pushdown, join plan tuning, and small-file compaction; enabled the first predictive policy-lapse model. Established governance and auditability: Kerberos/Ranger auth, Apache Atlas lineage, and data contracts with reconciliation to the system of record. For external reports, adopted a governed T+2 days publish window (automated validation, reconciliation, audit sign-off); internal marts remained near-real-time. Raised BI performance: migrated to columnar Parquet/ORC, enforced partition pruning, and added precomputed aggregates; tuned Impala resource pools for fairness; BI scan P95 < 8s. Delivered the PetroChina service desk: Spring Boot microservices with Oracle + Redis and WebSocket updates; Huawei AICC CTI (screen-pop, click-to-dial, skill routing). Achieved 99.95% uptime; chatbot deflected ~30% of Tier-0 tickets and reduced MTTR. Tech: Java (J2EE, Spring Boot), Oracle, Redis, WebSocket, Apache Hadoop (HDFS/YARN), Apache Spark, Hive/Impala, Parquet/ORC, Sqoop, Oozie, Apache Atlas, Apache Ranger, Kerberos, Prometheus/Grafana.
Education
Justin Tang's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.




