Rehan Malik

Rehan Malik

Senior Generative AI Engineer @ Mars

About

6+ years architecting production GenAI systems that work at scale from autonomous agentic AI platforms to RAG pipelines across 2TB+ enterprise document repositories to real-time voice AI engines processing 500+ concurrent calls at zero latency.I've shipped $1M+ Generative AI transformations at Mars for regulated environments, built enterprise RAG architectures at scale, and deployed fine-tuned LLMs that cut API costs by 40%. Every system I build is production-grade: low-latency, cost-efficient, and governed for compliance.Key Outcomes:→ Reduced LLM API costs by 40% via LoRA/QLoRA/PEFT fine-tuning on open-source models→ Cut model inference latency by 25% with Python/CUDA + gRPC model serving→ Scaled AI infrastructure to handle 2TB+ datasets (+60% processing capacity)→ Built autonomous agentic AI platforms (LangChain Agents, LLM orchestration)→ Deployed enterprise RAG systems (FAISS, ChromaDB, PGVector) at scale→ Architected real-time voice AI on Kafka + WebSocket streaming (500+ concurrent calls)At Reallytics.ai, I lead the engineering team building these systems for enterprise clients. Previously: Sr GenAI Engineer at Mars ($1M+ regulated AI), Cloud Solution Architect at Cloud Kinetics (US enterprise), Sr ML Engineer at Afiniti ($200K+/month revenue systems).GenAI & LLMs: LangChain, LangGraph, CrewAI, RAG, Agentic AI, OpenAI, Claude/Anthropic, LLaMA, Mistral, Hugging Face, VLLM, Fine-Tuning (LoRA, QLoRA, PEFT), RLHF/PPO, Prompt EngineeringVector Search: FAISS, ChromaDB, Pinecone, PGVectorCloud & Infra: AWS (SageMaker, ECS, Lambda, Bedrock), Azure OpenAI, GCP, Kubernetes, Docker, TerraformData & Serving: Python, Kafka, PySpark, Airflow, dbt, FastAPI, gRPC, WebSockets, PostgreSQL, MongoDB, RedisML & Deep Learning: PyTorch, TensorFlow, OpenCV, scikit-learn, C++ (CUDA, Eigen)Master's in AI from LUMS. Building GenAI systems that don't break at scale.Open to: Senior AI Engineer, GenAI Engineer, AI Architect, and AI Infrastructure roles.Let's talk: rehanmalil99@gmail.com | DM open.

Country

Pakistan

City

Lahore

Industry

Information Technology & Services

Skill

CrewAI, Hugging Face Products, LangGraph, Real-time Monitoring, ETL Pipelines, Model Training, Feature Engineering, Data Quality, Speech Recognition, Infrastructure as code (IaC), Data Pipeline Design, Solution Architecture, ChromaDB, FAISS, Reinforcement Learning Human Feedback (RLHF), Model Serving / ML Inference, Model Evaluation, Vector Databases, AI Safety, Fine-Tuning (LLMs)

Experience

Mars

Senior Generative AI Engineer

Mars

LinkedIn
2024-10 - Present · 2 yrs

Led a $1M+ Generative AI enterprise transformation at Mars, architecting production GenAI systems for regulated environments from LLM orchestration design through production deployment with full governance. → Architected enterprise RAG pipelines using LangChain and vector search (FAISS, PGVector), grounding LLM outputs in domain-specific knowledge bases to eliminate hallucination in regulated workflows → Designed and implemented GenAI orchestration patterns including multi-agent frameworks, chain-of-thought workflows, and knowledge grounding mechanisms aligned with enterprise cloud architecture requirements → Built audit-ready data lineage and governance frameworks for GenAI deliverables, ensuring traceability and compliance standards required for regulated industries (healthcare, finance) → Led human-in-the-loop validation sessions with enterprise stakeholders, achieving 100% safe adoption of AI-generated outputs in production workflows → Evaluated and selected LLM providers (OpenAI, Anthropic, open-source) based on accuracy, cost, and compliance requirements reducing vendor lock-in risk while maintaining output quality → Collaborated across cross-functional business and engineering teams to translate ambiguous business requirements into structured GenAI technical specifications

reallytics.ai

Head of Artificial Intelligence

reallytics.ai

LinkedIn
2022-10 - Present · 4 yrs

Punjab, Pakistan

Architecting and deploying production GenAI execution engines and AI infrastructure for enterprise clients. Leading engineering team, technical strategy, and product architecture for an AI-first startup. → Designed and deployed autonomous agentic AI platforms powered by LangChain Agents and LLM orchestration executing complex end-to-end workflows (SEO, content, analytics) without human intervention → Built real-time concurrent voice AI infrastructure: zero-latency voice-to-data ingestion engines processing 500+ simultaneous calls using WebSockets, Kafka, and streaming architectures → Architected enterprise RAG pipelines using LangChain, FAISS, and ChromaDB for knowledge retrieval across 2TB+ structured and unstructured data repositories → Scaled AI cloud infrastructure to handle 2TB+ datasets, boosting system processing capacity by 60% through optimized pipeline architecture and data partitioning strategies → Built high-performance model serving layer with C++/CUDA modules (Eigen) and gRPC microservices, reducing inference latency by 25% and supporting 15% more concurrent users → Implemented RLHF frameworks using Proximal Policy Optimization (PPO) to improve LLM output quality and reliability for production GenAI products → Deployed fine-tuned LLMs using LoRA, QLoRA, and PEFT on AWS SageMaker, cutting hosted API costs by 40% while maintaining accuracy thresholds → Built real-time audio AI pipelines for speech-to-text, sentiment analysis, and conversational intelligence using streaming architectures on AWS → Set technical vision and engineering roadmap for the company's AI product suite, aligning architecture decisions with business growth targets

Cloud Kinetics USA

Cloud Solution Architect

Cloud Kinetics USA

LinkedIn
2024-8 - 2025-12 · 1 yr 5 mos

Designed cloud-native AI solutions and GenAI automation architectures for US and North American enterprise clients, translating complex business problems into production-ready AI systems. → Architected and deployed AI automation tools and intelligent data pipelines on AWS, Azure, and GCP integrating LLM-powered workflows with existing enterprise ecosystems → Designed cloud infrastructure for AI/ML workloads including model training environments, inference endpoints, and automated retraining pipelines → Built custom GenAI-powered automation strategies replacing manual business processes, including intelligent document processing, automated reporting, and AI-driven decision support systems → Managed weekly architecture reviews with enterprise clients, aligning AI adoption roadmaps with business KPIs and cloud infrastructure scaling plans → Maintained enterprise-grade security and performance standards across all AI cloud deployments while reducing manual operational overhead by 50%+

Verticiti

Senior ML/AI Engineer

Verticiti

LinkedIn
2023-12 - 2024-7 · 8 mos

Built production GenAI systems and LLM infrastructure for enterprise clients from RAG pipeline architecture through fine-tuned model deployment and serving at scale. → Architected scalable RAG pipelines using LangChain, FAISS, and ChromaDB for context-aware enterprise search across 2TB+ document repositories, reducing manual research time by 70% → Deployed fine-tuned LLMs (LLaMA, Mistral) using LoRA, QLoRA, and PEFT on AWS SageMaker, reducing API costs by 40% versus hosted model providers while maintaining accuracy → Built agentic AI workflows connecting LLMs to internal APIs, databases, and messaging platforms for autonomous task execution — eliminating manual intervention in multi-step business processes → Served models at production scale using VLLM with optimized GPU utilization via CUDA, containerized with Docker on AWS ECS/ECR for auto-scaling inference → Integrated C++ modules into inference pipelines, lowering model latency by 25% and enabling real-time conversational AI with 35% faster response times → Designed speech-to-text and audio AI processing pipelines for real-time streaming use cases, powering conversational intelligence features → Built high-performance APIs using FastAPI and gRPC powering conversational AI systems backed by OpenAI models, increasing user satisfaction by 30%

Afiniti

Senior Machine Learning Engineer

Afiniti

LinkedIn
2022-10 - 2023-11 · 1 yr 2 mos

Built and managed production ML systems and AI-driven prediction engines for enterprise accounts generating $200K+ monthly revenue. → Designed and optimized production ML pipelines processing real-time data at scale, ensuring model accuracy, pipeline reliability, and automated retraining for $1M+ client accounts → Built feature engineering and data generation workflows for ML model training, improving prediction accuracy and reducing time-to-production for new model iterations → Developed real-time AI monitoring dashboards using SQL, PostgreSQL, Grafana, and Tableau — enabling proactive model drift detection and performance alerting → Evaluated and benchmarked model performance against client KPIs, delivering data-driven recommendations that directly impacted revenue outcomes → Operated in a 24/7 production ML environment, collaborating across engineering and data science teams via Jira and Slack for rapid incident resolution

Afiniti

AI Product Engineer

Afiniti

LinkedIn
2021-4 - 2022-10 · 1 yr 7 mos

Production AI engineer supporting enterprise ML systems, real-time prediction models, and AI-driven call routing in a 24/7 mission-critical environment. → Maintained and monitored production ML pipelines and AI inference systems, ensuring 99.5%+ uptime and data integrity across enterprise AI platforms → Built churn risk prediction models integrated into AI-powered call routing logic, optimizing sales agent assignment to maximize conversion rates and revenue → Conducted real-time AI system monitoring using SQL, PostgreSQL, Grafana, and Tableau — identifying data quality issues before they impacted model accuracy → Managed incident response workflows via JIRA, collaborating cross-functionally to resolve AI pipeline failures with minimal business impact

Merycure

Junior python Engineer

Merycure

LinkedIn
2020-2 - 2021-4 · 1 yr 3 mos

United States

Built data pipelines and applied ML models to IoT sensor data for predictive analytics across thousands of connected devices. → Designed ingestion pipelines processing IoT sensor telemetry data at scale using Python, PostgreSQL, and real-time streaming workflows → Applied machine learning techniques (clustering, anomaly detection, time-series forecasting) to IoT data for predictive maintenance insights → Built interactive dashboards in Power BI for executive reporting and AI-driven decision support

Education

Lahore University of Management Sciences

Lahore University of Management Sciences

LinkedIn

Artificial Intelligence

2018 - 2020 · 2 yrs

Master's Degree, Artificial Intelligence (2018 - 2020) Coursework: Machine Learning, Deep Learning, Natural Language Processing, Data Engineering, Distributed Systems, Statistical Modeling

COMSATS University Islamabad

COMSATS University Islamabad

LinkedIn

Computer Science

2014 - 2018 · 4 yrs

Bachelor's Degree, Computer Science (2014 - 2018) Coursework: Algorithms, Data Structures, Database Systems, Software Engineering, Operating Systems

Rehan Malik's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.