Rehan Malik
Senior Generative AI Engineer @ Mars
About
6+ years architecting production GenAI systems that work at scale from autonomous agentic AI platforms to RAG pipelines across 2TB+ enterprise document repositories to real-time voice AI engines processing 500+ concurrent calls at zero latency.I've shipped $1M+ Generative AI transformations at Mars for regulated environments, built enterprise RAG architectures at scale, and deployed fine-tuned LLMs that cut API costs by 40%. Every system I build is production-grade: low-latency, cost-efficient, and governed for compliance.Key Outcomes:→ Reduced LLM API costs by 40% via LoRA/QLoRA/PEFT fine-tuning on open-source models→ Cut model inference latency by 25% with Python/CUDA + gRPC model serving→ Scaled AI infrastructure to handle 2TB+ datasets (+60% processing capacity)→ Built autonomous agentic AI platforms (LangChain Agents, LLM orchestration)→ Deployed enterprise RAG systems (FAISS, ChromaDB, PGVector) at scale→ Architected real-time voice AI on Kafka + WebSocket streaming (500+ concurrent calls)At Reallytics.ai, I lead the engineering team building these systems for enterprise clients. Previously: Sr GenAI Engineer at Mars ($1M+ regulated AI), Cloud Solution Architect at Cloud Kinetics (US enterprise), Sr ML Engineer at Afiniti ($200K+/month revenue systems).GenAI & LLMs: LangChain, LangGraph, CrewAI, RAG, Agentic AI, OpenAI, Claude/Anthropic, LLaMA, Mistral, Hugging Face, VLLM, Fine-Tuning (LoRA, QLoRA, PEFT), RLHF/PPO, Prompt EngineeringVector Search: FAISS, ChromaDB, Pinecone, PGVectorCloud & Infra: AWS (SageMaker, ECS, Lambda, Bedrock), Azure OpenAI, GCP, Kubernetes, Docker, TerraformData & Serving: Python, Kafka, PySpark, Airflow, dbt, FastAPI, gRPC, WebSockets, PostgreSQL, MongoDB, RedisML & Deep Learning: PyTorch, TensorFlow, OpenCV, scikit-learn, C++ (CUDA, Eigen)Master's in AI from LUMS. Building GenAI systems that don't break at scale.Open to: Senior AI Engineer, GenAI Engineer, AI Architect, and AI Infrastructure roles.Let's talk: rehanmalil99@gmail.com | DM open.
Pakistan
Lahore
Information Technology & Services
CrewAI, Hugging Face Products, LangGraph, Real-time Monitoring, ETL Pipelines, Model Training, Feature Engineering, Data Quality, Speech Recognition, Infrastructure as code (IaC), Data Pipeline Design, Solution Architecture, ChromaDB, FAISS, Reinforcement Learning Human Feedback (RLHF), Model Serving / ML Inference, Model Evaluation, Vector Databases, AI Safety, Fine-Tuning (LLMs)
Experience

Senior Generative AI Engineer
Led a $1M+ Generative AI enterprise transformation at Mars, architecting production GenAI systems for regulated environments from LLM orchestration design through production deployment with full governance. → Architected enterprise RAG pipelines using LangChain and vector search (FAISS, PGVector), grounding LLM outputs in domain-specific knowledge bases to eliminate hallucination in regulated workflows → Designed and implemented GenAI orchestration patterns including multi-agent frameworks, chain-of-thought workflows, and knowledge grounding mechanisms aligned with enterprise cloud architecture requirements → Built audit-ready data lineage and governance frameworks for GenAI deliverables, ensuring traceability and compliance standards required for regulated industries (healthcare, finance) → Led human-in-the-loop validation sessions with enterprise stakeholders, achieving 100% safe adoption of AI-generated outputs in production workflows → Evaluated and selected LLM providers (OpenAI, Anthropic, open-source) based on accuracy, cost, and compliance requirements reducing vendor lock-in risk while maintaining output quality → Collaborated across cross-functional business and engineering teams to translate ambiguous business requirements into structured GenAI technical specifications

Head of Artificial Intelligence
Punjab, Pakistan
Architecting and deploying production GenAI execution engines and AI infrastructure for enterprise clients. Leading engineering team, technical strategy, and product architecture for an AI-first startup. → Designed and deployed autonomous agentic AI platforms powered by LangChain Agents and LLM orchestration executing complex end-to-end workflows (SEO, content, analytics) without human intervention → Built real-time concurrent voice AI infrastructure: zero-latency voice-to-data ingestion engines processing 500+ simultaneous calls using WebSockets, Kafka, and streaming architectures → Architected enterprise RAG pipelines using LangChain, FAISS, and ChromaDB for knowledge retrieval across 2TB+ structured and unstructured data repositories → Scaled AI cloud infrastructure to handle 2TB+ datasets, boosting system processing capacity by 60% through optimized pipeline architecture and data partitioning strategies → Built high-performance model serving layer with C++/CUDA modules (Eigen) and gRPC microservices, reducing inference latency by 25% and supporting 15% more concurrent users → Implemented RLHF frameworks using Proximal Policy Optimization (PPO) to improve LLM output quality and reliability for production GenAI products → Deployed fine-tuned LLMs using LoRA, QLoRA, and PEFT on AWS SageMaker, cutting hosted API costs by 40% while maintaining accuracy thresholds → Built real-time audio AI pipelines for speech-to-text, sentiment analysis, and conversational intelligence using streaming architectures on AWS → Set technical vision and engineering roadmap for the company's AI product suite, aligning architecture decisions with business growth targets

Cloud Solution Architect
Designed cloud-native AI solutions and GenAI automation architectures for US and North American enterprise clients, translating complex business problems into production-ready AI systems. → Architected and deployed AI automation tools and intelligent data pipelines on AWS, Azure, and GCP integrating LLM-powered workflows with existing enterprise ecosystems → Designed cloud infrastructure for AI/ML workloads including model training environments, inference endpoints, and automated retraining pipelines → Built custom GenAI-powered automation strategies replacing manual business processes, including intelligent document processing, automated reporting, and AI-driven decision support systems → Managed weekly architecture reviews with enterprise clients, aligning AI adoption roadmaps with business KPIs and cloud infrastructure scaling plans → Maintained enterprise-grade security and performance standards across all AI cloud deployments while reducing manual operational overhead by 50%+

Senior ML/AI Engineer
Built production GenAI systems and LLM infrastructure for enterprise clients from RAG pipeline architecture through fine-tuned model deployment and serving at scale. → Architected scalable RAG pipelines using LangChain, FAISS, and ChromaDB for context-aware enterprise search across 2TB+ document repositories, reducing manual research time by 70% → Deployed fine-tuned LLMs (LLaMA, Mistral) using LoRA, QLoRA, and PEFT on AWS SageMaker, reducing API costs by 40% versus hosted model providers while maintaining accuracy → Built agentic AI workflows connecting LLMs to internal APIs, databases, and messaging platforms for autonomous task execution — eliminating manual intervention in multi-step business processes → Served models at production scale using VLLM with optimized GPU utilization via CUDA, containerized with Docker on AWS ECS/ECR for auto-scaling inference → Integrated C++ modules into inference pipelines, lowering model latency by 25% and enabling real-time conversational AI with 35% faster response times → Designed speech-to-text and audio AI processing pipelines for real-time streaming use cases, powering conversational intelligence features → Built high-performance APIs using FastAPI and gRPC powering conversational AI systems backed by OpenAI models, increasing user satisfaction by 30%

Senior Machine Learning Engineer
Built and managed production ML systems and AI-driven prediction engines for enterprise accounts generating $200K+ monthly revenue. → Designed and optimized production ML pipelines processing real-time data at scale, ensuring model accuracy, pipeline reliability, and automated retraining for $1M+ client accounts → Built feature engineering and data generation workflows for ML model training, improving prediction accuracy and reducing time-to-production for new model iterations → Developed real-time AI monitoring dashboards using SQL, PostgreSQL, Grafana, and Tableau — enabling proactive model drift detection and performance alerting → Evaluated and benchmarked model performance against client KPIs, delivering data-driven recommendations that directly impacted revenue outcomes → Operated in a 24/7 production ML environment, collaborating across engineering and data science teams via Jira and Slack for rapid incident resolution

AI Product Engineer
Production AI engineer supporting enterprise ML systems, real-time prediction models, and AI-driven call routing in a 24/7 mission-critical environment. → Maintained and monitored production ML pipelines and AI inference systems, ensuring 99.5%+ uptime and data integrity across enterprise AI platforms → Built churn risk prediction models integrated into AI-powered call routing logic, optimizing sales agent assignment to maximize conversion rates and revenue → Conducted real-time AI system monitoring using SQL, PostgreSQL, Grafana, and Tableau — identifying data quality issues before they impacted model accuracy → Managed incident response workflows via JIRA, collaborating cross-functionally to resolve AI pipeline failures with minimal business impact

Junior python Engineer
United States
Built data pipelines and applied ML models to IoT sensor data for predictive analytics across thousands of connected devices. → Designed ingestion pipelines processing IoT sensor telemetry data at scale using Python, PostgreSQL, and real-time streaming workflows → Applied machine learning techniques (clustering, anomaly detection, time-series forecasting) to IoT data for predictive maintenance insights → Built interactive dashboards in Power BI for executive reporting and AI-driven decision support
Education
Rehan Malik's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.



