
Bhavana Akula
AI / ML engineer @ Scale AI
About
AI/ML Engineer with expertise in building production-grade Generative AI, RAG, and Conversational AI systems at scale across enterprise, fintech, and multilingual voice-assistant environments. Proven success deploying low-latency LLM platforms, optimizing distributed training, and enabling secure, compliant AI adoption in mission-critical settings. Experience spans Scale AI, Stripe, and Amazon, delivering advanced LLM systems with multilingual capabilities, streaming context, fraud-aware reasoning, human-feedback loops, and automated model evaluation. Passion for combining applied research, scalable engineering, and MLOps to move cutting-edge AI from prototype to production. Specialties include: LLMs with RAG, PEFT/QLoRA, multi-agent architectures Distributed fine-tuning and GPU-accelerated pipelines Real-time semantic retrieval over millions of documents Streaming AI with Kafka/Kinesis + vector DBs (FAISS, Pinecone) Evaluation, drift detection, and human-in-the-loop systems Security and compliance for regulated industries (SOC2, GDPR, PCI-DSS) Ultra-low-latency inference with ONNX, TorchScript, quantization Tech Stack: PyTorch • Hugging Face • LangChain • SageMaker • Kubernetes • MLflow • FastAPI • Snowflake • FAISS • Pinecone Focus areas: Real-time AI products that are fast, safe, and scalable Responsible AI and secure enterprise deployment Multilingual NLP & low-resource language innovation Continuous improvement via automated feedback and evaluation loops Open to collaborations and conversations on next-gen LLM systems, enterprise AI platforms, and applied research that pushes the boundaries of multimodal and multilingual intelligence
United States
San Jose
Computer Software
Generative AI, MLOps, Retrieval-Augmented Generation (RAG), Machine learning engineering, Large Language Models (LLM), Event-driven ML workflows, Telemetry-triggered retraining, CI/CD — AWS CodePipeline, Docker, Terraform, Blue-green deployments & rollback automation, TorchServe, AWS ECS / EKS / Lambda, Sub-100ms multilingual inference at scale, AWS EC2, Lambda, Kinesis Streams, SageMaker, CloudWatch, CloudWatch, SageMaker Model Monitor, Self-healing pipelines & alerts, Differential privacy, Fairness in multilingual systems, India data protection compliance
Experience

AI / ML engineer
California, United States
• Partnered with cross-functional teams (operations, compliance, and internal stakeholders) to identify high-impact AI use cases and translate business workflows into scalable LLM-powered applications • Led end-to-end development of a Retrieval-Augmented Generation (RAG) platform enabling enterprise users to query internal knowledge bases using natural language, improving information accessibility and decision-making • Designed and optimized prompt strategies and evaluation pipelines using human feedback and automated scoring, improving response accuracy and reliability in production • Built AI-powered automation workflows and agent-based systems to handle multi-step tasks, reducing manual effort and improving operational efficiency • Developed scalable document ingestion, embedding, and semantic retrieval pipelines using FAISS, supporting real-time responses across millions of documents • Deployed production-grade FastAPI microservices with low-latency inference (<200ms), ensuring reliable performance under high user load • Integrated AI systems with enterprise APIs and internal platforms, enabling seamless adoption across business teams • Implemented responsible AI practices including prompt validation, output filtering, and data privacy safeguards to ensure secure and compliant AI usage • Conducted A/B testing and iterative improvements based on user feedback, increasing system accuracy and user satisfaction

Data Scientist
India
• Engineered and deployed a multilingual Alexa conversational assistant using mBERT, ROBERTa, and internal GPT-style transformer models with PyTorch Lightning and Hugging Face Transformers, enabling real-time multilingual dialogue understanding and improving response accuracy by 28% across Hindi, Tamil, and Telugu. • Fine-tuned transformer models with parameter-efficient tuning (LoRA, PEFT, adapters) using DeepSpeed on AWS SageMaker Distributed Training, reducing model fine-tuning cycles by 35% and supporting on-demand retraining triggered by live telemetry signals. • Built and integrated low-latency semantic retrieval pipelines using FAISS (HNSW indexing) on AWS EC2, Lambda, and Kinesis Streams, supporting instant contextual retrieval and intent disambiguation during real-time Alexa interactions, improving CTR by 22% in A/B experiments. • Designed event-driven ML pipelines with Apache Airflow (MWAA), AWS Step Functions, and SageMaker Pipelines, automating data ingestion, retraining, and deployment in response to user interaction feedback, data drift, and latency thresholds, reducing manual intervention by 60%. • Implemented CI/CD pipelines using AWS CodePipeline, Docker, and Terraform, enabling continuous integration, automated model validation, and blue-green deployments across multiple Alexa environments with zero downtime and built-in rollback. • Deployed real-time inference microservices with TorchServe and ONNX Runtime on AWS ECS, EKS, and Lambda, achieving sub-100ms response times for multilingual Alexa skills and fault-tolerant scaling across millions of concurrent users • Established observability and proactive monitoring with CloudWatch, Prometheus, and SageMaker Model Monitor, tracking latency, drift, and performance degradation in production and enabling self-healing retraining workflows for GenAl models.
Bhavana Akula's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.


