Daksh Patel
Data Engineer @ Amazon
About
Machine Learning & Data Engineer at AWS FinTech | USC Alumnus | Aspiring AI Researcher I spend my days at AWS FinTech building the financial "ground truth" for the cloud, specifically in Cost Allocation and Scenario Modeling. In this world, accuracy is deterministic; if a billion-dollar billing simulation fails, it isn’t an "edge case," it’s a systemic failure. Yet, I am deeply troubled by the "Complexity Gap" in modern AI. My experience at Amazon, combined with my academic roots at USC (where I TAed for Professor Reza Rajati), has shown me that as our systems grow more over-engineered, we lose the "signal" for the "noise." The Research Question that drives me: Why is it that we can model millions of AWS billing scenarios with precision, but we still can't build a music recommendation engine that actually "reads the brain" instead of just echoing past patterns? I am intrigued by the failure of "Black Box" models to align with true human intent. I want to strip away algorithmic bloat to find the mathematical "truth" underneath creating systems that don't just predict, but collaborate with human discovery. I document my search for systemic transparency at: medium.com/@helloitsdaksh007. Research Interests: Scenario Modeling & Simulation, Information Bottleneck Theory, and Human-Centric Discovery.
United States
Seattle
Computer Software
Machine Learning, Amazon Web Services (AWS), Data Engineering, SQL, Large Language Models (LLM), Tranformers, Attention Mechanisms, Text Generation, Deep Learning, Applied Machine Learning, Object Detection, Fine Tuning, Generative AI, Optimization, Natural Language Processing (NLP), Data Science, AWS Glue, AWS SageMaker, Text Analytics, PyTorch
Experience

Data Engineer
Seattle, Washington, United States
– Built an automated scenario modeling engine using Python and SQL to conduct high-fidelity ”what-if” simulations, reducing latency by 85% and enabling proactive threat/impact analysis. – Architected a migration of 20TB+ of financial workloads to ATLAS with 100% integrity, developing automated validation frameworks to resolve 99% of data discrepancies for security auditing. – Refactored PySpark ETL pipelines to process billions of records daily, optimizing compute efficiency by 30% and ensuring zero production downtime for mission-critical financial systems. – Collaborated with cross-functional teams to modernize legacy proprietary logic into scalable SQL/Spark architectures, improving end-to-end data delivery speeds by 25%

Machine Learning / Data Engineer
Los Angeles, California, United States
- Developed a viral content ranking algorithm, utilizing XClip embeddings for videos and ResNet-50 embeddings for images, enhancing the prediction of content popularity by analyzing feature correlations. - Automated metadata augmentation to enrich content tagging and classification, improving the precision of recommendation models and reducing manual intervention. - Optimized multimodal data pipelines for scalable content processing, ensuring seamless integration of structured metadata, embeddings, and classification outputs for downstream applications.

Machine Learning / Data Engineer Intern
Los Angeles, California, United States
- Led large-scale metadata analysis on 9M+ images and 10M+ videos, developing automated data pipelines for content extraction, classification, and retrieval, enhancing searchability and recommendation systems. - Engineered and fine-tuned computer vision models for content analysis, leveraging deep learning techniques and metadata insights to optimize feature selection, classification accuracy, and processing efficiency. - Designed a data analysis pipeline to identify key content trends, improving feature selection and model robustness across two major processing modules, contributing to better prediction accuracy. - Integrated FAISS-based similarity search, enabling efficient image and video retrieval through optimized embedding-based indexing, reducing computational overhead and improving response times.

Graduate Teaching Assistant (Course Producer)
Los Angeles, California, United States
Course Producer – DSCI 552: Machine Learning for Data Science Under Prof. Mohammad Reza Rajati - Managed and responded to student queries on Piazza, ensuring clarity on course materials and assignments. - Provided dedicated support during help hours, assisting 350+ students in understanding complex machine-learning concepts. - Facilitated discussions and reinforced foundational ML principles, improving student engagement and comprehension.

Graduate Teaching Assistant (Course Producer)
Los Angeles, California, United States
Course Producer For DSCI 552 Machine Learning For Data Science under Prof. Mohammad Reza Rajati. - Proficiently grade assignments and provide constructive feedback to students. - Offer dedicated support during help hours, aiding students in understanding complex concepts to about 450+ students.

Graduate Teaching Assistant (Course Producer)
Los Angeles, California, United States
Course Producer For DSCI 552 Machine Learning For Data Science under Prof. Mohammad Reza Rajati. - Proficiently grade assignments and provide constructive feedback to students. - Offer dedicated support during help hours, aiding students in understanding complex concepts to about 250+ students.

Research Assistant
Los Angeles, California, United States
Radiomics Lab, Department of Radiology - Led research to optimize medical diagnosis with PyTorch, enhancing accuracy through a pipeline integrating classification with Foundation models (SAM, SAM2, MEDSAM) for segmentation and comparing results with nnU-Net - Implemented cutting-edge deep learning techniques to optimize feature extraction and classification processes, leading to enhanced interpretability and efficiency of medical imaging systems. - Developed a novel data preprocessing technique aimed at mitigating imbalance and reducing confusion within the model during training, resulting in improved classification performance and enhanced model robustness.

AI/ML Engineer
Los Angeles, California, United States
- Developed and fine-tuned custom Large Language Models (LLMs) using PyTorch, designing a chatbot capable of understanding and responding in the style of anime characters while providing detailed story-related information. This enhanced user engagement by 25%, improving interactive query responses. - Integrated AWS services, including EC2, Lambda, and S3, to deploy and scale the chatbot efficiently, ensuring low-latency query processing and optimized server load. - Engineered an interactive text-to-speech (TTS) system, utilizing NVIDIA Tacotron 2 for voice generation, enabling anime-character-based speech synthesis for a more immersive user experience. - Implemented model optimization techniques, including quantization and distillation, reducing the TTS response time and overcoming latency challenges, ensuring smooth, real-time interaction between the chatbot and anime avatar. - Designed and deployed an inference pipeline, integrating LLM-based natural language understanding with TTS processing, facilitating seamless user interactions with voice-enabled AI.

NLP Research/AI Engineer Intern
Los Angeles, California, United States
- Optimized text-to-speech (TTS) conversion speed by implementing quantization, model distillation, and efficient batching techniques, significantly reducing inference time and ensuring low-latency responses. - Overcame latency challenges by fine-tuning NVIDIA Tacotron 2 and integrating parallel processing, enabling real-time voice generation with minimal delay. - Enhanced the interaction between the chatbot and anime avatar, ensuring a seamless and natural conversational experience through synchronized speech and text processing.

Machine Learning Engineer
Siksti (Rebranded to Slikk)
Bengaluru, Karnataka, India
- Led core backend operations, managing the development, maintenance, and optimization of critical systems to ensure seamless functionality and scalability. - Engineered and automated catalog and inventory upload scripts, streamlining data processing workflows and reducing processing time by 80%, improving operational efficiency. - Optimized API performance, implementing caching strategies, query optimization, and load balancing, which resulted in a 66% reduction in response time and significantly enhanced the user experience. - Deployed and managed cloud-based infrastructure, leveraging AWS services (EC2, S3, RDS, Lambda) to ensure high availability, scalability, and security for backend operations. - Implemented CI/CD pipelines to automate deployment, reducing manual intervention and improving system reliability. - Developed and maintained database architectures using PostgreSQL, optimizing queries and indexing strategies for improved data retrieval efficiency.

ML executive
Chennai, Tamil Nadu, India
- Developed a Deep Learning model for Digit Recognition, deploying it over Heroku using Flask API. - Created a 3-Layered Neural Network from scratch using NumPy, showcasing strong technical skills. - Collaborated on a Healthcare Project, contributing to impactful research and development.

Data Scientist
Sixth Consultancy
Chennai, Tamil Nadu, India
- Created a Script to handle streaming data of the stock market and reviewed it every minute with a script written in python. - Creating a script that allows storing a particular company's data in the database every minute from the API.
Education

Applied Data Science
Coursework: • DSCI 552 Machine Learning for Data Science • DSCI 551 Foundation of Databases Management • DSCI 553 Foundations and Applications of Data Mining • ISE 529 Predictive Analytics • CSCI 566 Deep Learning and Its Applications • DSCI 517 Research Methods and Analysis for User Studies • DSCI 531 Fairness in AI • DSCI 599 Machine Learning in Security • ENGR 597 Internship in Engineering • ENGR 598 Internship in Engineering Credits: 32.5
Daksh Patel's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.




