Kavya Shah
Database Engineer @ San Diego Supercomputer Center
About
I am Kavya Shah, a dedicated and enthusiastic student currently pursuing my Bachelor's degree in Math and Computer Science with a minor in Data science at University of California, San Diego. With a strong foundation in programming, data structures, and algorithms, I am passionate about leveraging technology to solve real-world problems. My experience includes working on various projects that enhance my skills in software development and data analysis. I am committed to continuous learning and staying updated with the latest advancements in the tech industry, aiming to contribute effectively to any team or project I am part of.
United States
San Diego
Higher Education
LangChain, Deep agents, Model Context Protocol (MCP), Tool calls, Communication, Educational Leadership, Problem Solving, Convolutional Neural Networks (CNN), Image classification, Data Analysis, Machine Learning, Pandas (Software), Java, Python (Programming Language)
Experience

Database Engineer
-Engineered a secure data migration pipeline to extract large-scale congenital heart disease datasets from AWS DynamoDB using Boto3, ensuring full compliance with data use agreements and de-identification requirements -Transformed semi-structured clinical and MRI metadata into standardized JSON schemas and migrated them into PostgreSQL using Dockerized workflows -Built reproducible database dumps and containerized pipelines to support scalable downstream analytics and cross-institutional research access -Collaborated with researchers to validate data integrity and consistency across 1,000+ patient records and multi-center clinical sources

Product Management Intern
San Diego, California, United States
- Led sprint planning and backlog prioritization, achieving 98% task completion and improving delivery consistency across the team. - Allocated burndown charts and facilitated Agile ceremonies to enhance transparency, ownership, and team morale. - Collaborated with developers and designers to define user stories, refine technical requirements, and align features with product goals. -Oversaw integration across front-end (JavaScript, HTML/CSS) and back-end (Python, SQL) components to improve usability and performance. -Conducted user research and testing to translate feedback into actionable enhancements, strengthening product-market fit.

FitLife Assistant
Providing customer service Answer questions both in-person, on the telephone and electronically about all FitLife programs Distribute rosters and collect staff timesheets Learn Recreation Point-of-Sale (POS) system Assist at info tables and outreach events

Undergraduate Research Assistant- Data Science&Bioinformatics
San Diego, California, United States
- Analyze single-cell RNA-seq and ATAC-seq datasets to study immune responses in liver and colon cancer - Use Python, Scanpy, and Velocyto to generate UMAP visualizations and explore RNA velocity dynamics - Investigate integration of multi-omics data (gene expression + chromatin accessibility) for biological insight - Work closely with PI on dataset curation, bioinformatics pipelines, and literature-based validation - Apply data science techniques to real-world biomedical research in a cancer immunology context

Global Ambassador Program (GAP) mentor
Guided International and out of state incoming students at UCSD. Provided friendly support in order to engage our incoming undergraduate international and out-of-state students and help them create solid relationships within our campus community.

DataHacks Team
La jolla
-Collaborated with a 20-person board to organize and execute DataHacks, UC San Diego’s annual data science hackathon, attracting 200+ participants. -Helped design hackathon challenges using real-world and pop culture datasets to encourage creative data storytelling and technical excellence. -Led planning for hands-on workshops on data preprocessing, Python libraries (pandas, scikit-learn), and GitHub basics to support participant success. -Contributed to securing sponsorships, managing vendor relationships, and coordinating logistics to ensure a seamless event experience. -Promoted a welcoming, inclusive environment by fostering collaboration between participants from diverse academic and professional backgrounds.

Machine Learning Intern
Ahmedabad, Gujarat, India
Data Analysis and Preprocessing: Conducted thorough data cleaning, normalization, and feature engineering to prepare datasets for modeling. Employed techniques such as handling missing values, outlier detection, and encoding categorical variables to ensure robust model performance. Model Development and Implementation: Developed and implemented various machine learning models including Linear Regression, Classification, Logistic Regression, and Support Vector Machines (SVM), Convolutional Neural Networks(CNNs). Utilized Python libraries such as Scikit-Learn, Pandas, and NumPy for efficient model building and evaluation. Evaluation Metrics: Employed evaluation metrics like Confusion Matrices and Classification Reports to assess model performance. These metrics provided comprehensive insights into model accuracy, precision, recall, F1-score, and other critical performance indicators. Hyperparameter Tuning: Optimized model performance through hyperparameter tuning using techniques such as Grid Search and Random Search. Visualization and Reporting: Created detailed visualizations using Matplotlib and Seaborn to represent data trends and model performance. Developed comprehensive reports and presentations to communicate findings and insights to stakeholders.
Education

Mathematics and Computer Science
Minor in Data Science Relevant Coursework- MATH 18 Linear Algebra MATH 20B Calculus for Science and Engineering MATH 20C Calculus and Analytic Geometry MATH 20D Introduction to Differential Equations MATH 20E Vector Calculus MATH 109 Mathematical Reasoning MATH 154 Discrete Mathematics and Graph Theory MATH 183 Statistical Methods MATH 103A Modern Algebra I - Group Theory MATH 189 Data Analysis and Inference CSE 11 Accelerated Introduction to Programming CSE 12 Basic Data Structures and Object-Oriented Design CSE 29 Systems Programming and Software Tools CSE 100 Advanced Data Structures CSE 101 Design and Analysis of Algorithms CSE 105 Theory of Computation CSE 110 Software Engineering CSE 156 Statistical Natural Language Processing DSC 10 Principles of Data Science DSC 40A Theoretical Foundations of Data Science I DSC 80 Practice of Data Science
Kavya Shah's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.




