Sneha Maurya
Data Scientist @ IBM
About
Hi! I'm Sneha, a recent grad from Columbia's MS in Data Science. I've worked next action prediction for code completion at IBM, manufacturing ML at NXP, and multimodal RAG at Columbia. I built and published a clinical LLM evaluation benchmark; 22 models, expert-in-the-loop validation(arXiv Pre-print). I work across post-training, eval pipelines, and vertical agents. I care about observability over agent behavior, structured reasoning, and tools that augment, Expand Developer experiences. I wold love to connect with you if you are someone hiring or working in the field to build something together!📧 sm5755@columbia.edu
United States
New York City Metropolitan Area
Computer Software
Team work, Software Systems Engineering, Java, GPU, Extract, Transform, Load (ETL), Operating Systems, Project Management, Knowledge Engineering, R Shiny, Multimodal LLMs, CUDA & GPU Optimization, Cloud-based ML Deployment (Docker, AWS), Transformer Quantization & Compilation, Python (PyTorch, LangChain, Streamlit), streamlit, BGE-M3, Gaussian Mixture Model, SQLite, PostgreSQL, Beautiful Soup
Experience

Data Scientist
New York, United States
• Developed a curated dataset of semantic code edits by mining and normalizing multi-file Java commits, using tree-sitter to extract AST structure and represent changes consistently across repositories. • Fine-tuned Qwen-3 Code(30B) with graph-aware attention and cross-file context encoding to improve next-action prediction in code completion. Developed a full evaluation pipeline using AST edit-distance, semantic checks, and compile-time validation to assess model behavior on large, real-world edit corpora. Capstone Project @ Data Science Institute | Columbia University Advisor - Atin Sood

Teaching Assistant II
New York, New York, United States
As a TA I ran weekly lab sessions and open office hours where I walked students step-by-step through data manipulation in pandas, SQL joins, visualization, and building simple ML models for statistical problems. I reviewed code live in Jupyter notebooks, helped debug pipelines, and explained statistical ideas through hands-on examples so concepts clicked. I also prepared reusable Jupyter lab templates and short how-to guides for common issues, which made project iteration faster and reduced students’ debugging time.
Data Scientist
Austin, TX
I led two projects focused on improving tool reliability analysis and knowledge retention across multiple fabrication lines. I built a GenAI-assisted NLP workflow in Dataiku to cluster 10K+ EMS logs from CMP and Films tools, highlight recurring downtime themes, and automate failure summarization using prompt-engineered LLMs. I paired this with interactive Power BI Pareto dashboards to help engineers prioritize tool issues in real time. I also implemented an R Shiny ingestion pipeline to standardize 200+ Kaizen, 8D, and Lessons Learned reports into structured JSON and centralize them in Teradata, improving cross-line knowledge sharing and audit traceability.

Teaching Assistant II
New York, United States
Teaching Assistant II for two courses to MBAs and Executive MBAs at the department of Decision Risk and Operations. Both the courses are Taught by Professor Mattan Griffel. - Python for MBAs - Intro to Databases for Business Analysis

Graduate Student Researcher
New York, United States
Led the development of a multimodal Retrieval-Augmented Generation (RAG) system to extract workforce-related insights from CSRD sustainability reports. Fine-tuned vision-language models including BGE-M3, Qwen2.5-VL, and GMM-based retrieval to process 300+ pages of complex PDF layouts—spanning OCR, tables, and diagrams with significant structured extraction accuracy. The system adapts to diverse reporting styles and regulatory queries, ensuring alignment with disclosure standards. Also built a Streamlit-based interface for query resolution and relevance visualization, significantly reducing manual review time.

Machine Learning Engineer
Mumbai
• Built a GPT-4 based clinical NLP system that reads raw medical reports, retrieves relevant context with RAG, and generates clear, patient-friendly summaries; worked iteratively with clinicians to refine prompts and automated checks so outputs are clinically sensible and easy for patients to act on. • Engineered a secure, production AWS pipeline (S3 - Lambda - inference - versioned outputs) with CI/CD, monitoring and secrets/key management, and partnered with marketing to surface insights via SQL/Tableau dashboards for operational use.

Research Intern
Pittsburgh, Pennsylvania, United States
Researched joint particle-picking and 3D orientation estimation for cryo-ET subtomograms under the guidance of Dr. Min Xu and Mostofa Rafid Uddin. Worked with GUM-net and downstream alignment methods, experimenting with rotation-equivariant feature learning and pose-refinement strategies to improve the robustness of 3D structural analysis.

Machine Learning Intern
Singapore
• Led the setup and optimization of Ambari clusters at Amazon Web Services, enhancing data processing performance and reliability. • Designed and deployed robust data pipelines, streamlining data flow for critical business applications. • Integrated Big Data tools like Hadoop, Kafka, and Hive, gaining hands-on expertise in scalable data management.
Education

Data Science
Fall (2025) Capstone (IBM Research) - Next action prediction in code generation High performing Machine Learning LLMs for Business Applications Spring(2025) Deep Learning for NLP Machine Learning for Data Science Statistical Inference and Modelling Advanced Spoken Language Processing Financial Structuring for Data Science Fall (2024) Applied Machine Learning Probability & Stat for Data Science Algorithms for Data Science Exploratory Data Analysis & Visualisations

Attended an intensive NUS lecture series (faculty + industry speakers) on deep learning, model efficiency, and deployment; completed hands-on AWS workshops covering data storage and inference pipelines. Built and presented a short project: trained a lightweight spatio-temporal model for webcam gesture recognition and demonstrated an end-to-end AWS pipeline (S3 → preprocessing → inference API) during the program.

Computer Science and Engineering
Relevant Courses: Operating Systems, System Design, Database Management System, Data Structures and Algorithms, Statistical Machine Learning, Natural Language Processing, Software Management
Sneha Maurya's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.




