Chetan Chauhan
Data Engineer III @ Fred Hutch
About
As a Data Science graduate student at Northeastern University, I am highly passionate about building innovative products and learning new technologies. With 6+ years of work experience as a Data Engineer, I have delivered end-to-end project deployments in several Cloud Migration Projects. During my productive years at various projects, I have strengthened my technical skills in various fields of Data Science, including Cloud Computing, Pyspark in Glue Job, Databricks, Amazon Web Services, Java, Python, Advance SQL, Apache Hive, Apache Pig, Hadoop, Hbase, and Power BI.In addition, I have expertise in Kimball's Dimensional modeling techniques in data warehouse intelligence, such as Slowly Changing Dimension logic (all types of SCDs on Kimball-style star schema), Basic and Advanced Dimension table techniques, Dimension Hierarchy techniques, Basic and Advanced Fact table techniques, and Integration of files and Job Scheduling.
-
United States
Computer Software
Data Architects, Data Loading, ETL Testing, AWS s3, Salesforce.com Administration, API integration, Neural Networks, Pruning, Security, Interpersonal Skills, Intercultural Communication, Dashboards, Application Programming Interfaces (API), AWS Step Functions, Amazon Athena, SCD, Identity and Access Management (IAM), Amazon Dynamodb, Snowflake Cloud, Algorithms
Experience

ML Research Assistant
Boston, Massachusetts, United States
• Streamlined complex SNP genotype data analysis algorithms, enhancing efficiency and accuracy in genomic research. • Integrated machine learning techniques to predict significant genetic sequences, significantly speeding up the data analysis process. • Contributed to a groundbreaking approach in genomic data analysis, paving the way for significant scientific discoveries and insights. • Grateful for the mentorship of Professor Justin Kennedy, advancing the future of genomic research and understanding of complex diseases.

Senior Data Engineer (via Infosys)
Worked for Hyundai Capital America as a client, where I led critical data projects that streamlined processes and enhanced efficiency. • Spearheaded the ETL migration to AWS Cloud Data Warehouse using PySpark, improving data pipelines for ML applications. • Directed project architecture, optimized Redshift tables, developed Glue jobs, and automated processes, reducing processing time by 23%. • Leveraged AWS Glue, Redshift, and CloudWatch for efficient data processing, automation, and monitoring, providing real-time data insights and ensuring robust data workflows. • Collaborated with QA to automate ETL testing using Apache Airflow, reducing testing time and costs, and ensuring timely project delivery.

Data Engineer (via Infosys)
Worked for U.S. Polo Assn. as a client, where I played a pivotal role in data analysis and ETL process optimization. • Worked with ETL developers to translate business requirements into precise ETL job specifications, documenting data mappings, dictionaries, and lineage diagrams to ensure integrity of data assets using Snowflake and Star schema. • Utilized Oracle SQL Developer Data Modeler for comprehensive data modeling, creating logical and physical data models that supported diverse business processes and analytical needs. • Successfully migrated DataStage jobs to AWS Data Pipelines, designing and implementing strategies for complex ETL processes using S3(Bucket policies, Lifecycle policies), Glue, Lambda, Python, and MongoDB, ensuring seamless transition and minimal disruption to business operations. • Assisted the QA team in running tests to ensure business rules on transformed data and ETL processes, utilizing JIRA for tracking.
Education

Data Science
• Conducted cutting-edge ML research as a research assistant at Northeastern University, focusing on estimating the dimension of the latent manifold of the MNIST Digits dataset using UMAP, a dimensionality reduction technique. • Demonstrated expertise in evaluating clustering algorithms and determining the optimal number of clusters using external indices and statistical methods. • Pre-processed large volume datasets, performed Apache Hive for analysis and querying of big data, Pig for data transformation, and R language to visualize the data as part of a sport analysis project. • Developed test plans to improve the processing of data by adding Machine Learning Algorithms to evaluate more accurately, resulting in a significant increase of 42% reliable data of players for the team management.
Chetan Chauhan's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.






