Santhosh G

Santhosh G

Sr.Hadoop Developer @ State Street

About

• 8+ years of professional experience in Requirements Analysis, Design, Development and Implementation of Java, J2EE and Big Data technologies. • 4+ years of experience in using Hadoop and its ecosystem components like HDFS, Map Reduce, Yarn, Spark, Hive, Pig, HBase, Oozie, Flume, Storm and Sqoop. • In depth understanding of Hadoop Architecture and its various components such as Job Tracker, Task Tracker, Name Node, Data Node, Resource Manager and Map Reduce concepts • Strong experience creating real time data streaming solutions using Apache Spark Core, Spark SQL and Data Frames. • Hands on experience with Spark streaming to receive real time data using Kafka. • Developed Simple to complex Map/reduce streaming jobs using Java language • Developed Hive and Pig scripts for handling business transformations and analyzing data. • Developed Sqoop scripts for large dataset transfer between Hadoop and RDBMs. • Experience in job workflow scheduling and monitoring tools like Oozie. • Experience in HBase Cluster Setup and Implementation • Experience using various Hadoop Distributions (Cloudera, Hortonworks, Amazon AWS) to fully implement and leverage new Hadoop features. • Proficient in using Cloudera Manager, an end-to-end tool to manage Hadoop operations in Cloudera Cluster • Good knowledge and experience of Real time streaming technologies Spark and Kafka

Country

United States

City

Quincy

Industry

Financial Services

Skill

-

Experience

State Street

Sr.Hadoop Developer

State Street

LinkedIn
2017-3 - Present · 9 yrs 7 mos

Quincy MA

• Convert existing Mart creation application from MYSQL to Hive. • Performed dynamic transformation on Hive external tables to create the data mart. • Developed data pipeline using Spark, Hive, Sqoop and Kafka to ingest data into Hadoop platform for analysis. • Developed Spark code using Scala and Spark-SQL/Streaming for streaming of data. • Performed Data Ingestion from multiple internal clients using Apache Kafka. • Ingested data into HDFS using Sqoop and analyzed the data using Spark (Dataframes,DataSets and Spark-SQL), and series of Hive scripts to produce summarized results from Hadoop to downstream systems. • Used Spark Dataframes,DataSets, Spark-SQL extensively • Developed simple to complex Spark streaming jobs using Scala for real time processing • Created microservice to read from kafka stream and storing it in s3 or file system. • Developed an application from scratch and wrote code on KafkaProducer and Consumer as per the requirement. • Implemented installation and configuration of multi-node cluster on the cloud using Amazon Web Services (AWS) on EC2. • Used cloud computing on the multi-node cluster and deployed Hadoop application on cloud S3 and used Elastic Map Reduce (EMR) to run a Map-reduce. • Reading the file from s3 Storage using spark streaming and storing to s3 inparquet format and query using Athena. • Created Java component for Athena to query back from s3 parquet files. • Exported the data sets from Hive to RDBMS using Oozie. • Implemented business logic in Hive and written UDF’s to process the data for analysis. • Implemented dynamic partitions, bucketing and compression techniques in Hive External Tables and optimized the query performance. • Created HBase tables and column families to store the user event data. • Used Cloudera Manager to manage and monitor Hadoop Stack. • Actively participated in analysis, requirement gathering and development of various POC's for various business units.

Progressive Insurance

Sr.Hadoop Developer

Progressive Insurance

LinkedIn
2015-8 - Present · 11 yrs 2 mos

Colorado Springs, Colorado Area

• Implemented best practices for the full software development life cycle including coding standards, code reviews, source control management and build processes. • Processed data ingested into HDFS using Sqoop, custom HDFS Adaptors and analyzed the data using Spark, Hive, MapReduce and produced summary results from Hadoop to downstream systems. • Create/Modify Shell scripts for scheduling data cleansing scripts and ETL loading process. • Developed Spark applications to perform all the data transformations on User behavioral data coming from multiple sources. • Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS using Scala (Prototype). • Implemented Spark using Scala and SparkSQL for faster testing and processing of data. • Used Sqoop to import the data to Hadoop Distributed File System (HDFS) from RDBMS. • Created components like Hive UDFs for missing functionality in HIVE for analytics. • Worked on various performance optimizations like using distributed cache for small datasets, Partition, Bucketing in Hive and Map Side joins. • Used Oozie and Oozie coordinators to deploy end to end data processing pipelines and scheduling the work flows.

Strategic Solutions Group (SSG)

Sr.Hadoop Developer

Strategic Solutions Group (SSG)

LinkedIn
2014-1 - 2015-7 · 1 yr 7 mos

Needham,MA

• Developed simple to complex Map Reduce jobs using Java language for processing and validating the data. • Developed data pipeline using Sqoop, Spark, Map Reduce, and Hive to ingest, transform and analyze operational data. • Developed Map Reduce and Spark jobs to summarize and transform+ raw data. • Implemented Spark using Scala and utilizing Data frames and Spark SQL API for faster processing of data. • Used Spark for interactive queries, processing of streaming data and integration with popular NoSQL database for huge volume of data. • Real time streaming the data using Spark with Kafka • Handled importing data from different data sources into HDFS using Sqoop and also performing transformations using Hive, Map Reduce and then loading data into HDFS. • Exported the analyzed data to the relational databases using Sqoop, to further visualize and generate reports for the BI team. • Collecting and aggregating large amounts of log data using Flume and staging data in HDFS for further analysis • Analyzed the data by performing Hive queries (Hive QL) and running Pig scripts (Pig Latin) to study customer behavior. • Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting. • Developed Hive scripts in Hive QL to de-normalize and aggregate the data. • Created HBase tables and column families to store the user event data. • Scheduled and executed workflows in Oozie to run Hive and Pig jobs. • Used Impala to read, write and query the Hadoop data in Hive.

Verizon

Hadoop Developer

Verizon

LinkedIn
2012-10 - 2013-12 · 1 yr 3 mos

Irving , TX

• Worked with highly unstructured and semi structured data of 2 Petabytes in size. • Involved in full life cycle of the project from Design, Analysis, logical and physical architecture modeling, development, Implementation, testing. • Created Hive Tables, loaded transactional data from Teradata using Sqoop. • Developed MapReduce (YARN) jobs for cleaning, accessing and validating the data. • Created and worked Sqoop jobs with incremental load to populate Hive External tables. • Developed optimal strategies for distributing the web log data over the cluster importing and exporting the stored web log data into HDFS and Hive using Sqoop. • Extensive experience in writing Pig scripts to transform raw data from several data sources into forming baseline data. • Implemented Hive Generic UDF’s to incorporate business logic into Hive Queries. • Analyzed the web log data using the HiveQL to extract number of unique visitors per day, page views, visit duration, most visited page on website. • Integrated Oozie with the rest of the Hadoop stack supporting several types of Hadoop jobs out of the box (such as Map-Reduce, Pig, Hive, and Sqoop) as well as system specific jobs (such as Java programs and shell scripts). • Creating Hive tables and working on them using Hive QL. • Designed and Implemented Partitioning (Static, Dynamic), Buckets in HIVE. • Involved in End-to-End implementation of ETL logic. • Worked on Cluster co-ordination services through Zookeeper. • Monitored workload, job performance and capacity planning using Cloudera Manager. • Involved in build applications using Maven and integrated with CI servers like Jenkins to build jobs. • Exported the analyzed data to the RDBMS using Sqoop for to generate reports for the BI team. • Worked collaboratively with all levels of business stakeholders to architect, implement and test Big Data based analytical solution from disparate sources. • Involved in Agile methodologies, daily scrum meetings, spring planning.

Renaissance Softlabs

Java Developer

Renaissance Softlabs

LinkedIn
2011-8 - 2012-9 · 1 yr 2 mos

Hyderabad Area, India

• Actively Participated in JAD (Joint application development) sessions for requirements gathering and documenting business process. • Used JSP, Struts, JSTL tags, Java Script for building dynamic web pages. Added tag libraries like Display tag, Tiles and Validator for extended flexible page design. (For more flexible page design introduced tag libraries like Display tag and Validator tags.) • Incorporated J2EE design Patterns (Business Delegate, Singleton, Data Access Object, Data Transfer Object, MVC) for the Middle Tier development. • Data access framework by Spring is used for automatically acquiring and releasing database resources and exception handling by spring data access hierarchy for better handling of database connections with JDBC. • Established communication among external systems using Web Services (SOAP). • Implemented several JUnit test cases. • Implemented a logging application, Web Logging for better trace the data flow on application server using Log4J. • Used Clear Case as a version control of the application with developments streams. • Worked with team of Developers and Testers to resolve the issues with the server timeouts and database connection pooling issues. Initiated Profiling using RAD for finding the Objects memory leaks

Misys

Java Developer

Misys

LinkedIn
2009-11 - 2011-7 · 1 yr 9 mos

Bengaluru Area, India

• Responsible for coordinating on-site and off-shore development teams in various phases of the project. • Involved in developing dynamic Jsp and doing page validations using Java Script. • Involved in database schema design and review meetings. • Designed a nightly build process for updating the catalogue and intimating the user of the pending authorization. • Used automated test scripts and tools to test the application in various phases. Coordinated with Quality Control teams to fix issues that were identified • Involved in writing Stored Procedures using Oracle. • Responsible for building projects in deployable files (WAR files and JAR files) • Designed and developed base classes, framework classes and common re-usable components. • Involved in performance tuning, debugging production problems during testing and deployment phases of the project • Involved in re-factoring the existing components to meet the current application requirements • Used various Java and J2EE APIs including JDBC, XML, Servlet, JSP, and JavaBean. • Support Production Team members in the development and testing of production Implementation Plans and the Midrange group during Migrations. • Involved in testing, maintenance and production support of the application. • Responded to requests from Technical Team members to prepare a TAR and configured files for Production migration

Punjab National Bank

Java Developer

Punjab National Bank

LinkedIn
2008-7 - 2009-10 · 1 yr 4 mos

New Delhi Area, India

• Gathered and analyzed the requirements and converted them into User Requirement specifications and Functional Requirement Specifications • Involved in Full Software Development Life Cycle (SDLC). Used Agile Methodology to develop the entire application • Designed and implemented the User interface using HTML, CSS,JavaScript,JQuery and SQL Server. • Developed Interfaces using JSP based on the Users, Roles and Permissions. Screen options were displayed on User permissions. This was coded using Custom Tags in JSP using Tag Libraries. • Created web services using Advanced J2EE technologies to communicate with external systems. • Involved in the UI development, including layout and front-end coding per the requirements of the client. Used HTML4, CSS, Bootstrap, JavaScript. • Implemented client side validations using JQuery and used JSF core, PrimeFaces and FLEX to develop UI pages • Used Hibernate along with Spring Framework to integrate with Oracle database • Built complex SQL queries and ETL scripts for data extraction and analysis to define the application requirements. • Used DOM and SAX parsers with JAXP API. • Implementing JUnit test cases to test Java classes. • Utilized Rational Clear case for version control of the application. This involved creating development streams and defect streams. • Utilized WSAD for developing the application

Santhosh G's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.