Santhosh G
Sr.Hadoop Developer @ State Street
About
• 8+ years of professional experience in Requirements Analysis, Design, Development and Implementation of Java, J2EE and Big Data technologies. • 4+ years of experience in using Hadoop and its ecosystem components like HDFS, Map Reduce, Yarn, Spark, Hive, Pig, HBase, Oozie, Flume, Storm and Sqoop. • In depth understanding of Hadoop Architecture and its various components such as Job Tracker, Task Tracker, Name Node, Data Node, Resource Manager and Map Reduce concepts • Strong experience creating real time data streaming solutions using Apache Spark Core, Spark SQL and Data Frames. • Hands on experience with Spark streaming to receive real time data using Kafka. • Developed Simple to complex Map/reduce streaming jobs using Java language • Developed Hive and Pig scripts for handling business transformations and analyzing data. • Developed Sqoop scripts for large dataset transfer between Hadoop and RDBMs. • Experience in job workflow scheduling and monitoring tools like Oozie. • Experience in HBase Cluster Setup and Implementation • Experience using various Hadoop Distributions (Cloudera, Hortonworks, Amazon AWS) to fully implement and leverage new Hadoop features. • Proficient in using Cloudera Manager, an end-to-end tool to manage Hadoop operations in Cloudera Cluster • Good knowledge and experience of Real time streaming technologies Spark and Kafka
United States
Quincy
Financial Services
-
Experience

Sr.Hadoop Developer
Quincy MA
• Convert existing Mart creation application from MYSQL to Hive. • Performed dynamic transformation on Hive external tables to create the data mart. • Developed data pipeline using Spark, Hive, Sqoop and Kafka to ingest data into Hadoop platform for analysis. • Developed Spark code using Scala and Spark-SQL/Streaming for streaming of data. • Performed Data Ingestion from multiple internal clients using Apache Kafka. • Ingested data into HDFS using Sqoop and analyzed the data using Spark (Dataframes,DataSets and Spark-SQL), and series of Hive scripts to produce summarized results from Hadoop to downstream systems. • Used Spark Dataframes,DataSets, Spark-SQL extensively • Developed simple to complex Spark streaming jobs using Scala for real time processing • Created microservice to read from kafka stream and storing it in s3 or file system. • Developed an application from scratch and wrote code on KafkaProducer and Consumer as per the requirement. • Implemented installation and configuration of multi-node cluster on the cloud using Amazon Web Services (AWS) on EC2. • Used cloud computing on the multi-node cluster and deployed Hadoop application on cloud S3 and used Elastic Map Reduce (EMR) to run a Map-reduce. • Reading the file from s3 Storage using spark streaming and storing to s3 inparquet format and query using Athena. • Created Java component for Athena to query back from s3 parquet files. • Exported the data sets from Hive to RDBMS using Oozie. • Implemented business logic in Hive and written UDF’s to process the data for analysis. • Implemented dynamic partitions, bucketing and compression techniques in Hive External Tables and optimized the query performance. • Created HBase tables and column families to store the user event data. • Used Cloudera Manager to manage and monitor Hadoop Stack. • Actively participated in analysis, requirement gathering and development of various POC's for various business units.

Sr.Hadoop Developer
Colorado Springs, Colorado Area
• Implemented best practices for the full software development life cycle including coding standards, code reviews, source control management and build processes. • Processed data ingested into HDFS using Sqoop, custom HDFS Adaptors and analyzed the data using Spark, Hive, MapReduce and produced summary results from Hadoop to downstream systems. • Create/Modify Shell scripts for scheduling data cleansing scripts and ETL loading process. • Developed Spark applications to perform all the data transformations on User behavioral data coming from multiple sources. • Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS using Scala (Prototype). • Implemented Spark using Scala and SparkSQL for faster testing and processing of data. • Used Sqoop to import the data to Hadoop Distributed File System (HDFS) from RDBMS. • Created components like Hive UDFs for missing functionality in HIVE for analytics. • Worked on various performance optimizations like using distributed cache for small datasets, Partition, Bucketing in Hive and Map Side joins. • Used Oozie and Oozie coordinators to deploy end to end data processing pipelines and scheduling the work flows.

Sr.Hadoop Developer
Needham,MA
• Developed simple to complex Map Reduce jobs using Java language for processing and validating the data. • Developed data pipeline using Sqoop, Spark, Map Reduce, and Hive to ingest, transform and analyze operational data. • Developed Map Reduce and Spark jobs to summarize and transform+ raw data. • Implemented Spark using Scala and utilizing Data frames and Spark SQL API for faster processing of data. • Used Spark for interactive queries, processing of streaming data and integration with popular NoSQL database for huge volume of data. • Real time streaming the data using Spark with Kafka • Handled importing data from different data sources into HDFS using Sqoop and also performing transformations using Hive, Map Reduce and then loading data into HDFS. • Exported the analyzed data to the relational databases using Sqoop, to further visualize and generate reports for the BI team. • Collecting and aggregating large amounts of log data using Flume and staging data in HDFS for further analysis • Analyzed the data by performing Hive queries (Hive QL) and running Pig scripts (Pig Latin) to study customer behavior. • Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting. • Developed Hive scripts in Hive QL to de-normalize and aggregate the data. • Created HBase tables and column families to store the user event data. • Scheduled and executed workflows in Oozie to run Hive and Pig jobs. • Used Impala to read, write and query the Hadoop data in Hive.

Hadoop Developer
Irving , TX
• Worked with highly unstructured and semi structured data of 2 Petabytes in size. • Involved in full life cycle of the project from Design, Analysis, logical and physical architecture modeling, development, Implementation, testing. • Created Hive Tables, loaded transactional data from Teradata using Sqoop. • Developed MapReduce (YARN) jobs for cleaning, accessing and validating the data. • Created and worked Sqoop jobs with incremental load to populate Hive External tables. • Developed optimal strategies for distributing the web log data over the cluster importing and exporting the stored web log data into HDFS and Hive using Sqoop. • Extensive experience in writing Pig scripts to transform raw data from several data sources into forming baseline data. • Implemented Hive Generic UDF’s to incorporate business logic into Hive Queries. • Analyzed the web log data using the HiveQL to extract number of unique visitors per day, page views, visit duration, most visited page on website. • Integrated Oozie with the rest of the Hadoop stack supporting several types of Hadoop jobs out of the box (such as Map-Reduce, Pig, Hive, and Sqoop) as well as system specific jobs (such as Java programs and shell scripts). • Creating Hive tables and working on them using Hive QL. • Designed and Implemented Partitioning (Static, Dynamic), Buckets in HIVE. • Involved in End-to-End implementation of ETL logic. • Worked on Cluster co-ordination services through Zookeeper. • Monitored workload, job performance and capacity planning using Cloudera Manager. • Involved in build applications using Maven and integrated with CI servers like Jenkins to build jobs. • Exported the analyzed data to the RDBMS using Sqoop for to generate reports for the BI team. • Worked collaboratively with all levels of business stakeholders to architect, implement and test Big Data based analytical solution from disparate sources. • Involved in Agile methodologies, daily scrum meetings, spring planning.

Java Developer
Hyderabad Area, India
• Actively Participated in JAD (Joint application development) sessions for requirements gathering and documenting business process. • Used JSP, Struts, JSTL tags, Java Script for building dynamic web pages. Added tag libraries like Display tag, Tiles and Validator for extended flexible page design. (For more flexible page design introduced tag libraries like Display tag and Validator tags.) • Incorporated J2EE design Patterns (Business Delegate, Singleton, Data Access Object, Data Transfer Object, MVC) for the Middle Tier development. • Data access framework by Spring is used for automatically acquiring and releasing database resources and exception handling by spring data access hierarchy for better handling of database connections with JDBC. • Established communication among external systems using Web Services (SOAP). • Implemented several JUnit test cases. • Implemented a logging application, Web Logging for better trace the data flow on application server using Log4J. • Used Clear Case as a version control of the application with developments streams. • Worked with team of Developers and Testers to resolve the issues with the server timeouts and database connection pooling issues. Initiated Profiling using RAD for finding the Objects memory leaks

Java Developer
Bengaluru Area, India
• Responsible for coordinating on-site and off-shore development teams in various phases of the project. • Involved in developing dynamic Jsp and doing page validations using Java Script. • Involved in database schema design and review meetings. • Designed a nightly build process for updating the catalogue and intimating the user of the pending authorization. • Used automated test scripts and tools to test the application in various phases. Coordinated with Quality Control teams to fix issues that were identified • Involved in writing Stored Procedures using Oracle. • Responsible for building projects in deployable files (WAR files and JAR files) • Designed and developed base classes, framework classes and common re-usable components. • Involved in performance tuning, debugging production problems during testing and deployment phases of the project • Involved in re-factoring the existing components to meet the current application requirements • Used various Java and J2EE APIs including JDBC, XML, Servlet, JSP, and JavaBean. • Support Production Team members in the development and testing of production Implementation Plans and the Midrange group during Migrations. • Involved in testing, maintenance and production support of the application. • Responded to requests from Technical Team members to prepare a TAR and configured files for Production migration

Java Developer
New Delhi Area, India
• Gathered and analyzed the requirements and converted them into User Requirement specifications and Functional Requirement Specifications • Involved in Full Software Development Life Cycle (SDLC). Used Agile Methodology to develop the entire application • Designed and implemented the User interface using HTML, CSS,JavaScript,JQuery and SQL Server. • Developed Interfaces using JSP based on the Users, Roles and Permissions. Screen options were displayed on User permissions. This was coded using Custom Tags in JSP using Tag Libraries. • Created web services using Advanced J2EE technologies to communicate with external systems. • Involved in the UI development, including layout and front-end coding per the requirements of the client. Used HTML4, CSS, Bootstrap, JavaScript. • Implemented client side validations using JQuery and used JSF core, PrimeFaces and FLEX to develop UI pages • Used Hibernate along with Spring Framework to integrate with Oracle database • Built complex SQL queries and ETL scripts for data extraction and analysis to define the application requirements. • Used DOM and SAX parsers with JAXP API. • Implementing JUnit test cases to test Java classes. • Utilized Rational Clear case for version control of the application. This involved creating development streams and defect streams. • Utilized WSAD for developing the application
Santhosh G's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.

