
Sayeed Mohammed
Big Data Engineer @ PNC
About
Hadoop Developer with several years of experience in the field of Bigdata and Data Engineering. Experience in AWS Services like S3, ELB, SQS and EC2. Expertise in building end to end components using EMR tools, Microservices. Ability to analyze complex problems and develop effective solutions for them. Excellent communication skills working with the client and the team members at all levels within an organization
United States
Santa Clara
Information Technology & Services
Python (Programming Language), Shell Scripting, Bash, Ksh, Pandas, Hadoop, Apache Spark, Apache Kafka, Java, Unix, hadoop developer, spark developer , kafka , kafka java, unix, Linux, Big Data, bigaws, Apache Spark Streaming, Scala
Experience

Big Data Engineer
Pennsylvania, United States
Responsibilities Worked on developing architecture document and proper guidelines Responsible in Installation and Configuration of Hadoop Eco system components using CDH 5.2 Distribution. Responsible to manage data coming from different sources and involved in HDFS maintenance and loading of structured and unstructured data. Processed Multiple Data sources input to same Reducer using Generic Writable and Multi Input format. Worked Big data processing of clinical and non-clinical data using Map Reduce. Visualize the HDFS data to customer using BI tool with the help of Hive ODBC Driver. Customized BI tool for manager team that perform Query analytics using HiveQL. Imported data using Sqoop to load data from MySQL to HDFS on regular basis. Created Partitions, Buckets based on State to further process using Bucket based Hive joins. Created Hive Generic UDF's to process business logic that varies based on policy. Moved Relational Data base data using Sqoop into Hive Dynamic partition tables using staging tables. Experienced in Monitoring Cluster using Cloudera manager. Involved in Discussions with business users to gather the required knowledge. Capable of creating real time data streaming solutions and batch style large scale distributed computing applications using Apache Spark, Spark Streaming, Kafka and Flume. Analyzing the requirements to develop the framework. Designed and developed architecture for data services ecosystem spanning Relational, NoSQL and Big Data technologies. Loaded and transformed large sets of structured, semi structured and unstructured data using Hadoop/Big Data concepts. Developed Java Spark streaming scripts to load raw files and corresponding. Processed metadata files into AWS S3 and Elasticsearch cluster. Developed Python Scripts to get the recent S3 keys from Elasticsearch. Elaborated Python Scripts to fetch/get S3 files using Boto3 module.

Big Data Engineer
JPMC
Chicago, Illinois, United States
• Responsible for building scalable distributed data solutions using Hadoop. • Responsible for managing and scheduling Jobs on a Hadoop cluster. • Responsible for loading the customers Data from SAS to MSSQL 2016 and perform data massaging, mining & cleansing then export to HDFS and Hive using Sqoop • Written PIG scripts to process the Credit Card and Debit Card Transactions for Active customers by joining the data from HDFS and Hive using HCatalog for various merchants • Written Python UDFs to process the RegEx and return the valid Merchant codes and names using streaming • Responsible for creating data pipeline using Kafka, Spark Streaming • Loading data from UNIX file system to HDFS and vice versa. • Improving the performance and optimization of existing algorithms in Hadoop using Spark context, Spark-SQL and Spark YARN. • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, and Python. • Developed POC using Spark SQL and MLlib libraries along with Kafka and other tools as per requirement then deployed on the Yarn cluster.

Big Data Engineer
Sunnyvale, California, United States
• Developed Spark Applications by using Java and Implemented Apache Spark data processing project to handle data from various RDBMS and Streaming sources. • Worked with the Spark for improving performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Spark MLlib, Data Frame, Pair RDD's, Spark YARN. • Used Spark Streaming APIs to perform transformations and actions on the fly for building common learner data model which gets the data from Kafka in Near real time and persist it to Cassandra. • Consumed XML messages using Kafka and processed the xml file using Spark Streaming to capture UI updates. • Developed Preprocessing job using Spark Data frames to flatten Json documents to flat file. • Load D-Stream data into Spark RDD and do in memory data Computation to generate Output response.

Data Engineer
United States
● Responsible for building an Enterprise Data Lake to bring ML ecosystem capabilities to production and make it readily consumable for data scientists and business users. ● Processing and transforming the data using AWS EMR to assist the Data Science team as per business requirement. ● Developing Spark applications for cleaning and validation of the ingested data into the AWS cloud. ● Working on fine-tuning Spark applications to improve the overall processing time for the pipelines. ● Implement simple to complex transformation on Streaming Data and Datasets. ● Work on analysing Hadoop cluster and different big data analytic tools including Hive, Spark, Python, Sqoop, flume, Oozie.6 ● Use Spark Streaming to stream data from external sources using Kafka service and responsible for migrating the code base from Cloudera Platform to Amazon EMR and evaluated Amazon eco systems components like RedShift, Dynamo DB. ● Perform configuration, deployment, and support of cloud services in Amazon Web Services (AWS). ● Designing and building multi-terabyte, full end-to-end Data Warehouse infrastructure from the ground up on Confidential Redshift. ● Design Develop and test ETL Processes in AWS Glue to migrate Campaign data from external sources like S3, ORC/Parquet/Text Files into AWS Redshift. ● Migrate an existing on-premises application to AWS. ● Build and configure a virtual data centre in the Amazon Web Services cloud to support Enterprise Data Warehouse hosting including Virtual Private Cloud, Security Groups, Elastic Load Balancer. ● Implement data ingestion and handling clusters in real time processing using Kafka. ● Develop Spark Programs using Scala and Java API's and performed transformations and actions on RDD's. ● Develop Spark application for filtering Json source data in AWS S3 and store it into HDFS with partitions and used spark to extract schema of Json files

Data Engineer
Bloomington, Illinois, United States
● Worked on development of data ingestion pipelines using ETL tool, Talend & bash scripting with big data technologies including but not limited to Hive, Impala, Spark, Kafka, and Talend. ● Experience in developing scalable & secure data pipelines for large datasets. ● Gathered requirements for ingestion of new data sources including life cycle, data quality check, transformations, and metadata enrichment. ● Importing data from MS SQL server and Teradata into HDFS using Sqoop. ● Supported data quality management by implementing proper data quality checks in data pipelines. ● Enhancing Data Ingestion Framework by creating more robust and secure data pipelines. ● Implemented data streaming capability using Kafka and Talend for multiple data sources. ● Responsible for maintaining and handling data inbound and outbound requests through big data platform. ● Working knowledge of cluster security components like Kerberos, Sentry, SSL/TLS etc. ● Worked with multiple storage formats (Avro, Parquet) and databases (Hive, Impala, Kudu). ● Involved in the development of agile, iterative, and proven data modeling patterns that provide flexibility. ● Created Oozie workflows to automate and productionize the data pipelines. ● Troubleshooted user's analyses bugs (JIRA and IRIS Ticket). ● Involved in developing spark applications to perform ELT kind of operations on the data. ● Worked with SCRUM team in delivering agreed user stories on time for every Sprint. ● Worked on analyzing and resolving the production job failures in several scenarios. ● Implemented UNIX scripts to define the use case workflow and to process the data files and automate the jobs. ● Knowledge on implementing the JILs to automate the jobs in production cluster.

Java Software Developer
Hyderabad, Telangana, India
Involved in designing of Front-end, Implementing Functionality with Business Logic. Mentoring and grooming juniors technically as well as professionally on Agile practices and Java/J2EE development issues. Allowing department access to accounts receivable information. Subject to security and authorization constraints. Involved in designing and Code Reviews. Assisted in troubleshooting architectural problems. Design and development of report generation by using velocity framework. Interacting with business and reporting teams for requirement gathering, configure walkthrough and UAT. Generating monthly Sales report, Collection Report and Receivable-aging reports. Prepared use cases designed and developed object models and class diagrams. Worked on one of the most critical modules for project, right from the beginning phase which included requirement gathering, analysis, design, review, and development. Module lead located to another location had KT from him about roughly 2 weeks, Lead was absorbed by client. Took initiative in building a new team of more than 6 members with proper knowledge transfer sessions assigning and managing tasks with JIRA. Learned Backbone JS and worked with UI team on UI enhancements. Working with BA, QA to identify and fix bugs, raise new feature and enhancements. Was greatly appreciated by client with appreciation certificate and client bonus of 10k and 50k, respectively. Analyse the generated Junit and add proper asserts and make it more code specific along with increasing the code coverage. This helped to boast my product knowledge as well as my Junit writing skills. Addressed issues related to application integration and compatibility. Performed enhancements to User Group Tree using DOJO and JSON to meet business.
Sayeed Mohammed's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.


