Hosenara Khatun

Hosenara Khatun

Senior Data Engineer @ DBS Bank

About

I’m an accomplished Data Engineer with proven expertise in designing, building, and optimising scalable ETL/ELT pipelines and modern cloud-based data platforms. I specialise in transforming legacy analytics systems into high-performance, cloud-native architectures that drive business intelligence, automation, and decision-making. Skilled in PySpark, Spark, Databricks, SQL, and Python, I have hands-on experience delivering end- to-end data solutions on AWS and GCP, including services such as S3, Glue, EMR, Lambda, Redshift, Athena, Dataproc, and BigQuery. I’m passionate about architecting data lakes and lakehouse environments that enable reliable, governed, and scalable analytics. My work focuses on data quality, lineage, and governance, leveraging tools like Collibra for metadata management and implementing validation and reconciliation frameworks to ensure data accuracy and trust. I’m also experienced in performance tuning, workflow automation, and CI/CD pipelines using Airflow, Jenkins, GitLab, and Bitbucket, enabling faster, automated, and reliable data delivery. With a strong foundation in big data processing, data warehousing, and distributed computing frameworks, I bring both technical depth and strategic vision to every project. My goal is to bridge data engineering and business impact, delivering clean, actionable, and analytics-ready datasets that empower organisations to make smarter decisions at scale. Core Skills : Data Engineering | ETL/ELT | PySpark | Spark | Databricks | AWS | GCP | SQL | Python | Hadoop | Hive | Data Lakehouse | Collibra | Airflow | Jenkins | GitLab | Data Governance | Metadata Management | CI/CD | Automation | Big Data | Data Quality | Performance Tuning Reach out to me to discuss exciting opportunities.

Country

Singapore

City

Singapore

Industry

Information Technology & Services

Skill

Joomla!, WordPress Design, MySQL , AJAX, CodeIgniter, Apache Spark, Scala, Databricks, AWS, GIT, Hadoop, Hive, GCP, Python, Cassandra, Performance Tuning, Pandas (Software), Pyth (Programming Language), Data Warehousing, Airflow, Linux, Schema, MySQL, Continuous Integration (CI), Agile Methodologies, Data Lakes, Kubernetes, Data Management, Data Analysis

Experience

DBS Bank

Senior Data Engineer

DBS Bank

LinkedIn
2023-4 - 2025-4 · 2 yrs 1 mo

Singapore

Led the migration of banking analytics systems and designed multi-layered data pipelines, orchestrating scalable workloads and automating workflows to improve operational efficiency. Implemented systems to enhance reporting and communication, ensured data quality and governance for dashboards, collaborated on post-migration validation, and developed documentation to support knowledge sharing and onboarding. Significant Achievements: • Migrated a large-scale banking analytics application from SAS to PySpark, improving processing efficiency by 40%. • Achieved 99.9% reconciliation accuracy through rigorous validation. • Built multi-layered data pipelines in Databricks using Medallion Architecture (Bronze, Silver, Gold), accelerating ingestion and transformation speed by 35% for millions of records. • Developed and orchestrated workloads on Kubernetes (AWS EKS) using Application Load Controller (ALC) strategies, ensuring scalability, reliability, and cost-efficient resource utilisation. • Automated workflows with Airflow and Jenkins, reducing manual intervention. • Implemented automated email alerting systems in the ADA Data Ops Platform, reducing reporting delays and enabling proactive stakeholder communication. • Validated and enhanced data quality for Control Tower dashboards in QlikView/QlikSense, achieving close to total accuracy and improving metadata governance with Collibra. • Conducted post-migration validation with Superset in collaboration with business analysts, resolving significant levels of critical data issues within SLA. • Leveraged Python (Pandas) for data manipulation and ad hoc analyses, providing rapid insights to meet dynamic business needs. • Created comprehensive documentation of data engineering workflows and infrastructure, fostering for knowledge sharing and accelerating new engineer on boarding by 20%.

Innova Solutions

Data Engineer

Innova Solutions

LinkedIn
2022-6 - 2023-2 · 9 mos

India

Responsible for designing, building, and maintaining scalable pipelines and architectures to support data collection, storage, and analysis. Collaborated with data analysts, data scientists, and stakeholders to ensure reliable data flow and accessibility. Optimised ETL processes, integrating data from multiple sources, ensuring quality and governance, and implementing best practices for performance and security. Supported the deployment of analytics and reporting solutions, contributing to data-driven decision-making and efficiency. Significant Achievements: • Transitioned a legacy Oracle system to Spark-Scala on AWS EMR, enabling scalable analytics on health insurance claims and policy data, reducing processing time by 40%, and achieving total accuracy for downstream BI reporting in AWS S3. • Enhanced ETL workflows by developing Python-based AWS Lambda functions integrated with Glue and Athena, decreasing pipeline runtime by 30% and lowering infrastructure costs. • Designed customer lifecycle pipelines using Python, Spark, Pandas, and Boto3, processing 10M+ records a month, delivering accurate insights into customer onboarding, activity, and churn. • Tuned SQL queries to improve data extraction, transformation, and validation across large datasets, refining query reliability. • Joined forces with engineering and business teams to troubleshoot and resolve complex data issues, improving pipeline uptime to 99.95% and cutting incident response time by half. • Directed data operations by automating S3-to-local data transfers via AWS CLI for testing and ensuring version-controlled development using Bitbucket.

FIGmd, Inc.

Big Data Developer

FIGmd, Inc.

LinkedIn
2019-11 - 2022-5 · 2 yrs 7 mos

India

Developed and maintained large-scale data processing systems to manage and analyse high-volume, high-velocity datasets. Transformed data pipelines using big data technologies to ensure efficient data ingestion, transformation, and storage. Integrated diverse data sources, implemented quality checks, and ensured system scalability. Significant Achievements: • Developed the Datamart module for the Janus Project using Spark Scala, extracting and transforming clinical data from Hive and PostgreSQL, applying business logic, and persisting results in Cassandra with robust quality validations. • Augmented Spark jobs on Google Cloud Hadoop clusters, improving scalable processing speed by 35% through efficient logging (log4j). • Designed and executed ETL workflows in Informatica PowerCenter, migrating Oracle source data with advanced transformations (Expression, Lookup, Router, Update Strategy), error handling, and SCD Type 1 management, driving enriched datasets. • Led proof-of-concept initiatives to enhance Hadoop cluster performance and resource utilisation, achieving a 20% reduction in operating costs. • Developed Python modules for PostgreSQL data operations, including loading/dumping element mapping tables and validating Cassandra schema compatibility across database versions. • Applied Linux CLI and Shell scripting for automation and operational monitoring, reducing manual effort and saving up to 10 hours weekly. • Enhanced Cassandra security and performance by configuring authentication, integrating with Spark for authorised data access, and evolving schemas to support new product features. • Enabled advanced analytics by creating external Hive tables mapped to Google Cloud Storage (GCS) and leveraging BigQuery for structured data modelling and querying. • Streamlined release management using GitLab, ensuring 100% on-time deployments, improving release note accuracy.

Virasat Solutions

Web Developer

Virasat Solutions

LinkedIn
2017-12 - 2019-2 · 1 yr 3 mos

India

Enhanced WordPress search functionality by integrating the PubMed API, enabling users to access research articles and improve knowledge discovery. Designed dynamic, responsive websites using PHP, WordPress, Moodle, MySQL, and Drupal, providing feature-rich platforms to improve usability. Significant Achievements: • Successfully integrated Moodle and WordPress ecosystems, synchronising courses with WooCommerce products and implementing single sign-on (SSO), resulting in a seamless, frictionless user experience and increased course engagement. • Automated certification and support workflows by connecting Moodle course certificates to WordPress and integrating the Zoho Help Desk API, significantly reducing response times and improving user satisfaction. • Enhanced Drupal and WordPress platforms with advanced functionality, including dynamic pricing, custom page templates, and API/payment gateway integrations, enabling scalable, high-performance digital solutions that supported business growth.

Weavers Web

Web Developer

Weavers Web

LinkedIn
2017-7 - 2017-12 · 6 mos

India

Developed and launched custom WordPress websites, including e-commerce platforms with WooCommerce and integrated payment gateways, to enable seamless online transactions and drive higher sales conversions. Built and customised WordPress themes, plugins, and page templates, delivering tailored functionality, improved usability, and enhanced customer engagement across client sites. Significant Achievements: • Optimised and managed MySQL databases to support dynamic content, product catalogues, and course data, ensuring fast performance, high availability, and data integrity. • Delivered superior web solutions that significantly improved website performance, increased user engagement, and strengthened online presence.

Techwelfare

Developer

Techwelfare

LinkedIn
2014-9 - 2017-6 · 2 yrs 10 mos

India

Engineered dynamic websites using Core PHP and frameworks like CodeIgniter, delivering tailored CMS, e-commerce, and classified ad solutions. Integrated secure third-party services (mobile recharge, PayPal, SMS gateways), streamlining online payments and real-time notifications and increasing transaction rates. Significant Achievements: • Enhanced user engagement and retention by optimising front-end performance through responsive, dynamic interfaces built with Ajax, JavaScript, and jQuery, significantly improving load times and interaction smoothness. • Delivered measurable growth by developing and deploying high-performance e-commerce platforms on OpenCart and classified systems on OsClass, enabling clients to scale operations and increase online revenue. • Engineered and customised CMS solutions on WordPress and Joomla with optimised MySQL backends, improving site performance, simplifying content management, and ensuring long-term scalability for diverse client needs.

Education

NIIT Institute

NIIT Institute

LinkedIn

Computer Software Engineering

2008-6 - 2010-6 · 2 yrs 1 mo
Kuvempu University, Shankaraghatta, Shimoga

Kuvempu University, Shankaraghatta, Shimoga

LinkedIn

Information Technology

Hosenara Khatun's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.