Hosenara Khatun
Senior Data Engineer @ DBS Bank
About
I’m an accomplished Data Engineer with proven expertise in designing, building, and optimising scalable ETL/ELT pipelines and modern cloud-based data platforms. I specialise in transforming legacy analytics systems into high-performance, cloud-native architectures that drive business intelligence, automation, and decision-making. Skilled in PySpark, Spark, Databricks, SQL, and Python, I have hands-on experience delivering end- to-end data solutions on AWS and GCP, including services such as S3, Glue, EMR, Lambda, Redshift, Athena, Dataproc, and BigQuery. I’m passionate about architecting data lakes and lakehouse environments that enable reliable, governed, and scalable analytics. My work focuses on data quality, lineage, and governance, leveraging tools like Collibra for metadata management and implementing validation and reconciliation frameworks to ensure data accuracy and trust. I’m also experienced in performance tuning, workflow automation, and CI/CD pipelines using Airflow, Jenkins, GitLab, and Bitbucket, enabling faster, automated, and reliable data delivery. With a strong foundation in big data processing, data warehousing, and distributed computing frameworks, I bring both technical depth and strategic vision to every project. My goal is to bridge data engineering and business impact, delivering clean, actionable, and analytics-ready datasets that empower organisations to make smarter decisions at scale. Core Skills : Data Engineering | ETL/ELT | PySpark | Spark | Databricks | AWS | GCP | SQL | Python | Hadoop | Hive | Data Lakehouse | Collibra | Airflow | Jenkins | GitLab | Data Governance | Metadata Management | CI/CD | Automation | Big Data | Data Quality | Performance Tuning Reach out to me to discuss exciting opportunities.
Singapore
Singapore
Information Technology & Services
Joomla!, WordPress Design, MySQL , AJAX, CodeIgniter, Apache Spark, Scala, Databricks, AWS, GIT, Hadoop, Hive, GCP, Python, Cassandra, Performance Tuning, Pandas (Software), Pyth (Programming Language), Data Warehousing, Airflow, Linux, Schema, MySQL, Continuous Integration (CI), Agile Methodologies, Data Lakes, Kubernetes, Data Management, Data Analysis
Experience

Senior Data Engineer
Singapore
Led the migration of banking analytics systems and designed multi-layered data pipelines, orchestrating scalable workloads and automating workflows to improve operational efficiency. Implemented systems to enhance reporting and communication, ensured data quality and governance for dashboards, collaborated on post-migration validation, and developed documentation to support knowledge sharing and onboarding. Significant Achievements: • Migrated a large-scale banking analytics application from SAS to PySpark, improving processing efficiency by 40%. • Achieved 99.9% reconciliation accuracy through rigorous validation. • Built multi-layered data pipelines in Databricks using Medallion Architecture (Bronze, Silver, Gold), accelerating ingestion and transformation speed by 35% for millions of records. • Developed and orchestrated workloads on Kubernetes (AWS EKS) using Application Load Controller (ALC) strategies, ensuring scalability, reliability, and cost-efficient resource utilisation. • Automated workflows with Airflow and Jenkins, reducing manual intervention. • Implemented automated email alerting systems in the ADA Data Ops Platform, reducing reporting delays and enabling proactive stakeholder communication. • Validated and enhanced data quality for Control Tower dashboards in QlikView/QlikSense, achieving close to total accuracy and improving metadata governance with Collibra. • Conducted post-migration validation with Superset in collaboration with business analysts, resolving significant levels of critical data issues within SLA. • Leveraged Python (Pandas) for data manipulation and ad hoc analyses, providing rapid insights to meet dynamic business needs. • Created comprehensive documentation of data engineering workflows and infrastructure, fostering for knowledge sharing and accelerating new engineer on boarding by 20%.

Data Engineer
India
Responsible for designing, building, and maintaining scalable pipelines and architectures to support data collection, storage, and analysis. Collaborated with data analysts, data scientists, and stakeholders to ensure reliable data flow and accessibility. Optimised ETL processes, integrating data from multiple sources, ensuring quality and governance, and implementing best practices for performance and security. Supported the deployment of analytics and reporting solutions, contributing to data-driven decision-making and efficiency. Significant Achievements: • Transitioned a legacy Oracle system to Spark-Scala on AWS EMR, enabling scalable analytics on health insurance claims and policy data, reducing processing time by 40%, and achieving total accuracy for downstream BI reporting in AWS S3. • Enhanced ETL workflows by developing Python-based AWS Lambda functions integrated with Glue and Athena, decreasing pipeline runtime by 30% and lowering infrastructure costs. • Designed customer lifecycle pipelines using Python, Spark, Pandas, and Boto3, processing 10M+ records a month, delivering accurate insights into customer onboarding, activity, and churn. • Tuned SQL queries to improve data extraction, transformation, and validation across large datasets, refining query reliability. • Joined forces with engineering and business teams to troubleshoot and resolve complex data issues, improving pipeline uptime to 99.95% and cutting incident response time by half. • Directed data operations by automating S3-to-local data transfers via AWS CLI for testing and ensuring version-controlled development using Bitbucket.

Big Data Developer
India
Developed and maintained large-scale data processing systems to manage and analyse high-volume, high-velocity datasets. Transformed data pipelines using big data technologies to ensure efficient data ingestion, transformation, and storage. Integrated diverse data sources, implemented quality checks, and ensured system scalability. Significant Achievements: • Developed the Datamart module for the Janus Project using Spark Scala, extracting and transforming clinical data from Hive and PostgreSQL, applying business logic, and persisting results in Cassandra with robust quality validations. • Augmented Spark jobs on Google Cloud Hadoop clusters, improving scalable processing speed by 35% through efficient logging (log4j). • Designed and executed ETL workflows in Informatica PowerCenter, migrating Oracle source data with advanced transformations (Expression, Lookup, Router, Update Strategy), error handling, and SCD Type 1 management, driving enriched datasets. • Led proof-of-concept initiatives to enhance Hadoop cluster performance and resource utilisation, achieving a 20% reduction in operating costs. • Developed Python modules for PostgreSQL data operations, including loading/dumping element mapping tables and validating Cassandra schema compatibility across database versions. • Applied Linux CLI and Shell scripting for automation and operational monitoring, reducing manual effort and saving up to 10 hours weekly. • Enhanced Cassandra security and performance by configuring authentication, integrating with Spark for authorised data access, and evolving schemas to support new product features. • Enabled advanced analytics by creating external Hive tables mapped to Google Cloud Storage (GCS) and leveraging BigQuery for structured data modelling and querying. • Streamlined release management using GitLab, ensuring 100% on-time deployments, improving release note accuracy.

Web Developer
India
Enhanced WordPress search functionality by integrating the PubMed API, enabling users to access research articles and improve knowledge discovery. Designed dynamic, responsive websites using PHP, WordPress, Moodle, MySQL, and Drupal, providing feature-rich platforms to improve usability. Significant Achievements: • Successfully integrated Moodle and WordPress ecosystems, synchronising courses with WooCommerce products and implementing single sign-on (SSO), resulting in a seamless, frictionless user experience and increased course engagement. • Automated certification and support workflows by connecting Moodle course certificates to WordPress and integrating the Zoho Help Desk API, significantly reducing response times and improving user satisfaction. • Enhanced Drupal and WordPress platforms with advanced functionality, including dynamic pricing, custom page templates, and API/payment gateway integrations, enabling scalable, high-performance digital solutions that supported business growth.

Web Developer
India
Developed and launched custom WordPress websites, including e-commerce platforms with WooCommerce and integrated payment gateways, to enable seamless online transactions and drive higher sales conversions. Built and customised WordPress themes, plugins, and page templates, delivering tailored functionality, improved usability, and enhanced customer engagement across client sites. Significant Achievements: • Optimised and managed MySQL databases to support dynamic content, product catalogues, and course data, ensuring fast performance, high availability, and data integrity. • Delivered superior web solutions that significantly improved website performance, increased user engagement, and strengthened online presence.

Developer
India
Engineered dynamic websites using Core PHP and frameworks like CodeIgniter, delivering tailored CMS, e-commerce, and classified ad solutions. Integrated secure third-party services (mobile recharge, PayPal, SMS gateways), streamlining online payments and real-time notifications and increasing transaction rates. Significant Achievements: • Enhanced user engagement and retention by optimising front-end performance through responsive, dynamic interfaces built with Ajax, JavaScript, and jQuery, significantly improving load times and interaction smoothness. • Delivered measurable growth by developing and deploying high-performance e-commerce platforms on OpenCart and classified systems on OsClass, enabling clients to scale operations and increase online revenue. • Engineered and customised CMS solutions on WordPress and Joomla with optimised MySQL backends, improving site performance, simplifying content management, and ensuring long-term scalability for diverse client needs.
Hosenara Khatun's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.


