Kagnana ITH Portfolio
← Go back to projects

CHU Big data

October 2025
Big Data Data Lakehouse Delta Lake Apache Spark Trino PostgreSQL MinIO Apache Superset Docker Python SQL FISA INFO A4

Healthcare data lakehouse architecture for analyzing medical consultation and hospitalization rates using Delta Lake, Spark, and Trino.

CHU Big data

About the project

A comprehensive Big Data architecture for healthcare analytics that processes and analyzes medical data from various sources. The system implements a Lakehouse pattern using Delta Lake as the storage layer, Apache Spark for distributed data processing, and Trino as the SQL query engine. PostgreSQL serves as the initial data source, with MinIO providing object storage simulating a data lake. The architecture enables analysis of consultation rates, hospitalization trends, patient demographics, and satisfaction metrics across different regions and time periods. The solution includes data ingestion pipelines, transformation workflows, and visualization dashboards in Apache Superset. All components are containerized using Docker for easy deployment and scalability.

Key Features

  • Lakehouse architecture with Delta Lake storage
  • Distributed data processing with Apache Spark
  • SQL query engine via Trino
  • PostgreSQL for initial data ingestion
  • MinIO object storage simulating data lake
  • Containerized deployment with Docker
  • Data analysis on consultation rates by region/time
  • Hospitalization trend analysis with demographic breakdowns
  • Satisfaction metrics visualization
  • Real-time data processing pipeline
  • Secure data storage and access control
  • Scalable architecture for healthcare analytics
  • Automated data ingestion workflows
  • Interactive dashboards in Apache Superset

Used technologies

Delta Lake Apache Spark Trino PostgreSQL MinIO Apache Superset Python Docker SQL JDBC/ODBC

Gallery

Screenshot 1
Screenshot 2
Screenshot 3