Sitelet https://github.com/mragetsars/Deep-Learning-Workflows-for-Data-Science
Skip to content

Latest commit

 

History

5 Commits

Folders and files

Repository files navigation

Deep Learning Workflows for Data Science

Introduction to Data Science – University of Tehran – Department of Computer Engineering

Language Notebook Framework Status

Overview

This repository contains Deep Learning Workflows for Data Science, a Python and Jupyter Notebook implementation of three deep learning tasks: MLP regression on California Housing data, CNN-based fruit and vegetable image classification with robustness and Grad-CAM analysis, and RNN/LSTM knowledge tracing on ASSISTments2017. This project was developed as the Fourth Assignment for the Introduction to Data Science course at the University of Tehran.

The project follows a complete experimental pipeline, including dataset loading, exploratory data analysis, train-only preprocessing, model design, training or checkpoint loading, quantitative evaluation, visualization, robustness analysis, explainability, and final interpretation.

The repository preserves the original assignment PDF, the submitted analytical report, standardized notebooks, trained checkpoints, generated figures, and machine-readable experiment artifacts.

Project Objectives

  • ✅ Train and compare baseline and feature-engineered MLP models for California Housing price regression.
  • ✅ Build a four-block CNN image classifier with train-only normalization, augmentation, robustness testing, and Grad-CAM explainability.
  • ✅ Implement RNN and LSTM knowledge-tracing models using chronological student interaction sequences from ASSISTments2017.
  • ✅ Evaluate the models with task-appropriate metrics, including MAE, RMSE, R², accuracy, precision, recall, F1-score, AUC, confusion matrices, and robustness drops.
  • ✅ Preserve reproducible notebooks, trained checkpoints, generated plots, JSON metrics, and a final PDF report.

Methodology

1️⃣ Task 1: MLP Regression on California Housing

The first notebook performs tabular regression on California Housing data. It inspects distributions, correlations, outliers, and target behavior; splits the data into train, validation, and test subsets; fits scalers on training data only; trains a baseline Multi-Layer Perceptron (MLP); then trains an enhanced MLP using engineered ratio, logarithmic, and geographic-distance features.

2️⃣ Task 2: CNN Classification with Robustness and Grad-CAM

The second notebook implements a ten-class fruit and vegetable classifier. It analyzes RGB channel statistics, computes class-level color-distance matrices, uses train-only channel normalization, applies training-only augmentation, trains a modular Convolutional Neural Network (CNN), and evaluates the model under baseline, grayscale, channel-swapped, and noisy test scenarios. Grad-CAM visualizations are used to inspect correct and incorrect predictions.

3️⃣ Task 3: RNN Knowledge Tracing

The third notebook models student performance as a chronological sequence-learning problem. It uses student-level splitting, training-only categorical encoders and numerical scaling, padded fixed-length interaction sequences, shifted previous-correctness inputs to avoid target leakage, and recurrent models for binary correctness prediction. A bonus LSTM model is trained and compared with the vanilla RNN.

4️⃣ Evaluation and Analysis

The project reports model-level metrics, learning curves, confusion matrices, residual plots, robustness comparison tables, ROC curves, and qualitative prediction examples. The final report summarizes the empirical results and discusses limitations such as censored housing targets, color dependence in image classification, and the difficulty of modeling long-term educational dependencies.

Repository Structure

The project is organized as follows:

data-science-assignment-4/
├── artifacts/            # JSON histories and machine-readable metric summaries
├── data/                 # Included tabular datasets and Task 2 dataset instructions
├── description/          # Original assignment specification
├── figures/              # Generated plots, matrices, curves, Grad-CAM outputs, and visual results
├── models/               # Trained PyTorch checkpoints retained for reproducibility
├── notebooks/            # Standardized Jupyter notebooks for Tasks 1, 2, and 3
├── reports/              # Submitted final PDF report
├── .gitignore            # Git ignore rules for Python, Jupyter, and local datasets
├── Makefile              # Convenience commands for installation and notebook verification
├── README.md             # Project documentation
└── requirements.txt      # Python dependencies

Setup & Usage

Create a Python environment and install the dependencies:

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Launch the notebooks from the repository root:

jupyter lab

Task 1 and Task 3 can be executed with the included CSV files and checkpoints:

make verify-task1
make verify-task3

Task 2 requires the external image dataset. Place it under data/fruits_and_vegetables/ or set TASK2_DATASET_ROOT before opening the notebook:

export TASK2_DATASET_ROOT=/path/to/fruits_and_vegetables
jupyter lab

Results

Task Model / Scenario Main Result
Task 1 Baseline MLP MAE = 0.4118, RMSE = 0.5774, R² = 0.7479
Task 1 Enhanced MLP MAE = 0.3526, RMSE = 0.5049, R² = 0.8072
Task 2 CNN baseline color test Accuracy = 91.92%, Macro F1 = 0.9193
Task 2 CNN grayscale challenge Accuracy = 14.14%, accuracy drop = 77.78 percentage points
Task 2 CNN channel-swap challenge Accuracy = 25.25%, accuracy drop = 66.67 percentage points
Task 3 RNN knowledge tracing Accuracy = 0.8020, AUC = 0.8895, F1 = 0.7808
Task 3 LSTM bonus model Accuracy = 0.7963, AUC = 0.8867, F1 = 0.7665

The metrics above are taken from the JSON artifacts in artifacts/. The Task 2 image dataset is not included, but the trained checkpoint, figures, and saved metrics are retained for inspection.

Notes

  • The Task 2 image archive is intentionally excluded because of size; see data/task2_dataset_readme.md for the expected directory layout.
  • The notebooks use FORCE_RETRAIN = False by default, so existing checkpoints are loaded when available.
  • The final PDF report is preserved in reports/final_report.pdf; small numeric differences may appear between the static report and the latest notebook artifact files because model training is stochastic.

Contributors

This project was developed as a team effort for the Introduction to Data Science course at the University of Tehran.

About

Python and Jupyter Notebook implementation of three deep learning workflows: MLP regression, CNN image classification with robustness and Grad-CAM analysis, and RNN/LSTM knowledge tracing. Developed as the Fourth Assignment for the Introduction to Data Science course at the University of Tehran.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages