Introduction to Data Science – University of Tehran – Department of Computer Engineering
This repository contains Deep Learning Workflows for Data Science, a Python and Jupyter Notebook implementation of three deep learning tasks: MLP regression on California Housing data, CNN-based fruit and vegetable image classification with robustness and Grad-CAM analysis, and RNN/LSTM knowledge tracing on ASSISTments2017. This project was developed as the Fourth Assignment for the Introduction to Data Science course at the University of Tehran.
The project follows a complete experimental pipeline, including dataset loading, exploratory data analysis, train-only preprocessing, model design, training or checkpoint loading, quantitative evaluation, visualization, robustness analysis, explainability, and final interpretation.
The repository preserves the original assignment PDF, the submitted analytical report, standardized notebooks, trained checkpoints, generated figures, and machine-readable experiment artifacts.
- ✅ Train and compare baseline and feature-engineered MLP models for California Housing price regression.
- ✅ Build a four-block CNN image classifier with train-only normalization, augmentation, robustness testing, and Grad-CAM explainability.
- ✅ Implement RNN and LSTM knowledge-tracing models using chronological student interaction sequences from ASSISTments2017.
- ✅ Evaluate the models with task-appropriate metrics, including MAE, RMSE, R², accuracy, precision, recall, F1-score, AUC, confusion matrices, and robustness drops.
- ✅ Preserve reproducible notebooks, trained checkpoints, generated plots, JSON metrics, and a final PDF report.
The first notebook performs tabular regression on California Housing data. It inspects distributions, correlations, outliers, and target behavior; splits the data into train, validation, and test subsets; fits scalers on training data only; trains a baseline Multi-Layer Perceptron (MLP); then trains an enhanced MLP using engineered ratio, logarithmic, and geographic-distance features.
The second notebook implements a ten-class fruit and vegetable classifier. It analyzes RGB channel statistics, computes class-level color-distance matrices, uses train-only channel normalization, applies training-only augmentation, trains a modular Convolutional Neural Network (CNN), and evaluates the model under baseline, grayscale, channel-swapped, and noisy test scenarios. Grad-CAM visualizations are used to inspect correct and incorrect predictions.
The third notebook models student performance as a chronological sequence-learning problem. It uses student-level splitting, training-only categorical encoders and numerical scaling, padded fixed-length interaction sequences, shifted previous-correctness inputs to avoid target leakage, and recurrent models for binary correctness prediction. A bonus LSTM model is trained and compared with the vanilla RNN.
The project reports model-level metrics, learning curves, confusion matrices, residual plots, robustness comparison tables, ROC curves, and qualitative prediction examples. The final report summarizes the empirical results and discusses limitations such as censored housing targets, color dependence in image classification, and the difficulty of modeling long-term educational dependencies.
The project is organized as follows:
data-science-assignment-4/
├── artifacts/ # JSON histories and machine-readable metric summaries
├── data/ # Included tabular datasets and Task 2 dataset instructions
├── description/ # Original assignment specification
├── figures/ # Generated plots, matrices, curves, Grad-CAM outputs, and visual results
├── models/ # Trained PyTorch checkpoints retained for reproducibility
├── notebooks/ # Standardized Jupyter notebooks for Tasks 1, 2, and 3
├── reports/ # Submitted final PDF report
├── .gitignore # Git ignore rules for Python, Jupyter, and local datasets
├── Makefile # Convenience commands for installation and notebook verification
├── README.md # Project documentation
└── requirements.txt # Python dependencies
Create a Python environment and install the dependencies:
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtLaunch the notebooks from the repository root:
jupyter labTask 1 and Task 3 can be executed with the included CSV files and checkpoints:
make verify-task1
make verify-task3Task 2 requires the external image dataset. Place it under data/fruits_and_vegetables/ or set TASK2_DATASET_ROOT before opening the notebook:
export TASK2_DATASET_ROOT=/path/to/fruits_and_vegetables
jupyter lab| Task | Model / Scenario | Main Result |
|---|---|---|
| Task 1 | Baseline MLP | MAE = 0.4118, RMSE = 0.5774, R² = 0.7479 |
| Task 1 | Enhanced MLP | MAE = 0.3526, RMSE = 0.5049, R² = 0.8072 |
| Task 2 | CNN baseline color test | Accuracy = 91.92%, Macro F1 = 0.9193 |
| Task 2 | CNN grayscale challenge | Accuracy = 14.14%, accuracy drop = 77.78 percentage points |
| Task 2 | CNN channel-swap challenge | Accuracy = 25.25%, accuracy drop = 66.67 percentage points |
| Task 3 | RNN knowledge tracing | Accuracy = 0.8020, AUC = 0.8895, F1 = 0.7808 |
| Task 3 | LSTM bonus model | Accuracy = 0.7963, AUC = 0.8867, F1 = 0.7665 |
The metrics above are taken from the JSON artifacts in artifacts/. The Task 2 image dataset is not included, but the trained checkpoint, figures, and saved metrics are retained for inspection.
- The Task 2 image archive is intentionally excluded because of size; see
data/task2_dataset_readme.mdfor the expected directory layout. - The notebooks use
FORCE_RETRAIN = Falseby default, so existing checkpoints are loaded when available. - The final PDF report is preserved in
reports/final_report.pdf; small numeric differences may appear between the static report and the latest notebook artifact files because model training is stochastic.
This project was developed as a team effort for the Introduction to Data Science course at the University of Tehran.