What drew me to data science was realising that data makes patterns visible that would otherwise stay hidden, and that those patterns can support real decisions. I build ML pipelines and experiment with deep learning architectures where the goal is not just exploration, but findings that can actually be used, whether to improve a product, plan capacity, or inform a decision.
Languages & core: Python | SQL
ML / Data Science: scikit-learn | XGBoost | SARIMAX | Optuna | pandas | NumPy
Deep Learning: PyTorch | CNNs | RNNs | Transformers | Attention | Transfer Learning
Experiment Tracking: Weights & Biases
Automation: UiPath (RPA)
causal-inference-heart-catheter-mortality
Effect of right heart catheterization on 30-day mortality in 5,735 ICU patients from the SUPPORT study, estimated with a causal forest, doubly robust AIPW and overlap weights. Adjustment halves the raw difference from 7.4 to 3.5 percentage points, and a sensitivity analysis benchmarked against randomized trials shows that a hidden confounder 1.2 times as strong as the recorded diagnosis would close the remaining gap.
time-series-forecasting-pipeline
End-to-end time-series forecasting pipeline with a feature selection comparison (RFE, Lasso, Random Forest), expanding-window validation and five models: linear regression, Ridge, Lasso, SARIMAX and XGBoost, with hyperparameters tuned in Optuna.
causal-inference-smoking-weight
Causal effect of quitting smoking on 11-year weight gain in the NHEFS cohort, estimated with a causal forest, IPW and AIPW along an eight-step pipeline. The adjusted effect is 3.4 kg compared to 2.5 kg from the naive comparison, with IPW reproducing the published Hernán and Robins estimate and 5 of 5 robustness checks passed.
image-captioning-flickr8k
Five captioning architectures from CNN-LSTM to Transformer decoder on Flickr8k, evaluated with BLEU and METEOR across greedy and beam search, with attention visualisation.
cnn-architecture-study
Eight controlled CNN experiments covering depth, regularization, optimizer comparison, and transfer learning. Pretrained ResNet18 reaches 94.7% validation accuracy, compared to 63.2% from scratch.

