This project uses the Healthy Lifestyle Dataset to classify individuals as healthy or unhealthy based on their lifestyle habits.
We build an ML pipeline that includes preprocessing, data balancing, and model selection with cross-validation.
- Source: Healthy Lifestyle Dataset on Kaggle
- Description: Contains lifestyle-related features such as:
- Physical Activity
- Sleep Patterns
- Diet Habits
- Stress Levels
- Smoking & Alcohol Consumption
along with a target variable indicating health status.
- Task: Binary classification —
HealthyvsUnhealthy.
- Pipeline with Imbalanced Learning
- Used
imblearn.Pipelineto chain together transformations, balancing, and classification.
- Used
- Data Preprocessing
- Applied transformations (
TF) for handling categorical and numerical data.
- Applied transformations (
- Class Balancing
- Used
SMOTENor similar balancing technique to handle class imbalance.
- Used
- Model Selection
- Compared:
RandomForestClassifierAdaBoostClassifierXGBClassifier
- Compared:
- Hyperparameter Tuning
- Implemented
GridSearchCVwith multiple parameter grids.
- Implemented
- Model Evaluation
- Performed 5-Fold Cross-Validation for reliable performance estimation.
- Calculated mean accuracy across folds.
- Generated Confusion Matrix and Classifiaction report
- Final Prediction
- Trained the best model on the full training data and generated predictions on
X_test.
- Trained the best model on the full training data and generated predictions on
# Clone this repository
git clone https://github.com/Fatimibee/PyVerse
cd Machine Learning
# Install dependencies
pip install -r requirements.txt