Downloads ยท 30 days
0
Cristi232/Animal_Faces
Animal_Faces is a machine learning model from Cristi232. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Animal Faces\ DeepLearning ComputerVision Project
Downloads ยท 30 days
0
Access
Public
Updated Mar 13, 2025
Repo size
3.8 GB
Likes
0
Public
Click a slice to open those files.
.pth3.4 GB ยท 88%
From the Hugging Face model README
Animal Faces
DeepLearning ComputerVision Project
FMI, Master An II, BDTS (505), DeepLearning
Our source dataset is called Animal Faces and it contains aproximate of
16,130 high-quality images at 512ร512 resolution (cat, dog, wild).
Source link
https://www.kaggle.com/datasets/andrewmvd/animal-faces/data
This readme.md is a copy form Animal_Faces.docx in the same folder

The collection of images is organized in folders (cat, dog, wild), the kaggle source has a training and validation folder. The images are majority clean headshots of cats, dogs, foxes, wolfs, tigers, lions, other felines, in jpg format.


The dataset provided contained only training and validation folders, here is where I have made a slight modification, by creating a test folder with 500 each images moved from train.



Kaggle source Local Setup Final Local Dataset
Train and evaluate 3 deep models (EfficientNet, ViT, DenseNet)
Other adjacent objectives:
Train a Fourth model
Run predictions on the test dataset
Tune of parameters, grid search, gather graphs and confusion matrix
What it is, how it works.
In traditional neural networks each layer only receives input from the previous layer. Example diagram below.

Densely Connected Convolutional Network (DenseNet) is a deep learning architecture where each layer gets input from all preceding layers, designed for image classification and other computer vision tasks.
Example diagram below.

This design provides multiple benefits: fewer parameters, greater computational efficiency, and enhanced generalization.
Our model variant is DenseNet121, the 121 part comes form the fact that it has a depth of 121(layers).
In our first iteration and practice runs, there was used the pretrained model. But for the following presentation and charts, we have a untrained model in all scenarios.

Model training was done on a gaming laptop with the following configuration

In the DenseNet_train.py file, I have added system probing functions with the purpose to gain knowledge on the system's capabilities. Example: Is Gpu available, Gpu's limit for batch sizes test, number of workers test and others. Not all are presented in the final code.
HyperParameters
Our model uses the following hyperparameters:
[Learning Rate (lr)]{.underline} -- Controls how fast the model updates its weights during training.
[Batch Size (batch_size)]{.underline} -- Determines how many training samples are processed in one iteration.
[Weight Decay (weight_decay)]{.underline} -- A regularization technique that prevents overfitting by penalizing large weight values, ensuring smoother and more generalizable models.
[Optimizer (optimizer)]{.underline} -- Defines how weights are updated based on gradients. AdamW, used in our case.
[Learning Rate Scheduler (scheduler)]{.underline} -- Dynamically adjusts the learning rate during training to maintain stable convergence and avoid premature stagnation.
[Epochs (epochs)]{.underline} -- The number of times the model goes through the entire dataset. More epochs typically improve learning but can lead to overfitting if too high.
Hyperparameter Tuning, Grid Search, and Performance Evaluation

These hyperparameters are tested to find the best combination.
After we define our hyper parameter grid, we reach make our first optimisation in the data loader. Specifically we resize the image to 128x128 since it is provided as 512x512. Our system can not process efficiently those sizes so in our transformations we perform this resize. Other transformations to be mentioned, is that we normalize the images by the mean and standard deviation.
We normalize by scaling pixel values from [0, 255] = [0, 1] to the pythorch [C, H, W] format.

As next step we use only the training and validation dataset and keep
the test only for prediction testing.
And begin the training to find the best combination of hyperparameters.

Grid contains 3 lr, 3 batch_sizes, 2 weight_decays. Means 3 x 3 x 2 = 18
Total combinations
Total training duration over 2.5h
Other hyperparameters

Adam optimizer with weight decay regularization
A learning rate scheduler that reduces the learning rate after a set number of epochs
GradScales is an automatic mixed precision (AMP) tool in PyTorch. It reduces memory usage by using FP16 (half-precision) floating-point calculations where possible.

We provided the batch sizes 32, 64, 128 and the learning rates 0.001, 0.0005, 0.0001.
Observations:

Other generated charts based on our csv.

Prediction


Test Folder
===== test =====
cat = 566
dog = 525
wild = 525
Total in test = 1616
Also found cat image that model predicted to be a dog ๐

Other charts

What it is, how it works.
ViT (Vision Transformer) is a deep learning model designed for image recognition tasks, it applies Transformer architectures (originally designed for NLP) to images.
Works by spliting the image into fixed-size patches (e.g., 16x16 or 32x32 pixels), then each patch is flattened into a 1D vector and projected into a higher-dimensional space using a linear transformation.
Example diagram below.

Our model variant is vit_b_16, the 16 part comes form the fact that it patches the image into 16x16 and the base version has a depth of 86M parameters.
This design provides multiple benefits: performs well on large datasets,
better than CNNs on complex images (once trained properly).
The challenge on a untrained model is that it requires multiple epochs to achive better accuracy and a large dataset. Also a different strategy for hyperparameters.
Since the image sizes will me double, 224 instead of the 128 as before we need to reduce the batch size for faster processing.
Pretrained models likely will perform very well no matter the size of the epochs but we will adjust the following:
On the scheduler, adjust the steps to 1 and the gamma to 0.25, meaning that it will learn/adjust on each epoch.

Because of the larger image sizes and system performace, we will use a
reduced hyperparameter grid. This will allow us to spot the issues, tune
the parameters and restart if needed.

Observations made:

Model starts with a low accuracy as expected (55%), but improves by 8%,
7%, 5% per epoch. Progression indicates fast gains early, diminishing
returns later.

We provided the batch sizes 8, 32 and the learning rates 3e-05, 0.0001.

Prediction

Other charts


What it is, how it works.
EfficientNet is a family of convolutional neural networks (CNNs). It is designed to achieve high accuracy while being computationally efficient.
Unlike traditional CNN architectures, which are scaled manually by increasing depth (more layers), width (more channels), or input resolution (larger images), EfficientNet introduces a unique scaling method called "Compound Scaling" that balances all three dimensions efficiently.
Example diagram below.

Our model variant is EfficientNet-B0, the B0 part comes from being the smallest base version and has a depth of 5.3M parameters.

This design provides multiple benefits: Outperforms ResNet, Inception,
and DenseNet with fewer parameters, EfficientNet achieves higher
accuracy with lower computational cost.
Will revert to the image sizes of 128 as before and the previous batch sizes for faster processing.
Pretrained models likely will perform very well no matter the size of the epochs but we will adjust the following:
On the scheduler, adjust the steps to 1 and the gamma to 0.15, meaning that it will learn/adjust on each epoch.

Learning rates will be very different, such that we will see the large
variations in the results.


Model starts with a high accuracy as expected (71%), but improves up to 90% .
We provided the batch sizes 32, 128 and the learning rates 0.001, 0.0005.


Other generated charts based on our csv.
Prediction

Other charts


Pretrained model

Fixed hyperparameters


