Downloads · 30 days
11
33% of all-time downloads
BechirTrabelsi1/optLLM-small
optLLM-small is a machine learning model from BechirTrabelsi1. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
11
33% of all-time downloads
All-time downloads
33
Public
Parameters
77M
309 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors308 MB · 99%
From the Hugging Face model README
This model is a fine-tuned version of the google/flan-t5-small on a synthetic dataset designed for supply chain optimization problems.
The primary goal of this model is to classify the type of supply chain optimization problem (e.g., Newsvendor, EOQ, VRP) described in the input and extract the relevant variables necessary for solving these problems.
The idea is that anyone, even with no background in optimisation , can interact with a solver and solve his problem: for example imagine a salesperson in a store with no experience in optimization, he could describe their problem in natural language and obtain an optimal solution
google/flan-t5-smallThis model is designed to be used in the following scenarios:
The model can classify and extract variables for the following supply chain optimization problems:
The model was fine-tuned using a synthetic dataset with the following steps:
google/flan-t5-small model was fine-tuned on this dataset, with a focus on accurately classifying the problem type and extracting relevant variables.The model has shown good performance on the synthetic dataset, with a high accuracy in classifying problem types and extracting variables. However, performance on real-world data may vary, and further fine-tuning or validation on real-world datasets is recommended.
I tried to make the model dynamic and agnostic the number of params/variables in the problem , however due to a bug I’m still trying to identify the model always ignores the last product in the newsvendor problem (open to ideas if someone has a clue why is this happening )
Next steps would be fixing the bug with the newsvendor problem , and developing the routing interface to communicate with the appropriate solvers to find an optimal solution for the user
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| 0.0161 | 1.0 | 4000 | 0.0046 |