Downloads · 30 days
8
20% of all-time downloads
ldp72/Test-SmolLM-Marcel-codecarbon
Test-SmolLM-Marcel-codecarbon is a text generation model from ldp72. Use it when you need the model to write or continue text. It is set up for transformers.
This model was finetuned by performing instruct tuning on Telco domain datatsets.
Downloads · 30 days
8
20% of all-time downloads
All-time downloads
41
Public
Parameters
135M
538 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors269 MB · 98%
From the Hugging Face model README
This model was finetuned by performing instruct tuning on Telco domain datatsets.
This model can be used with the transformers library using pipeline abstraction as follows:
import torch
from transformers import pipeline
model_id = "ldp72/Test-SmolLM-Marcel-codecarbon"
pipe = pipeline(
"text-generation",
model=model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
messages = [
{"role": "system", "content": "You are chatbot specialized on Telco domain."},
{"role": "user", "content": "Can you give a sample of your specialized knowledge?"},
]
outputs = pipe(
messages,
max_new_tokens=256,
)
print(outputs[0]["generated_text"][-1])
[More Information Needed]
[More Information Needed]
[More Information Needed]
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
Use the code below to get started with the model.
[More Information Needed]
This model was finetuned with Orange internal fine tuning tools with the Docker Image tagged 0.1.2 in the registry and the following configuration file:
data:
dataset_name:
train:
- path: telco-lm/arxiv-abstract-generation-telco-instructions
revision: legacy
- path: telco-lm/synthetic-dsp.stackexchange.com-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-networkengineering.stackexchange.com-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-security.stackexchange.com-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-3gpp-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-5gamericas-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-huawei-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-itu-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-mef-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-ngmn-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-rfc-multi-task-telco-instructions
revision: legacy
- path: telco-lm/teleqna-mcqa-cot-telco-instructions
revision: legacy
- path: telco-lm/tii-huawei-qa-open-qa-telco-instructions
revision: legacy
validation_abstract_generation:
- path: telco-lm/arxiv-abstract-generation-telco-instructions
revision: legacy
split: validation
validation_general:
- path: telco-lm/slim-orca-multi-task-general-instructions
revision: legacy
split: validation
validation_synthetic:
- path: telco-lm/synthetic-dsp.stackexchange.com-multi-task-telco-instructions
revision: legacy
split: validation
- path: telco-lm/synthetic-security.stackexchange.com-multi-task-telco-instructions
revision: legacy
split: validation
- path: telco-lm/synthetic-networkengineering.stackexchange.com-multi-task-telco-instructions
revision: legacy
split: validation
- path: telco-lm/synthetic-technical-rfc-multi-task-telco-instructions
revision: legacy
split: validation
- path: telco-lm/synthetic-technical-3gpp-multi-task-telco-instructions
revision: legacy
split: validation
- path: telco-lm/synthetic-technical-5gamericas-multi-task-telco-instructions
revision: legacy
split: validation
- path: telco-lm/synthetic-technical-itu-multi-task-telco-instructions
revision: legacy
split: validation
- path: telco-lm/synthetic-technical-mef-multi-task-telco-instructions
revision: legacy
split: validation
- path: telco-lm/synthetic-technical-huawei-multi-task-telco-instructions
revision: legacy
split: validation
- path: telco-lm/synthetic-technical-ngmn-multi-task-telco-instructions
revision: legacy
split: validation
validation_telco_qa:
- path: telco-lm/tii-huawei-qa-open-qa-telco-instructions
revision: legacy
split: validation
validation_telco_qcm:
- path: telco-lm/teleqna-mcqa-cot-telco-instructions
revision: legacy
split: validation
debug: true
implementation_name: instructions
description:
contributors:
- email: loic.fosse@orange.com
first_name: Loïc
last_name: Fosse
- email: lionel.delphinpoulat@orange.com
first_name: Lionel
last_name: Delphin-Poulat
- email: ismael.rousseau@orange.com
first_name: Ismaël
last_name: Rousseau
domain: Telco
languages:
- en
model_name: ldp72/Test-SmolLM-Marcel-codecarbon
image:
version: 0.1.2
model:
attn_implementation: flash_attention_2
chat_template_tokenizer: HuggingFaceTB/SmolLM-135M-Instruct
model_name_or_path: HuggingFaceTB/SmolLM-135M-Instruct
trust_remote_code: true
training:
bf16: true
dataloader_num_workers: 4
dataloader_persistent_workers: true
dataloader_pin_memory: true
dataloader_prefetch_factor: 2
deepspeed: /config/zero3.json
disable_tqdm: true
eval_accumulation_steps: 1
eval_steps: 10
eval_strategy: steps
fp16: false
gradient_accumulation_steps: 2
gradient_checkpointing: true
group_by_length: false
learning_rate: 2.0e-05
log_level: debug
logging_dir: /outputs/Telco-SmolLM-135-Instruct-it-test-codecarbon-process-push/logs
logging_steps: 10
lr_scheduler_type: cosine
max_grad_norm: 1.0
max_steps: -1
num_train_epochs: 2
optim: paged_adamw_32bit
output_dir: /outputs/Telco-SmolLM-135-Instruct-it-test-codecarbon-process-push
per_device_eval_batch_size: 2
per_device_train_batch_size: 2
push_to_hub: false
report_to: tensorboard
save_steps: 0
save_strategy: epoch
save_total_limit: 1
seed: 42
torch_compile: false
training_type: instruct-tuning
use_liger_kernel: false
warmup_ratio: 0.05
weight_decay: 0.1
The model was trained on 1 gpus with at least 40GB on each gpu.
The model was trained using deepspeed with the following configuration file:
{
"fp16": {
"enabled": "auto",
"loss_scale": 0,
"loss_scale_window": 1000,
"initial_scale_power": 16,
"hysteresis": 2,
"min_loss_scale": 1
},
"bf16": {
"enabled": "auto"
},
"zero_optimization": {
"stage": 3,
"offload_optimizer": {
"device": "cpu",
"pin_memory": true
},
"offload_param": {
"device": "cpu",
"pin_memory": true
},
"overlap_comm": true,
"contiguous_gradients": true,
"sub_group_size": "1e9",
"reduce_bucket_size": "auto",
"stage3_prefetch_bucket_size": "auto",
"stage3_param_persistence_threshold": "auto",
"stage3_max_live_parameters": "1e9",
"stage3_max_reuse_distance": "1e9",
"stage3_gather_16bit_weights_on_model_save": true
},
"gradient_accumulation_steps": "auto",
"gradient_clipping": "auto",
"steps_per_print": 2000,
"train_batch_size": "auto",
"train_micro_batch_size_per_gpu": "auto",
"wall_clock_breakdown": false
}
This model was trained on the following datasets:
- path: telco-lm/arxiv-abstract-generation-telco-instructions
revision: legacy
- path: telco-lm/synthetic-dsp.stackexchange.com-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-networkengineering.stackexchange.com-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-security.stackexchange.com-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-3gpp-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-5gamericas-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-huawei-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-itu-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-mef-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-ngmn-multi-task-telco-instructions
revision: legacy
- path: telco-lm/synthetic-technical-rfc-multi-task-telco-instructions
revision: legacy
- path: telco-lm/teleqna-mcqa-cot-telco-instructions
revision: legacy
- path: telco-lm/tii-huawei-qa-open-qa-telco-instructions
revision: legacy
[More Information Needed]
SFTTrainer,other parameters were set as default:bf16: true
dataloader_num_workers: 4
dataloader_persistent_workers: true
dataloader_pin_memory: true
dataloader_prefetch_factor: 2
deepspeed: /config/zero3.json
disable_tqdm: true
eval_accumulation_steps: 1
eval_steps: 10
eval_strategy: steps
fp16: false
gradient_accumulation_steps: 2
gradient_checkpointing: true
group_by_length: false
learning_rate: 2.0e-05
log_level: debug
logging_dir: /outputs/Telco-SmolLM-135-Instruct-it-test-codecarbon-process-push/logs
logging_steps: 10
lr_scheduler_type: cosine
max_grad_norm: 1.0
max_steps: -1
num_train_epochs: 2
optim: paged_adamw_32bit
output_dir: /outputs/Telco-SmolLM-135-Instruct-it-test-codecarbon-process-push
per_device_eval_batch_size: 2
per_device_train_batch_size: 2
push_to_hub: false
report_to: tensorboard
save_steps: 0
save_strategy: epoch
save_total_limit: 1
seed: 42
torch_compile: false
use_liger_kernel: false
warmup_ratio: 0.05
weight_decay: 0.1
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
emissions.csv (emissions were computed using codecarbon)[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
BibTeX:
[More Information Needed]
APA:
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Thanks to Loïc Fosse, Lionel Delphin-Poulat, Ismaël Rousseau for adding this model.