Downloads · 30 days
29
25% of all-time downloads
lemms/openllm-small-extended-7k
openllm-small-extended-7k is a text generation model from lemms. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as gpl-3.0.
This is the OpenLLM Small Extended 7K model, a 35.8M parameter GPT-style language model trained for 7,000 steps on Wikipedia passages from the SQuAD dataset. This model represents the latest iteration of our small mod…
Downloads · 30 days
29
25% of all-time downloads
All-time downloads
114
Public
Repo size
169 MB
Likes
0
Public
Click a slice to open those files.
.bin168 MB · 100%
From the Hugging Face model README
This is the OpenLLM Small Extended 7K model, a 35.8M parameter GPT-style language model trained for 7,000 steps on Wikipedia passages from the SQuAD dataset. This model represents the latest iteration of our small model architecture with extended training.
huggingface/
├── config.json # Model configuration
├── generation_config.json # Generation parameters
├── pytorch_model.bin # Model weights (161MB)
├── tokenizer_config.json # Tokenizer configuration
├── tokenizer.model # SentencePiece tokenizer
└── load_hf_model.py # Loading script
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
# Load model and tokenizer
model_name = "path/to/huggingface"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)
# Generate text
prompt = "The history of artificial intelligence"
inputs = tokenizer(prompt, return_tensors="pt")
with torch.no_grad():
outputs = model.generate(
inputs.input_ids,
max_new_tokens=100,
temperature=0.7,
do_sample=True,
pad_token_id=tokenizer.pad_token_id
)
generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(generated_text)
from load_hf_model import load_openllm_model
# Load the model using our custom loader
model, tokenizer = load_openllm_model("path/to/huggingface")
# Generate text
prompt = "Explain quantum computing in simple terms"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(
inputs.input_ids,
max_new_tokens=150,
temperature=0.8,
top_p=0.9
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
# Start the FastAPI inference server
python core/src/inference_server.py \
--model_path exports/huggingface-7k/huggingface \
--port 8000
# Make API calls
curl -X POST "http://localhost:8000/generate" \
-H "Content-Type: application/json" \
-d '{
"prompt": "The future of renewable energy",
"max_tokens": 100,
"temperature": 0.7
}'
{
"max_length": 512,
"max_new_tokens": 256,
"temperature": 0.7,
"top_k": 40,
"top_p": 0.9,
"do_sample": true,
"pad_token_id": 0,
"eos_token_id": 1,
"bos_token_id": 2
}
{
"vocab_size": 32000,
"n_layer": 6,
"n_head": 8,
"n_embd": 512,
"block_size": 1024,
"dropout": 0.1,
"bias": true
}
# Test the model with a simple prompt
test_prompt = "Hello, how are you today?"
inputs = tokenizer(test_prompt, return_tensors="pt")
with torch.no_grad():
outputs = model.generate(
inputs.input_ids,
max_new_tokens=20,
temperature=0.7
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(f"Input: {test_prompt}")
print(f"Output: {response}")
| Model | Parameters | Training Steps | Context Length | Use Case |
|---|---|---|---|---|
| Small 4K | 35.8M | 4,000 | 1,024 | Basic text generation |
| Small 6K | 35.8M | 6,000 | 1,024 | Improved coherence |
| Small 7K | 35.8M | 7,000 | 1,024 | Extended training |
This model is dual-licensed:
See LICENSE and docs/LICENSES.md for full license information.
We welcome contributions to improve the model! Please see:
docs/CONTRIBUTING.md for contribution guidelinesdocs/CODE_OF_CONDUCT.md for community standardsFor questions, issues, or commercial licensing:
docs/ directoryAuthor: Louis Chua Bean Chong
Project: OpenLLM - Open Source Large Language Model
Version: 0.1.0
Last Updated: 2024