Downloads · 30 days
15
23% of all-time downloads
inference4j/tinyllama-1.1b-chat
tinyllama-1.1b-chat is a text generation model from inference4j. Use it when you need the model to write or continue text. It is set up for onnx. The card lists the license as apache-2.0.
ONNX export of TinyLlama-1.1B-Chat-v1.0 (1.1B parameters, FP16 weights) with KV cache support for efficient autoregressive generation.
Downloads · 30 days
15
23% of all-time downloads
All-time downloads
65
Public
Repo size
2.4 GB
Likes
0
Public
Click a slice to open those files.
.onnx_data2.4 GB · 100%
From the Hugging Face model README
ONNX export of TinyLlama-1.1B-Chat-v1.0 (1.1B parameters, FP16 weights) with KV cache support for efficient autoregressive generation.
Converted for use with inference4j, an inference-only AI library for Java.
try (var gen = OnnxTextGenerator.tinyLlama().build()) {
GenerationResult result = gen.generate("What is Java?");
System.out.println(result.text());
}
| Property | Value |
|---|---|
| Architecture | LlamaForCausalLM (1.1B parameters, 22 layers, 2048 hidden, 32 heads, 4 KV heads) |
| Task | Text generation (instruction-tuned, Zephyr chat template) |
| Precision | FP16 |
| Context length | 2048 tokens |
| Vocabulary | 32,000 tokens (SentencePiece BPE) |
| Chat template | Zephyr (`< |
| Original framework | PyTorch (transformers) |
| Export method | Hugging Face Optimum (with KV cache, FP16) |
This model is licensed under the Apache License 2.0. Original model by TinyLlama.