Downloads · 30 days
1.2K
100% of all-time downloads
paradigma-inc/limite-1b-base
limite-1b-base is a text generation model from paradigma-inc. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Limite 1B - Base is the post-pretraining checkpoint of the Limite 1B model family.
Downloads · 30 days
1.2K
100% of all-time downloads
All-time downloads
1.2K
Public
Parameters
1B
2.1 GB on disk
Likes
8
Public
Click a slice to open those files.
.safetensors2.1 GB · 99%
How the weights are stored.
BF161B · 100%
From the Hugging Face model README
Limite 1B - Base is the post-pretraining checkpoint of the Limite 1B model family.
This model card focuses on loading and running the checkpoint with Hugging Face Transformers. For the latest checkpoint and its full model card, see Limite 1B - Violetto.
Limite 1B - Base supports inference through the standard Hugging Face Transformers APIs. The custom architecture code is downloaded from this repository, so loading the model requires trust_remote_code=True.
Validated on an NVIDIA H100 with Python 3.12 · PyTorch 2.11.0 (CUDA 13.0) · Transformers 5.6.2. Other version and hardware combinations have not yet been formally qualified. SDPA is the portable default.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "paradigma-inc/limite-1b-base"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype="auto",
device_map="cuda",
attn_implementation="sdpa",
)
messages = [{"role": "user", "content": "Solve: If x + 3 = 8, what is x?"}]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
with torch.inference_mode():
output_ids = model.generate(**inputs, max_new_tokens=256)
answer = tokenizer.decode(
output_ids[0, inputs.input_ids.shape[1]:],
skip_special_tokens=True,
)
print(answer)
The Transformers implementation supports one complete model replica per process. Automatic or explicit model sharding, tensor parallelism, and pipeline parallelism are not currently supported.
For offline multi-GPU inference, independent replicas can be launched with one process per GPU using device_map={"": local_rank}. For production serving, continuous batching, and multi-GPU request scheduling, use the Limite vLLM integration.
| Attention backend | Cache | Status |
|---|---|---|
| SDPA | DynamicCache | Supported |
| SDPA | mixed global/sliding-window StaticCache | Supported |
| SDPA | StaticCache + torch.compile | Supported |
| FlashAttention 2 | DynamicCache | Supported |
| FlashAttention 2 | StaticCache | Unsupported; rejected with an explicit error |
FlashAttention 2 with StaticCache is intentionally rejected because that combination does not produce numerically correct logits for Limite's hybrid local/global attention layout. Use SDPA with StaticCache, or FlashAttention 2 with DynamicCache.
FlashAttention 2 was validated through Transformers' kernels-community/flash-attn2 integration with kernels==0.12.3. The separately installed native flash_attn package has not been independently qualified.
Official support currently covers inference. The Transformers implementation is differentiable and exposes the standard causal-language-model loss, but full training, gradient checkpointing, PEFT/LoRA, and distributed-training workflows have not yet been formally validated and are not part of the supported interface.
The model weights are released under Apache-2.0.