Downloads · 30 days
22
20% of all-time downloads
manueldeprada/sampling_with_kvcache_hf_helpers
sampling_with_kvcache_hf_helpers is a text generation model from manueldeprada. Use it when you need the model to write or continue text. It is set up for transformers.
A clean, hackable implementation of sampling (also called ancestral sampling or multinomial sampling) with full KV cache support. This is a simplified alternative to the complex generation mixin in transformers, desig…
Downloads · 30 days
22
20% of all-time downloads
All-time downloads
108
Public
Parameters
135M
269 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors269 MB · 99%
From the Hugging Face model README
A clean, hackable implementation of sampling (also called ancestral sampling or multinomial sampling) with full KV cache support. This is a simplified alternative to the complex generation mixin in transformers, designed for readability and ease of modification while maintaining full performance.
The implementation supports both sampling and greedy decoding modes, with optional temperature scaling and top-k/top-p filtering.
Most transformer LLM/VLM models trained for causal language modeling.
temperature (float): Sampling temperature (default: 1.0, higher = more random)top_k (int): Only consider top-k most probable tokens (default: None)top_p (float): Only consider tokens with cumulative probability <= top_p (default: None)do_sample (bool): Whether to use sampling (True, default) or greedy decoding (False)Logits processors are applied in sequence: temperature → softmax → top_k → top_p (same as HuggingFace's LogitProcessor system). Temperature scaling occurs before top-p filtering, affecting the probability distribution that top-p operates on.
For example, with temperature=1.0, top_p=0.9 might include tokens A, B, C. With temperature=0.5, probability mass is much more concentrated, so top_p=0.9 might only include token A.
When return_dict_in_generate=True, returns a dictionary with:
sequences: Generated token IDsscores: Log probabilities of sampled tokens (with temperature/sampling modifications)logprobs: Original model log probabilities (T=1, no modifications)
Otherwise, returns a tensor of generated token IDs.from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct", device_map="auto")
inputs = tokenizer(["The quick brown"], return_tensors="pt").to(model.device)
# Basic sampling
gen_out = model.generate(**inputs, custom_generate="manueldeprada/sampling_with_kvcache_hf_helpers", trust_remote_code=True)
# With temperature
gen_out = model.generate(**inputs, custom_generate="manueldeprada/sampling_with_kvcache_hf_helpers", temperature=0.8, trust_remote_code=True)
# With top-k
gen_out = model.generate(**inputs, custom_generate="manueldeprada/sampling_with_kvcache_hf_helpers", top_k=50, trust_remote_code=True)
# With top-p (nucleus sampling)
gen_out = model.generate(**inputs, custom_generate="manueldeprada/sampling_with_kvcache_hf_helpers", top_p=0.9, trust_remote_code=True)
# Greedy decoding (no sampling)
gen_out = model.generate(**inputs, custom_generate="manueldeprada/sampling_with_kvcache_hf_helpers", do_sample=False, trust_remote_code=True)
# Get detailed output with probabilities
gen_out = model.generate(
**inputs,
custom_generate="manueldeprada/sampling_with_kvcache_hf_helpers",
return_dict_in_generate=True,
trust_remote_code=True
)
print(f"Generated text: {tokenizer.batch_decode(gen_out['sequences'], skip_special_tokens=True)}")
print(f"Sampling scores: {gen_out['scores']}")
print(f"Model log probabilities: {gen_out['logprobs']}")