Downloads Β· 30 days
11
12% of all-time downloads
oflorez/Wayra-Perplexity-Estimator-55M-TensorRT
Wayra-Perplexity-Estimator-55M-TensorRT is a text classification model from oflorez. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
π A100-optimized TensorRT version of WayraPPL for high-throughput prediction of Perplexity.
Downloads Β· 30 days
11
12% of all-time downloads
All-time downloads
95
Public
Parameters
55.4M
351 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors221 MB Β· 63%
From the Hugging Face model README
π A100-optimized TensorRT version of WayraPPL for high-throughput prediction of Perplexity.
This model works on NVIDIA A100 GPUs with:
# Install requirements (A100 + CUDA 12.8+ required)
pip install -r tensorrt_requirements.txt
# Verify TensorRT installation
python -c "import tensorrt; print(tensorrt.__version__)" # Should be 10.13.x
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("oflorez/Wayra-Perplexity-Estimator-55M-TensorRT")
model = AutoModel.from_pretrained("oflorez/Wayra-Perplexity-Estimator-55M-TensorRT")
texts = ["Your text here"]
inputs = tokenizer(texts, return_tensors="pt", padding=True, truncation=True)
outputs = model(**inputs)
print(f"PPL: {outputs['ppl']}")
from tensorrt_inference import WayraPPLTensorRT
from transformers import AutoTokenizer
# Load TensorRT model (A100 required)
model = WayraPPLTensorRT("wayrappl_fp16_bs2048.engine")
tokenizer = AutoTokenizer.from_pretrained("oflorez/Wayra-Perplexity-Estimator-55M-TensorRT")
# High-throughput inference
texts = ["Your text here"] * 1000 # Large batch
inputs = tokenizer(texts, return_tensors="pt", padding=True, truncation=True)
outputs = model.infer(inputs['input_ids'].numpy(), inputs['attention_mask'].numpy())
PyTorch Model: Standard HuggingFace format
pytorch_model.bin - Model weightsconfig.json - Model configurationtokenizer.json - TokenizerTensorRT Engine: A100-optimized
wayrappl_fp16_bs2048.engine - TensorRT engine (A100 only)tensorrt_config.json - Engine configurationtensorrt_inference.py - Inference codetensorrt_requirements.txt - Dependencies| Model Type | Throughput | Latency | Memory |
|---|---|---|---|
| Llama 3 1B | ~200/sec | 50ms | 8GB |
| Wayra PyTorch | ~1,000/sec | 10ms | 4GB |
| Wayra TensorRT | ~50,000/sec | <1ms | 2GB |
"TensorRT engine not compatible"
nvidia-smi (should be 12.8+)python -c "import tensorrt" (should be 10.13.x)"CUDA out of memory"
@software{WayraPPL,
title={WayraPPL: High-Performance Perplexity Estimation of Data Novelty},
author={Omar U. Florez and LatamGPT Team},
year={2025},
url={https://huggingface.co/latam-gpt/Wayra-Perplexity-Estimator-55M}
}
Apache 2.0 - See LICENSE file
Note: This model is optimized for A100 GPUs. For other GPUs, use the PyTorch version or retrain the TensorRT engine for your specific hardware.