Downloads · 30 days
0
officialcyber88/torch-tensorcore-speedkit
torch-tensorcore-speedkit is a machine learning model from officialcyber88. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Enables TF32/BF16 Tensor Core fast paths in PyTorch via safe auto-detection, with auditable, reversible flag application and reproducible benchmarks.
Downloads · 30 days
0
Access
Public
Updated Jan 4, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.py47.7 KB · 38%
From the Hugging Face model README
Enables TF32/BF16 Tensor Core fast paths in PyTorch via safe auto-detection, with auditable, reversible flag application and reproducible benchmarks.
A reproducible performance protocol packaged as code.
It focuses on real levers PyTorch exposes:
Note: On NVIDIA, Tensor Cores are accessed through CUDA libraries under the hood; the point here is: you don't write CUDA.
For CUDA 11.8 (recommended for RTX 30/40 series):
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu118
For other CUDA versions, see PyTorch Get Started.
git clone https://github.com/yourusername/torch-tensorcore-speedkit.git
cd torch-tensorcore-speedkit
pip install -e .
# or with vision examples:
pip install -e ".[vision]"
Get instant feedback on your hardware:
python -m torch_speedkit.report --apply --restore
Run the verified benchmark:
python examples/benchmark_compare.py
python -m torch_speedkit.bench --config examples/configs/speed.yaml
python examples/train_toy_transformer.py --config examples/configs/speed.yaml
# or CIFAR10 ResNet if you have torchvision:
python examples/train_cifar10_resnet.py --config examples/configs/speed.yaml
See examples/configs/speed.yaml and src/torch_speedkit/config.py.
Key settings:
shape_paddingLinear(8192, 8192) × 3 layers, batch size 512| Configuration | ms/step | Speedup |
|---|---|---|
| Baseline (FP32, TF32 off) | 184.68 | 1.00× |
| SpeedKit (BF16, TF32 on) | 131.08 | 1.41× |
State changes verified:
torch.get_float32_matmul_precision(): highest → hightorch.backends.cuda.matmul.allow_tf32: False → TrueSpeedups are largest when:
Always benchmark on your specific model + batch size. Smaller workloads may show less gain.