Downloads · 30 days
12
21% of all-time downloads
PIPer-icml/PIPer-8B-RL-only
PIPer-8B-RL-only is a text generation model from PIPer-icml. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Downloads · 30 days
12
21% of all-time downloads
All-time downloads
58
Public
Parameters
8.2B
32.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors32.8 GB · 100%
From the Hugging Face model README
Democratizing environment setup with on-device sized models that match the performance of much larger proprietary systems
</div>Environment setup—the process of configuring systems to work with specific software projects—remains a persistent challenge in software engineering. PIPer addresses this by training specialized on-device models that can automatically generate correct Bash scripts for environment configuration.
Our approach combines:
| Model | Size | EnvBench avg@5 | Cost per 1M tokens |
|---|---|---|---|
| PIPer | 8B | 19.4 | $0.60 |
| GPT-4o | - | 19.4 | $15.00 |
| Qwen3-32B | 32B | 16.2 | $2.00 |
| Qwen3-8B | 8B | 2.6 | $0.60 |
🎉 PIPer achieves 9× improvement over its base model while matching GPT-4o performance at 25x lower cost

| Model | Description | HuggingFace Link |
|---|---|---|
| 🏅 PIPer (Full) | Complete SFT+RL trained model | PIPer-icml/PIPer-8B |
| 🎯 PIPer (RL-only) | RLVR checkpoint only | PIPer-icml/PIPer-8B-RL-only |
| 📚 PIPer (SFT-only) | Supervised fine-tuning only | PIPer-icml/PIPer-8B-SFT-only |
| Dataset | Description | HuggingFace Link |
|---|---|---|
| EnvBench Zero-shot RL | Training prompts and evaluation data | PIPer-icml/envbench-zeroshot-rl |
| Benchmark | Description | Metric | Our Result |
|---|---|---|---|
| EnvBench-Python | 329 Python repositories | pass@5 | 🏆 27/329 |
| Repo2Run | 420 Python repositories | pass@5 | 🏆 103/420 |
| Terminal-Bench | 80 terminal tasks | pass@10 | 4/80 |
This project is licensed under the MIT License - see the LICENSE file for details.