Downloads · 30 days
0
stupid-zwl/rtlmul
rtlmul is a machine learning model from stupid-zwl. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This model predicts Power, Performance (delay), and Area (PPA) of RTL designs synthesized for Skywater 130nm technology. Given a Verilog module, it performs step-by-step chain-of-thought reasoning about gate-level syn…
Downloads · 30 days
0
Access
Public
Updated Apr 7, 2026
Repo size
33.2 GB
Likes
0
Public
Click a slice to open those files.
.safetensors24.8 GB · 69%
From the Hugging Face model README
This model predicts Power, Performance (delay), and Area (PPA) of RTL designs synthesized for Skywater 130nm technology. Given a Verilog module, it performs step-by-step chain-of-thought reasoning about gate-level synthesis and outputs structured PPA estimates with [area], [delay], and [static_power] tags directly usable as reinforcement learning reward signals.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "/path/to/merged/model"
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True)
SYSTEM_PROMPT = (
"Your task is to estimate area, delay, and static power for RTL designs in Skywater 130nm technology node.\n"
"For the given RTL design, reason about the number and type of gates that would be present after synthesis, "
"then output all four tags:\n"
"<synth> ... </synth>\n"
"<area> ... [area]value[/area] </area>\n"
"<delay> ... [delay]value[/delay] </delay>\n"
"<static_power> ... [static_power]value[/static_power] </static_power>"
)
rtl_code = "module top_module (...); ..."
messages = [{"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": rtl_code}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=16384, do_sample=False)
result = tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
# Extract: [area]VALUE[/area], [delay]VALUE[/delay], [static_power]VALUE[/static_power]
Fine-tuned on 23,816 samples (small circuits, <1000 gates) to establish fundamental CoT reasoning capability.
SFT Training Configuration:
| Parameter | Value |
|---|---|
| Base Model | Qwen3-8B |
| LoRA Rank | 256 |
| LoRA Alpha | 512 |
| LoRA Target | all |
| Learning Rate | 5e-5 |
| Batch Size | 64 (8 GPU × 1 × 8 accum) |
| Cutoff Length | 8192 |
| Epochs | 3 |
| Precision | BF16 + DeepSpeed ZeRO-3 |
Refined via Group Relative Policy Optimization (GRPO) using the verl framework. The reward signal is computed from MAPE on [area], [delay], and [static_power] tags extracted from model outputs, enabling iterative improvement on hard circuits.
RL Training Configuration:
| Parameter | Value |
|---|---|
| Algorithm | GRPO (Group Relative Policy Optimization) |
| Framework | verl |
| Train Batch Size | 256 |
| Max Prompt Length | 3072 |
| Max Response Length | 4096 |
| Actor Learning Rate | 1e-6 |
| PPO Mini Batch Size | 32 |
| PPO Micro Batch Size per GPU | 2 |
| Rollout Samples (n) | 4 |
| KL Loss Coefficient | 0.0 (disabled) |
| Entropy Coefficient | 0 |
| Gradient Checkpointing | Enabled |
| Precision | BF16 |
| Rollout Engine | vLLM |
| Total Epochs | 5 |
RL Launch Command:
HF_HUB_OFFLINE=1 CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 python -m verl.trainer.main_ppo \
algorithm.adv_estimator=grpo \
data.train_files=/path/to/train.parquet \
data.val_files=/path/to/val.parquet \
data.train_batch_size=256 \
data.max_prompt_length=3072 \
data.max_response_length=4096 \
actor_rollout_ref.model.path=/path/to/metrex_merged_full \
actor_rollout_ref.actor.optim.lr=1e-6 \
actor_rollout_ref.actor.ppo_mini_batch_size=32 \
actor_rollout_ref.actor.ppo_micro_batch_size_per_gpu=2 \
actor_rollout_ref.actor.use_kl_loss=False \
actor_rollout_ref.actor.kl_loss_coef=0.0 \
actor_rollout_ref.actor.entropy_coeff=0 \
actor_rollout_ref.rollout.n=4 \
actor_rollout_ref.rollout.temperature=1.0 \
actor_rollout_ref.rollout.top_p=1.0 \
actor_rollout_ref.rollout.val_kwargs.temperature=1.0 \
actor_rollout_ref.rollout.val_kwargs.top_p=0.7 \
actor_rollout_ref.rollout.val_kwargs.do_sample=True \
actor_rollout_ref.rollout.val_kwargs.n=1 \
reward_model.reward_manager=metrex \
algorithm.use_kl_in_reward=False \
trainer.critic_warmup=0 \
trainer.project_name=MetRex-RL \
trainer.experiment_name=GRPO-Qwen3-8B-MetRex \
trainer.n_gpus_per_node=8 \
trainer.nnodes=1 \
trainer.val_before_train=True \
trainer.test_freq=5 \
trainer.save_freq=10 \
trainer.total_epochs=5
<synth>, <area>, <delay>, <static_power>)
## References
- [MetRex: A Benchmark for RTL Code Generation with LLMs](https://github.com/scale-lab/MetRex/tree/main) — Chain-of-thought PPA prediction baseline
- [ChipGPT: How Far Are We From Natural Language Hardware Design](https://arxiv.org/abs/2305.14019) — LLM-assisted hardware design framework
- [Data is All You Need: Finetuning LLMs for Chip Design via Automated Design-Data Augmentation](https://arxiv.org/abs/2403.11202) — Automated RTL data augmentation framework
- [verl: Versatile RL Framework](https://github.com/verl-project/verl) — GRPO training framework