Downloads · 30 days
15
35% of all-time downloads
Irfanuruchi/Qwen3-4B-Computer-Science-OpenVINO-INT4
Qwen3-4B-Computer-Science-OpenVINO-INT4 is a text generation model from Irfanuruchi. Use it when you need the model to write or continue text. It is set up for optimum. The card lists the license as apache-2.0.
This repository provides an OpenVINO INT4 version of Qwen3-4B-Computer-Science, optimized for efficient inference on Intel CPUs and other OpenVINO-supported hardware.
Downloads · 30 days
15
35% of all-time downloads
All-time downloads
43
Public
Repo size
2.3 GB
Likes
0
Public
Click a slice to open those files.
.bin2.3 GB · 99%
From the Hugging Face model README
This repository provides an OpenVINO INT4 version of Qwen3-4B-Computer-Science, optimized for efficient inference on Intel CPUs and other OpenVINO-supported hardware.
The model has been exported using Optimum Intel with OpenVINO IR format and INT4 weight compression, providing a significantly smaller footprint while maintaining strong performance for software engineering and computer science workloads.
| Property | Value |
|---|---|
| Base Model | Irfanuruchi/Qwen3-4B-Computer-Science |
| Architecture | Qwen3 |
| Parameters | ~4 Billion |
| Format | OpenVINO IR |
| Weight Compression | INT4 Asymmetric |
| Group Size | 128 |
| Framework | OpenVINO + Optimum Intel |
| Primary Device | CPU |
| License | Apache License 2.0 |
The model was exported using:
optimum-cli export openvino \
--model Irfanuruchi/Qwen3-4B-Computer-Science \
--task text-generation-with-past \
--weight-format int4 \
Qwen3-4B-Computer-Science-OpenVINO-INT4
Compression statistics:
pip install -U openvino optimum-intel transformers
from transformers import AutoTokenizer
from optimum.intel.openvino import OVModelForCausalLM
model_id = "Irfanuruchi/Qwen3-4B-Computer-Science-OpenVINO-INT4"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = OVModelForCausalLM.from_pretrained(
model_id,
device="CPU",
)
messages = [
{
"role": "system",
"content": "You are a computer science assistant."
},
{
"role": "user",
"content": "Explain Floyd's cycle detection algorithm."
},
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(
**inputs,
max_new_tokens=256,
do_sample=False,
)
response = tokenizer.decode(
outputs[0][inputs["input_ids"].shape[1]:],
skip_special_tokens=True,
)
print(response)
The exported model has been successfully validated using:
Validation confirmed successful generation of technically correct programming responses, including algorithm implementation and complexity analysis.
This model is intended for:
As with other large language models, outputs should be reviewed before production use. The model may occasionally:
INT4 compression may also introduce minor differences compared to higher-precision variants.
This model is distributed under the Apache License 2.0.
Please refer to the included LICENSE file for the complete license text and attribution requirements.