Downloads · 30 days
356
100% of all-time downloads
nanash66/typhoon-ocr1.5-2b-ROCMFP4-GGUF
typhoon-ocr1.5-2b-ROCMFP4-GGUF is a image-text-to-text model from nanash66. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This repository provides ROCmFP4 quantized GGUF weights for typhoon-ai/typhoon-ocr1.5-2b, calibrated specifically for high-fidelity Thai document OCR using Importance Matrix (iMatrix).
Downloads · 30 days
356
100% of all-time downloads
All-time downloads
356
Public
Repo size
2 GB
Likes
0
Public
Click a slice to open those files.
.gguf2 GB · 100%
From the Hugging Face model README
This repository provides ROCmFP4 quantized GGUF weights for typhoon-ai/typhoon-ocr1.5-2b, calibrated specifically for high-fidelity Thai document OCR using Importance Matrix (iMatrix).
The text backbone is quantized to Q4_0_ROCMFP4 (UE4M3-scale experimental with Q6_K token embeddings), while the Vision Projector (mmproj-f16.gguf) remains unquantized in full FP16 to ensure zero loss in visual document resolution.
[!TIP]
🚀 UPDATE (Sep 2026): 100% Native ROCm / HIP GPU Acceleration is LIVE!
Full GPU acceleration (Vision Transformer + LLM Backbone) is now supported natively via PR #112 (charlie12345/ROCmFPX#112)!
- No More CPU Fallback: You no longer need
--no-mmproj-offload! Both the vision encoder and LLM run 100% on the AMD GPU.- Blazing Speeds on AMD Radeon 890M (Strix Point):
- ⚡ Vision Prompt Processing:
327.8 tok/s(entire high-res document ingested in ~6–7 seconds)- 🚀 Text Generation (OCR Decode):
46.7 tok/s- 💾 Ultra-Low VRAM: Total memory footprint is only ~2.7 GB (allowing large batches or multi-turn OCR on 16GB–32GB unified memory laptops).
- Build Requirement: Until PR #112 is merged into upstream master releases, you must build the engine from source using PR #112 or the branch
nanashi66:fix-cublas-f16-pointer-mismatch. See the Build Instructions below.- Alternative (Plug-and-Play): If you prefer not to build from source, you can still use the pre-built Vulkan backend of
ROCmFPXwhich runs out of the box.
<table>...</table>), zero character hallucinations, and 100% identical Thai numerals.gfx1150 / gfx1100 / gfx1103): High inference speed (46.7 tok/s) and ultra-low VRAM footprint (~2.7 GB) on AMD Ryzen AI 9 HX 370 / Radeon 890M / 880M / 780M iGPUs.Hardware: AMD Ryzen AI 9 HX 370 / AMD Radeon 890M (16 CUs, RDNA 3.5, 32GB LPDDR5X UMA).
Test Asset: Multi-column Thai Police Rank Standard Classification Document (test_page2.png).
| Backend / Engine | Vision Projector (mmproj) | Vision Ingestion (Prompt) | Generation Speed (Decode) | Avg Full Page Time | Thai Character Fidelity |
|---|---|---|---|---|---|
| 🚀 Native ROCm / HIP 7.2 (With PR #112) | 🛡️ GPU Accelerated | 🏆 327.8 tok/s | 🏆 46.7 tok/s | ⏱️ ~12.8 s | 🎯 100% Perfect Match |
| 🔴 Vulkan (SPIR-V Shaders) | 🛡️ GPU Accelerated | ⚡ ~110–130 tok/s | 🚀 28.6 tok/s | ⏱️ ~29.1 s | 🎯 99.61% Match |
⏳ Legacy ROCm / HIP (--no-mmproj-offload) | 🐢 CPU (AVX-512) | ~40–50 tok/s | 🚀 45 tok/s | ⏱️ ~50.6 s | 🎯 99.61% Match |
| ❌ Unpatched ROCm / HIP (Without PR #112) | GPU (Buggy pointer) | Corrupted | N/A | Invalid | ⚠️ Severe Hallucinations |
In earlier releases of CUDA/HIP backends in llama.cpp and ROCmFPX, running multimodal vision models (such as Qwen2-VL, SigLIP, and Typhoon-OCR) on GPU caused output corruption and hallucinations.
In ggml-cuda.cu, intermediate activation buffer src1_ddf_i is allocated as 32-bit float (float *). In ggml_cuda_op_mul_mat_cublas, when multiplying matrices where src1->type == GGML_TYPE_F16, the code mistakenly bypassed the to_fp16_cuda conversion kernel and directly cast the pointer:
// Bug in upstream: Reading 32-bit floats as pairs of 16-bit halfs!
const half * src1_as_half = (const half *) src1_ddf_i;
This caused the GPU GEMM kernel to interpret 32-bit IEEE floats as pairs of 16-bit half floats, corrupting the first Vision Transformer embedding layer by over 40x magnitude.
PR #112 enforces explicit to_fp16_cuda type conversion into an allocated half-precision scratch buffer before invoking hipblasGemmEx / cublasGemmEx. This restores 100% mathematical accuracy on the GPU for all vision and multimodal models.
| File | Size | Description |
|---|---|---|
typhoon-ocr1.5-2b-Q4_0_ROCMFP4-imatrix.gguf | 1.08 GB | Quantized text model backbone (Q4_0_ROCMFP4 with Q6_K token embeddings) |
typhoon-ocr1.5-2b.mmproj-f16.gguf | 781 MB | Full FP16 multimodal vision projector (Required for image/PDF input) |
To get full GPU acceleration on AMD Radeon iGPUs/dGPUs under Windows 11 or Linux:
git clone https://github.com/charlie12345/ROCmFPX.git
cd ROCmFPX
# Fetch and checkout PR #112
git fetch origin pull/112/head:pr-112
git checkout pr-112
# Alternatively, clone the author's fork directly:
# git clone -b fix-cublas-f16-pointer-mismatch https://github.com/nanashi66/ROCmFPX.git
Windows (ROCm 7.2 / Visual Studio Build Tools + Ninja):
$env:PATH = "C:\Program Files\AMD\ROCm\7.2\bin;" + $env:PATH
cmake -B build-hip -G Ninja `
-DGGML_HIP=ON `
-DAMDGPU_TARGETS="gfx1150;gfx1100" `
-DCMAKE_BUILD_TYPE=Release `
-DLLAMA_BUILD_WEBUI=OFF
ninja -C build-hip bin/llama-server.exe bin/llama-cli.exe
Linux (ROCm 6.x / 7.x):
cmake -B build-hip -G Ninja \
-DGGML_HIP=ON \
-DAMDGPU_TARGETS="gfx1150;gfx1100" \
-DCMAKE_BUILD_TYPE=Release \
-DLLAMA_BUILD_WEBUI=OFF
ninja -C build-hip bin/llama-server bin/llama-cli
llama-server (Production Web API)Set the required environment variables and launch the server:
# Windows PowerShell
$env:HSA_OVERRIDE_GFX_VERSION = "11.5.0" # Required for Strix Point (Radeon 890M / gfx1150)
$env:GGML_CUDA_FORCE_MMQ = "1"
llama-server.exe `
-m typhoon-ocr1.5-2b-Q4_0_ROCMFP4-imatrix.gguf `
--mmproj typhoon-ocr1.5-2b.mmproj-f16.gguf `
-ngl 99 `
-c 8192 `
-ctk q8_0 -ctv q8_0 `
-fa on `
--port 8083 `
--host 0.0.0.0
[!NOTE]
-ctk q8_0 -ctv q8_0: Uses 8-bit quantized KV cache to halve VRAM usage with zero degradation.-fa on: Enables Flash Attention to compute long document context in linear memory.- Notice: Notice that
--no-mmproj-offloadis no longer needed!
llama-cli)To test directly from the command line on an image:
llama-cli.exe `
-m typhoon-ocr1.5-2b-Q4_0_ROCMFP4-imatrix.gguf `
--mmproj typhoon-ocr1.5-2b.mmproj-f16.gguf `
--image sample_document.png `
-ngl 99 `
-c 8192 `
-ctk q8_0 -ctv q8_0 `
-fa on `
-n 1024 `
-p "Extract all text from the image.\n\nInstructions:\n- Only return the clean Markdown.\n- Do not include any explanation or extra text.\n- You must include all information on the page.\n\nFormatting Rules:\n- Tables: Render tables using <table>...</table> in clean HTML format.\n- Equations: Render equations using LaTeX syntax with inline ($...$) and block ($$...$$)."
import base64
import requests
def ocr_document(image_path: str, server_url: str = "http://localhost:8083/v1/chat/completions"):
with open(image_path, "rb") as f:
img_b64 = base64.b64encode(f.read()).decode("utf-8")
prompt = """Extract all text from the image.
Instructions:
- Only return the clean Markdown.
- Do not include any explanation or extra text.
- You must include all information on the page.
Formatting Rules:
- Tables: Render tables using <table>...</table> in clean HTML format.
- Equations: Render equations using LaTeX syntax with inline ($...$) and block ($$...$$).
- Page Numbers: Wrap page numbers in <page_number>...</page_number>."""
payload = {
"model": "typhoon-ocr1.5-2b",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{img_b64}"}}
]
}
],
"temperature": 0.0,
"max_tokens": 4096
}
response = requests.post(server_url, json=payload, timeout=120)
return response.json()["choices"][0]["message"]["content"]
# Example:
# print(ocr_document("thai_legal_doc.png"))
typhoon-ai/typhoon-ocr1.5-2b.llama.cpp open-source community.If you find this ROCmFP4 quantization, empirical benchmark data, and Native ROCm/HIP GPU bugfixes helpful for your local AI workflows on AMD hardware, consider supporting our ongoing research and optimization efforts:
bc1pm9yqcj5m5jv3jzxkhv4uvrgl2weyq46m2jlzufzyfwyr4fmv8gwq548kw6Your support helps us continue optimizing open-weight models, submitting upstream ROCm/HIP patches, and developing high-efficiency inference recipes for AMD Strix Point and RDNA architectures!
Apache 2.0