Downloads · 30 days
38
9% of all-time downloads
prithivMLmods/Fara1.5-4B-FP8
Fara1.5-4B-FP8 is a image-text-to-text model from prithivMLmods. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
Fara1.5-4B-FP8 is a FP8 dynamic quantized version of microsoft/Fara1.5-4B, a multimodal computer use agent (CUA) for web browsers developed by Microsoft Research AI Frontiers. This quantization reduces model size and…
Downloads · 30 days
38
9% of all-time downloads
All-time downloads
405
Public
Parameters
4.5B
6.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors6.5 GB · 100%
How the weights are stored.
F8_E4M32.6B · 56%
From the Hugging Face model README
Fara1.5-4B-FP8 is a FP8 dynamic quantized version of microsoft/Fara1.5-4B, a multimodal computer use agent (CUA) for web browsers developed by Microsoft Research AI Frontiers. This quantization reduces model size and memory footprint while preserving the model's action-grounding and web-task-completion capabilities, making deployment more accessible on constrained hardware.
Fara1.5-4B is supervised fine-tuned from Qwen3.5-4B on synthetic web-task trajectories generated by FaraGen1.5. It observes the browser through screenshots and emits structured tool calls (click, type, scroll, visit URL, web search, and so on) to complete tasks end-to-end. It is vision-only at perception time, with coordinate-grounded actions and a critical-points safety design that pauses before personal info entry, payments, submissions, sign-ins, and other irreversible actions.
Quantization was performed using llmcompressor with the following recipe:
default_stage:
default_modifiers:
QuantizationModifier:
targets: [Linear]
ignore: ['re:.*lm_head', 're:.*embed_tokens$', 're:.*visual.*', 're:.*model.visual.*',
're:.*linear_attn.*']
scheme: FP8_DYNAMIC
bypass_divisibility_checks: false
Linear layers are quantized to FP8 with dynamic per-tensor activation scaling. The lm_head, embedding table, all vision-tower (visual) components, and linear_attn layers are excluded from quantization and remain at full precision to preserve grounding accuracy and multimodal fidelity.
| Base model | microsoft/Fara1.5-4B |
| Quantization scheme | FP8_DYNAMIC (Linear layers only) |
| Format | compressed-tensors |
| Original model size | ~9.11 GB |
| Compressed model size | ~6.54 GB |
| Size reduction | ~28% |
| Excluded from quantization | lm_head, embed_tokens, visual encoder, linear_attn |
| Developer (base model) | Microsoft Research AI Frontiers |
| Quantized by | prithivMLmods |
| Architecture | Multimodal decoder-only LM (image + text to text) |
| Parameters | 4B |
| Context length | 262,144 tokens |
| Inputs | User goal (text), current screenshot(s), prior agent thoughts and actions |
| Outputs | Chain-of-thought block followed by a tool-call block (XML-tagged) |
| License | MIT |
| Base model | Qwen3.5-4B (via Fara1.5-4B) |
Fara1.5-4B-FP8 is served through vLLM with native support for compressed-tensors FP8 checkpoints.
torch >= 2.11.0vllm >= 0.19.1vllm serve prithivMLmods/Fara1.5-4B-FP8 \
--max-model-len 262144 \
--limit-mm-per-prompt image=10
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
screenshot_b64 = capture_browser_screenshot() # base64 PNG
messages = [
{"role": "system", "content": FARA_SYSTEM_PROMPT},
{"role": "user", "content": [
{"type": "text", "text": "Book a table for 2 at a sushi place in Sunnyvale for Friday 7pm."},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{screenshot_b64}"}},
]},
]
response = client.chat.completions.create(
model="prithivMLmods/Fara1.5-4B-FP8",
messages=messages,
temperature=0.0,
max_tokens=2048,
)
print(response.choices[0].message.content)
The assistant reply contains chain-of-thought text followed by a <tool_call>{"name": "computer_use", "arguments": {...}}</tool_call> block. Execute the action in a sandboxed browser, capture the new screenshot, append the assistant turn and a tool/observation turn, and loop until the model emits action="terminate". Keep only the most recent 3 screenshots in the chat history. The screen resolution Fara is most commonly trained with is 1440x900; match this in your sandbox for the most reliable grounding.
Use the original Fara system prompt verbatim for best results, as documented in the base model card.
As with the base model, the safest way to run Fara1.5-4B-FP8 is inside a sandboxed environment such as MagenticLite, with Docker isolation, domain allow-lists, watch-mode monitoring, and an immediate pause control. Do not run this model with unrestricted browser access on a machine holding sensitive credentials or files.
Intended use, known limitations, safety considerations, and responsible AI guidance are unchanged from the base model. See the Fara1.5-4B model card for full details on critical-points safety design, evaluation results, training data, out-of-scope use, and responsible AI considerations. Quantization to FP8 may introduce minor numerical differences relative to the bf16 base model; downstream accuracy should be validated for production use cases.
Released under the MIT License, consistent with the base model.
Fara-1.5: Scalable Learning Environments for Computer Use Agents — Hugging Face Papers