Downloads · 30 days
48
19% of all-time downloads
TracNetwork/functiongemma-270m-it-intercomswap-v3
functiongemma-270m-it-intercomswap-v3 is a text generation model from TracNetwork. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as gemma.
IntercomSwap fine-tuned FunctionGemma model for deterministic tool-calling in BTC Lightning <- USDT Solana swap workflows.
Downloads · 30 days
48
19% of all-time downloads
All-time downloads
259
Public
Parameters
268M
1.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors928 MB · 50%
From the Hugging Face model README
IntercomSwap fine-tuned FunctionGemma model for deterministic tool-calling in BTC Lightning <-> USDT Solana swap workflows.
Intercom Swap is a fork of upstream Intercom that keeps the Intercom stack intact and adds a non-custodial swap harness for BTC over Lightning <> USDT on Solana via a shared escrow program, with deterministic operator tooling, recovery, and unattended end-to-end tests.
GitHub: https://github.com/TracSystems/intercom-swap
Base model: google/functiongemma-270m-it
./:
./nvfp4:
./gguf:
functiongemma-v3-f16.gguffunctiongemma-v3-q8_0.ggufpython -m vllm.entrypoints.openai.api_server \
--model TracNetwork/functiongemma-270m-it-intercomswap-v3 \
--host 0.0.0.0 \
--port 8000 \
--dtype auto \
--max-model-len 8192
Lower memory mode example:
python -m vllm.entrypoints.openai.api_server \
--model TracNetwork/functiongemma-270m-it-intercomswap-v3 \
--host 0.0.0.0 \
--port 8000 \
--dtype auto \
--max-model-len 4096 \
--max-num-seqs 8
./nvfp4)TensorRT-LLM example with explicit headroom (avoid consuming all VRAM):
trtllm-serve serve ./nvfp4 \
--backend pytorch \
--host 0.0.0.0 \
--port 8012 \
--max_batch_size 8 \
--max_num_tokens 16384 \
--kv_cache_free_gpu_memory_fraction 0.05
Memory tuning guidance:
--max_num_tokens first.--max_batch_size.--kv_cache_free_gpu_memory_fraction around 0.05 to preserve safety headroom../gguf)Q8_0 (recommended default balance):
llama-server \
-m ./gguf/functiongemma-v3-q8_0.gguf \
--host 0.0.0.0 \
--port 8014 \
--ctx-size 8192 \
--batch-size 256 \
--ubatch-size 64 \
--gpu-layers 12
F16 (higher quality, higher memory):
llama-server \
-m ./gguf/functiongemma-v3-f16.gguf \
--host 0.0.0.0 \
--port 8014 \
--ctx-size 8192 \
--batch-size 256 \
--ubatch-size 64 \
--gpu-layers 12
Memory tuning guidance:
--gpu-layers to reduce VRAM usage.--ctx-size to reduce RAM+VRAM KV-cache usage.q8_0 for general deployment, f16 for quality-first offline tests.From held-out evaluation for this release line:
62637550.013480.02012google/functiongemma-270m-it