Downloads · 30 days
0
0% of all-time downloads
TheStageAI/LFM2.5-350M
LFM2.5-350M is a text generation model from TheStageAI. Use it when you need the model to write or continue text. It is set up for coreml. The card lists the license as apache-2.0.
Downloads · 30 days
0
0% of all-time downloads
All-time downloads
28
Public
Repo size
2.5 GB
Likes
0
Public
Click a slice to open those files.
.png580 KB · 99%
From the Hugging Face model README

Original model: LiquidAI/LFM2.5-350M (Liquid AI)
On-device chat LLM (Liquid LFM2.5 350M) compressed and packaged by TheStage AI for Apple Silicon. Strong quality/speed trade-off for assistants that need more capacity than 230M.
| Parameters | 350M |
| Chat template | LFM2 (from bundle) |
| Runtime | TheStage Apple SDK (TheStageLLM) |
| HF engines | TheStageAI/LFM2.5-350M |
This repo hosts CoreML / MLX engine bundles for the TheStage Apple SDK. Point engines_path at the repo id; the SDK downloads, caches, and loads the engines automatically. Chat template, EOS / stop tokens, and KV-cache horizon all come from the bundle.
| Property | Value |
|---|---|
| Hardware | Apple Silicon Mac, or physical iPhone / iPad |
| macOS | 15.0+ |
| iOS | 18.0+ |
| Xcode | 16.0+ |
| Swift | 6.0+ |
| Flutter (optional) | 3.24+ |
Simulator is not supported — run on real Apple Silicon hardware.
Create a token at app.thestage.ai and pass it to the SDK:
try await TheStageAI.shared.initialize(apiToken: "th_…")
await TheStageFlutterSDK.initialize(api_token: 'th_…');
Token is checked online in initialize (once per app process when reachable). Inference runs fully on-device. Offline initialize fails — reconnect and call initialize again.
Docs: TheStage Apple SDK · LLM · Repo: TheStage Apple SDK
In Xcode: File → Add Package Dependencies…, paste https://github.com/TheStageAI/AppleSDK.git, and add the TheStageSDK product. Or in Package.swift:
.package(
url: "https://github.com/TheStageAI/AppleSDK.git",
exact: Version(1, 1, 0)
)
# pubspec.yaml
dependencies:
thestage_apple_sdk:
git:
url: https://github.com/TheStageAI/AppleSDK.git
path: plugin/thestage_apple_sdk
ref: 1.1.0
import TheStageSDK
try await TheStageAI.shared.initialize(apiToken: "th_…")
let llm = try await TheStageLLM(
engines_path: "TheStageAI/LFM2.5-350M"
)
let result = llm.infer(
prompt: "What is 2+2?",
system_prompt: "You are a helpful assistant.",
max_new_tokens: 64
)
print(result.text)
For sampling control, start from the bundle preset and override only what you need:
var config = llm.generation_defaults
config.max_new_tokens = 256
config.temperature = 0.7
config.top_p = 0.8
config.repetition_penalty = 1.1
let result = llm.infer(prompt: "List 5 facts about London.", config: config)
for await chunk in llm.infer_stream(
prompt: "Tell me a short story.",
max_new_tokens: 512
) {
if chunk.is_final {
print("\n--- \(chunk.tokens_per_second ?? 0) tok/s ---")
} else {
print(chunk.text, terminator: "")
}
}
import 'package:thestage_apple_sdk/thestage_apple_sdk.dart';
await TheStageFlutterSDK.initialize(api_token: 'th_…');
await TheStageFlutterSDK.start_model(
model_name: 'llm',
engines_path: 'TheStageAI/LFM2.5-350M',
);
final result = await TheStageFlutterSDK.infer(
model_name: 'llm',
input_json: {
'prompt': 'What is 2+2?',
'system_prompt': 'You are a helpful assistant.',
'max_new_tokens': 64,
},
);
print(result[0]['text']);
Streaming:
final stream = TheStageFlutterSDK.infer_stream(
model_name: 'llm',
input_json: {'prompt': 'Tell me a short story.', 'max_new_tokens': 512},
);
await for (final chunk in stream) {
if (chunk['is_final'] == true) break;
final delta = chunk['delta'] as String?;
if (delta != null) stdout.write(delta);
}
| Direction | Field | Notes |
|---|---|---|
| in | prompt | User message |
| in | system_prompt | Optional; bundle default if omitted |
| in | max_new_tokens | Cap on generated tokens |
| in | temperature / top_k / top_p / min_p / repetition_penalty / seed | Sampling; omit → bundle generation_defaults |
| out | text | Decoded reply |
| out | tokens_per_second, time_to_first_token, stop_reason | Metrics |
Release build on Apple M2 Max (NPU / ANE), macOS 26.2 (guidance, not an SLA). Warmed decode, 256 new tokens:
| Metric | Value |
|---|---|
| tok/s (medium prompt) | 178 |
| TTFT small → long (s) | 0.038 → 0.060 |
| Process mem (MB) | ~61 |
Always measure release builds on device.
This work builds on LFM2.5-350M by Liquid AI: LiquidAI/LFM2.5-350M.
This Apple Silicon package is produced by TheStage AI; it is not an official Liquid AI release.