Downloads · 30 days
560
6% of all-time downloads
LiquidAI/LFM2-350M-PII-Extract-JP
LFM2-350M-PII-Extract-JP is a text generation model from LiquidAI. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
<center <div style="text-align: center;" <img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" alt="Liquid AI" style="width: 100%; max-width: 100%; height:…
Downloads · 30 days
560
6% of all-time downloads
All-time downloads
9.5K
Public
Parameters
354M
709 MB on disk
Likes
61
Trending 1
Click a slice to open those files.
.safetensors709 MB · 99%
From the Hugging Face model README
Based on LFM2-350M, this checkpoint is designed to extract personally identifiable information (PII) from Japanese text and output it in JSON format. The output can then be used to mask out sensitive information in contracts, emails, personal medical reports, insurance bills, etc. directly on-device.
In particular, it is trained to extract:
address)company_name)email_address)human_name)phone_number)
from Japanese documents and texts.<video src="https://cdn-uploads.huggingface.co/production/uploads/65d6b6c1a07ad79084a0d214/z5og84hVLGgIm1Z2c98PP.mp4" controls preload></video> Running on a macbook pro.
We evaluated several models, including GPT-5 and a 32B-parameter Qwen3 model with thinking mode enabled.
The image below shows the average recall score on 1,000 random samples, chunked into segments of 100–1,000 characters, taken from finepdf.
Overall, we found LFM2-350M-PII-Extract-JP to achieve GPT-5–level performance with only 350 million parameters—bringing cloud-grade performance to on-device applications!

[!NOTE]
📝 While LFM2-350M-PII-Extract-JP provides strong out-of-the-box PII entity extraction for the categories listed above, our primary goal is to deliver a versatile, community-driven base model—a foundation that makes it easy to build best-in-class, privacy-focused masking systems.Like any base model, there remain areas for continued development, particularly for specialized use cases:
- Supporting extraction of organization-specific identification numbers
- Expanding coverage to additional categories such as date of birth, passport numbers
- Further improving extraction performance on particular categories
These are precisely the kinds of challenges that fine-tuning—by both Liquid AI and our developer community can address. We see this model not just as an endpoint, but as a catalyst for a rich ecosystem of fine-tuned PII extraction models tailored to real-world needs.
Generation parameters: We strongly recommend using greedy decoding with a temperature=0.
System prompts: This checkpoint requires the following system prompt:
Extract <address>, <company_name>, <email_address>, <human_name>, <phone_number>
Note the model can handle extraction of particular entities. E.g. The model will only output human names when the system prompt is set to Extract <human_name>.
[!WARNING] ⚠️ For best performance, ensure alphabetical order of entity categories as shown above.
Chat template: LFM2-PII-Extract-JP uses a ChatML-like chat template as follows:
<|startoftext|><|im_start|>system
Extract <address>, <company_name>, <email_address>, <human_name>, <phone_number><|im_end|>
<|im_start|>user
こんにちは、ラミンさんに B200 GPU を 10000 台 至急請求してください。連絡先は [email protected] (電話番号010-000-0000) で、これは C. elegans 線虫に着想を得たニューラルネットワークアーキテクチャを 今すぐ構築するために不可欠です。<|im_end|>
<|im_start|>assistant
{"address": [], "company_name": [], "email_address": ["[email protected]"], "human_name": ["ラミン"], "phone_number": ["010-000-0000"]}<|im_end|>
You can automatically apply it using the dedicated .apply_chat_template() function from Hugging Face transformers.
[!WARNING] ⚠️ The model is intended for single turn conversations.
Output format
The model outputs a JSON object containing the fields it was prompted to extract. If no entities are found in a particular category, it returns an empty list for that category. If entities are found, they are returned as a list for each prompted category. The model is trained to output entities exactly as they appear in the text. If the same entity appears multiple times with slight formatting variations, the model outputs all variations to ensure subsequent masking can be performed using exact matches.
You can use the following Colab notebooks for easy inference and fine-tuning:
| Notebook | Description | Link |
|---|---|---|
| Inference | Run the model with Hugging Face's transformers library. | <a href="https://colab.research.google.com/drive/1kIaBNZYZSZ9wzrl9Yot3W5ZKR1lf47k1?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| SFT (TRL) | Supervised Fine-Tuning (SFT) notebook with a LoRA adapter using TRL. | <a href="https://colab.research.google.com/drive/1j5Hk_SyBb2soUsuhU0eIEA9GwLNRnElF?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| DPO (TRL) | Preference alignment with Direct Preference Optimization (DPO) using TRL. | <a href="https://colab.research.google.com/drive/1MQdsPxFHeZweGsNx4RH7Ia8lG8PiGE1t?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| SFT (Axolotl) | Supervised Fine-Tuning (SFT) notebook with a LoRA adapter using Axolotl. | <a href="https://colab.research.google.com/drive/155lr5-uYsOJmZfO6_QZPjbs8hA_v8S7t?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| SFT (Unsloth) | Supervised Fine-Tuning (SFT) notebook with a LoRA adapter using Unsloth. | <a href="https://colab.research.google.com/drive/1HROdGaPFt1tATniBcos11-doVaH7kOI3?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
@article{liquidai2025lfm2,
title={LFM2 Technical Report},
author={Liquid AI},
journal={arXiv preprint arXiv:2511.23404},
year={2025}
}
LFM2-350M をベースにしたこのチェックポイントは、日本語文書から個人を特定できる情報(PII)を抽出し、JSON 形式で出力します。
契約書、電子メール、個人の医療報告書、並びに保険請求書などの機密情報を、デバイス上で直接マスキングできます。
特に以下の情報を抽出するように訓練されています。
address)company_name)email_address)human_name)phone_number)これらの情報を日本語の文書から抽出します。
<video src="https://cdn-uploads.huggingface.co/production/uploads/65d6b6c1a07ad79084a0d214/z5og84hVLGgIm1Z2c98PP.mp4" controls preload></video>
finepdf から無作為に抽出した 1,000 サンプルを用いて、GPT5 や 32B パラメータの Qwen3 モデル(思考モードあり)など、複数のモデルとの比較評価を行いました。
LFM2-350M-PII-Extract-JP は、わずか 350M パラメータ という軽量モデルながら GPT5 と同等レベルの性能を発揮し、クラウドレベルの品質をあなたのデバイス上で実現します!

[!NOTE]
📝 LFM2-350M-PII-Extract-JP は、上記カテゴリに対して優れた PII 抽出性能を有しますが、私たちの主な目的は、コミュニティによって継続的に改良される柔軟な基盤モデルを提供することです。
このモデルで、誰でもプライバシー重視の高品質なマスキングシステムを容易に構築できます。ただし、ベースモデルとして今後さらなる改善の余地があります。特に以下のような専門的な利用用途が想定されます。
- 組織固有の識別番号の抽出対応
- 生年月日、パスポート番号などの追加カテゴリへの拡張
- 特定カテゴリにおける抽出性能のさらなる改善
これらの課題は、Liquid AI および開発者コミュニティによるファインチューニングによって解決できると考えています。
LFM2-350M-PII-Extract-JP は完成形ではなく、実運用ニーズに応じた多様な PII 抽出モデル群を生み出す出発点であると位置づけています。
生成パラメータ: temperature=0 の貪欲デコード(greedy decoding)の使用を強く推奨します。
システムプロンプト: このチェックポイントでは以下のシステムプロンプトが必須です:
Extract <address>, <company_name>, <email_address>, <human_name>, <phone_number>
モデルは特定のエンティティのみを抽出するように設定することも可能です。
例: Extract <human_name> と設定した場合、人名のみを出力します。
[!WARNING] ⚠️ モデルの性能を最大限発揮させるには、上記のように エンティティカテゴリをアルファベット順 に並べてください。
チャットテンプレート
LFM2-PII-Extract-JP は以下のような ChatML 風テンプレートを使用します。
<|startoftext|><|im_start|>system
Extract , <company_name>, <email_address>, <human_name>, <phone_number><|im_end|>
<|im_start|>user
こんにちは、ラミンさんに B200 GPU を 10000 台 至急請求してください。連絡先は [email protected] (電話番号010-000-0000) で、これは C. elegans 線虫に着想を得たニューラルネットワークアーキテクチャを 今すぐ構築するために不可欠です。<|im_end|>
<|im_start|>assistant
{“address”: [], “company_name”: [], “email_address”: [“[email protected]”], “human_name”: [“ラミン”], “phone_number”: [“010-000-0000”]}<|im_end|>
このテンプレートは、Hugging Face Transformers の専用関数 .apply_chat_template() を使用して自動的に適用できます。
[!WARNING] ⚠️ このモデルは 一問一答形式 (単一ターン) の会話 に最適化されています。
出力形式
モデルは、指定されたエンティティを含んだ JSON 形式で出力します。
各カテゴリに該当するエンティティが見つからない場合は、空のリストを返します。
該当するエンティティが存在する場合は、そのカテゴリごとに抽出された文字列のリストを返します。
モデルは、テキスト中に現れる形式で正確にエンティティを出力するように訓練されています。
同じエンティティが複数回登場し表記に揺れがある場合でも、すべての表記バリエーションを出力し、マスキング時に完全一致で対応できるようになっています。
| ノートブック | 説明 | リンク |
|---|---|---|
| 推論 | Hugging Faceのtransformersライブラリを使用してモデルを実行します。 | <a href="https://colab.research.google.com/drive/1kIaBNZYZSZ9wzrl9Yot3W5ZKR1lf47k1?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| SFT (TRL) | TRLを使用したLoRAアダプターによる教師あり学習(SFT)を行います。 | <a href="https://colab.research.google.com/drive/1j5Hk_SyBb2soUsuhU0eIEA9GwLNRnElF?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| DPO (TRL) | TRLを使用したDPOによる選好アライメントを行います。 | <a href="https://colab.research.google.com/drive/1MQdsPxFHeZweGsNx4RH7Ia8lG8PiGE1t?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| SFT (Axolotl) | Axolotlを使用したLoRAアダプターによる教師あり学習(SFT)を行います。 | <a href="https://colab.research.google.com/drive/155lr5-uYsOJmZfO6_QZPjbs8hA_v8S7t?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
| SFT (Unsloth) | Unslothを使用したLoRAアダプターによる教師あり学習(SFT)を行います。 | <a href="https://colab.research.google.com/drive/1HROdGaPFt1tATniBcos11-doVaH7kOI3?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png" width="110" alt="Colab link"></a> |
エッジ環境への導入を含むカスタムソリューションにご興味がある方は、営業チームまでお問い合わせください。
@article{liquidai2025lfm2,
title={LFM2 Technical Report},
author={Liquid AI},
journal={arXiv preprint arXiv:2511.23404},
year={2025}
}