Downloads · 30 days
0
syaher/custom-ammar-agent
custom-ammar-agent is a machine learning model from syaher. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This dataset is designed for fine-tuning language models (e.g., Qwen2.5-Coder-7B-Instruct) to create an AI agent that understands the Ammar Jewels project (a Nuxt 4 + Drizzle/Turso monorepo) and can assist with coding…
Downloads · 30 days
0
Access
Public
Updated Aug 14, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.jsonl1.4 MB · 100%
From the Hugging Face model README
This dataset is designed for fine-tuning language models (e.g., Qwen2.5-Coder-7B-Instruct) to create an AI agent that understands the Ammar Jewels project (a Nuxt 4 + Drizzle/Turso monorepo) and can assist with coding, architecture, and domain-specific questions.
The dataset is a mixture of three sources, formatted as ShareGPT-style JSONL (each line is a JSON object with a "conversations" array containing user and assistant turns).
Chat (chat.jsonl)
Extracted from real conversation logs (7 sources: 4 .md exports and 3 .json attachments).
Contains technical discussions, architecture decisions, and debugging sessions in Bahasa Malaysia and English.
Cleaned to remove system noise, tool calls, and non-essential content.
Docs (docs.jsonl)
Auto-generated question-answer pairs from project documentation:
MASTERPLAN.mdAGENTS.mddocs/AUDIT-CORE-VERTIKAL.mddocs/BIO-SITE-BUSINESS-PLAN.mddocs/FRAPPE-BLUEPRINT.mddocs/GUARDRAILS.mddocs/PERLEMBAGAAN.mdQuestions are auto-templated (e.g., "Terangkan tentang: ...", "Apakah Fasa X?") and answers are the original text chunks.
Code (code.jsonl)
Extracted from write_file and patch tool calls in the conversation logs.
Each example consists of a natural language request (inferred from context or generated) and the actual code that was written or patched.
Covers: API endpoints, composables, utilities, components, schema modifications, and more.
| Split | # Examples | Approx. Size |
|---|---|---|
| Train | 721 | 1.3 MB |
| Validation | 38 | 47 KB |
| Total | 759 | ~1.4 MB |
Note: After deduplication, 759 unique examples remain from a gross of 762.
{"conversations":[{"role":"user","content":"Buat endpoint POST untuk stok/pembelian?"},{"role":"assistant","content":"import { eq } from 'drizzle-orm'\\nimport { stocks, purchases } from '@ammar/database'\\n\\ndefineEventHandler(async (event) => {\n const body = await readBody(event)\n // ... validation and DB insert\n})"},"source":"ui1","tool":"write_file","path":"apps/portal/server/api/stock/purchases.post.ts","lines":42}
The dataset is ready for supervised fine-tuning (SFT) or LoRA/QLoRA training.
Example with �� 🤗 Transformers and �� 🤗 PEER:
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForCausalLM, TrainingArguments, Trainer
dataset = load_dataset("syaher/custom-ammar-agent")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct")
# ... tokenize, set up Trainer, etc.
nuxt-template repository (Ammar Jewels monorepo) as of August 2026.scripts/dataset/ of the source repository.This dataset is released under the CC-BY-SA 4.0 license. Please attribute appropriately if used.
Generated automatically via dataset creation pipeline. For questions or issues, contact the dataset curator.