Downloads · 30 days
0
ParasiticRogue/Model-Tips-and-Tricks
Model-Tips-and-Tricks is a machine learning model from ParasiticRogue. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Greetings! This page intends to share some insights I've gathered from several years of experience testing Large Language Models (LLMs). The focus is primarily on optimizing LLMs for creative writing and roleplaying a…
Downloads · 30 days
0
Access
Public
Updated Apr 25, 2025
Repo size
—
Likes
13
Public
Click a slice to open those files.
.md36.4 KB · 96%
From the Hugging Face model README
Greetings! This page intends to share some insights I've gathered from several years of experience testing Large Language Models (LLMs). The focus is primarily on optimizing LLMs for creative writing and roleplaying applications, though these tips may also be beneficial for technical tasks as well.
These suggestions may not apply universally to every model or use case, nor do they guarantee optimal results in every instance. However, they should help improve your results and overall experience with the model, even if marginally, and assist you in navigating common pitfalls or misleading advice. Some points may be familiar to experienced users, but I aim to include less commonly discussed details too. Please note that I am still learning and open to feedback and corrections.
The Instruct Format (also known as an Instruct Template) is arguably one of the most crucial aspects for ensuring an LLM functions correctly. It defines the specific structure and special tokens used to separate different parts of the input – such as system prompts, user messages, and AI responses – during the model's training and subsequent interactions.
Some formats, like ChatML or Alpaca, are widely adopted across various models. Others are specific to certain model families, such as Llama 3 Instruct or Mistral Instruct. However, it's important to note that not all models released under a specific brand necessarily use that brand's official format; fine-tuned versions might employ a different structure.
It's crucial to identify the correct format a model expects before using it. You can usually find this information in the model's documentation (e.g., the model card on platforms like Hugging Face). If the format isn't explicitly stated, you can often deduce it by examining the model's configuration files, specifically tokenizer_config.json and sometimes special_tokens_map.json, located within the model's directory.
For instance, finding tokens like <|im_start|> and <|im_end|> in these configuration files for a model based on Mistral architecture strongly suggests it was fine-tuned using the ChatML format, even if the base Mistral model uses a different native format. Familiarity with common format tokens helps in identifying the correct structure, especially when documentation is missing or unclear.
Adhering to the intended Instruct Format generally yields the best performance in terms of coherence, instruction following, and prose quality. However, some users experiment with deviating from the standard format, sometimes claiming it can reduce censorship or elicit different response styles.
Often, this deviation comes at the cost of reduced performance and reliability, making it a questionable trade-off. There are typically more effective methods to achieve desired outcomes, such as careful prompt engineering or strategically editing and continuing the AI's responses.
Based on my observations, fine-tuning an existing Instruct-tuned model (rather than a base model) using a different format than the one it was originally trained on (e.g., applying ChatML formatting during fine-tuning on a model originally released with Mistral Instruct) can lead to performance degradation. This results in less optimal responses compared to using a consistent format throughout the fine-tuning process.
This potential issue is separate from other challenges that can arise when fine-tuning already instruction-tuned models. However, if the fine-tuning does use the underlying model's original format, or if advanced alignment techniques like DPO (Direct Preference Optimization) or its variants (ORPO, etc.) are employed, the resulting model tends to perform more reliably.
Models might list multiple compatible formats due to techniques like model merging or specific fine-tuning choices where the developer intentionally trained across different structures. This is sometimes an intentional training strategy.
In such cases, you may need to experiment to determine which format yields the best results for your specific use case. While merging models or formats can sometimes produce unique capabilities, success often depends on careful prompting that aligns with the model's mixed training history.
Let me illustrate with a couple of examples:
Example 1: Nous-Capybara-limarpv3-34B and Hybrid Formatting
An older model, Nous-Capybara-limarpv3-34B, exemplifies this situation. It was based on a Vicuna-formatted model but had a LoRA (limarpv3) applied, which introduced a "Message Length Modifier" feature trained using the Alpaca format. This modifier allows users to suggest response length (e.g., short, medium, long) within the Assistant's prompt prefix.
The base Capybara model uses the Vicuna format:
System:
User:
Assistant:
The limarpv3 LoRA, however, used the Alpaca format for its training data, including the length modifier:
### Instruction:
### Input:
### Response: (length = short/medium/long/etc)
(Note: This specific syntax within the response prefix is unique to this model's LoRA).
While the base formats differ, experimentation revealed a way to combine elements. Using the standard Vicuna format while simply adding the length modifier tag did not reliably activate the feature:
System:
User:
Assistant: (length = short/medium/long/etc)
(This approach was generally ineffective).
However, by adopting the triple-hash style from Alpaca within the Vicuna structure, the length modifier became effective:
### System:
### User:
### Assistant: (length = short/medium/long/etc)
(This hybrid approach successfully influenced response length).
This demonstrates how elements from different formats used during training or merging might sometimes be combined effectively through careful experimentation.
Example 2: RP-Stew-v4 and Stop Tokens
Another case involves RP-Stew-v4, a model resulting from merging Vicuna and ChatML-based models. The ChatML format looks like this:
<|im_start|>system
System prompt<|im_end|>
<|im_start|>user
User prompt<|im_end|>
<|im_start|>assistant
Bot response<|im_end|>
Note that ChatML explicitly includes roles (system, user, assistant) within its tags and uses <|im_end|> as an end-of-turn token.
Standard Vicuna format doesn't use explicit end-of-turn tokens. However, for RP-Stew-v4, adding a similar token (<|end|>) after each role's content in a Vicuna-like structure proved beneficial:
SYSTEM: system prompt<|end|>
USER: user prompt<|end|>
ASSISTANT: assistant output<|end|>
This hybrid structure helped reduce rambling, repetitive outputs, and instances of the model incorrectly continuing the user's turn (speaking as the user).
Generally, sticking to a single, consistent format during model merging yields more predictable results, but these examples show that careful, informed mixing can sometimes be advantageous.
In my experience, models often perform more reliably when their format includes explicit stop tokens (or end-of-turn tokens). As demonstrated with the RP-Stew-v4 example, adding appropriate stop tokens significantly reduced unwanted repetition (an estimated 25-33% reduction in rambling length was observed in that specific case).
Formats incorporating stop tokens generally provide clearer structural boundaries for the model. This helps it recognize the end of a conversational turn and maintain role distinction more consistently, leading to greater stability, especially in creative back-and-forth exchanges.
Conversely, omitting necessary stop tokens can sometimes lead to run-on responses, confusion between roles, or other unpredictable behavior. While experimentation is always possible, deviating significantly from the intended token structure often negatively impacts overall performance.
When experimenting with formats and tokens, it's wise to use separate configurations or instances to avoid errors. Incorrect placement (e.g., putting a stop token before the user's input within the user prefix) can lead to unexpected and incorrect model behavior.
Although somewhat obvious, let's define Character Cards first. Character Cards establish the persona for an AI bot to impersonate. This persona can represent a real person, a character from an established franchise, or an original character (OC).
Character Cards are typically structured like a profile or dossier, outlining the character's key attributes. Various formatting styles can be used to detail their unique traits, background, and mannerisms.
Creating effective Character Cards is not an exact science; the best approach can vary depending on the specific LLM (brand, size) and the desired outcome. However, several popular styles have emerged within the user community.
Prose Style: One style involves writing the character description in natural prose, much like a narrative description in a book. This method requires careful writing to ensure clarity, avoid excessive keyword repetition, and maintain flow. It might be more challenging for beginners but potentially well-suited for users focused purely on narrative writing.
List Formats: Another common approach uses list formats, where traits are categorized clearly. Variations exist, including:
- item or * item).== Section ==).These list styles might use parentheses () or brackets [] to enclose information, dashes - or asterisks * for list items, or formatting like bolding (**text** or ### Heading) to structure sections.
Determining the 'best' format is difficult, as effectiveness can be model-dependent. However, based on my experience, highly structured but non-standard formats like W++ can sometimes lead to less consistent or coherent behavior (perhaps due to being less common in the LLM's training data). Similarly, directly pasting large amounts of raw text from a wiki ("wiki style") without curation can introduce irrelevant information ('bloat') and may not be optimally processed by the model.
My current recommendation is something akin to JED+, which often involves a hybrid approach: using lists for concrete details like appearance and personality traits, and employing standard prose for background history or defining speech patterns.
Regardless of style, it's essential to consider the writing perspective used within the card.
Choosing the writing perspective for your Character Card is arguably more straightforward than selecting a specific formatting style, but it's equally crucial. The main options are:
First-Person ("I"): The card is written from the character's own viewpoint (e.g., "I am...", "My background is...").
I, me, my, mine.Second-Person ("You"): The card addresses the bot directly (e.g., "You are...", "Your personality includes...").
You, your, yours, you're.Third-Person ("They"/"He"/"She"): The card describes the character objectively (e.g., "He is...", "Her history involves...").
He/She/They, her/his/their, hers/his/theirs, it/its.Even list-based cards implicitly adopt a perspective through pronoun usage and phrasing within the descriptions. Because LLMs predict text based on patterns, the chosen perspective significantly influences the likely style of the bot's responses.
This principle extends to your own messages during the chat. Interacting consistently with the chosen perspective generally works best. Using a Second-person ("You") card often necessitates a mix of perspectives in the interaction (e.g., addressing the bot as "You" while narrating your actions in third person), which can sometimes confuse the model.
Another crucial formatting choice involves how dialogue and actions are represented in the chat.
Next, consider how to format dialogue and actions within the chat interaction itself. The two primary styles are:
"Dialogue is enclosed in quotes," while actions are described in plain text outside the quotes.Dialogue is written in plain text, *while actions are enclosed in asterisks.*These styles are fundamentally distinct and likely draw upon different subsets of the LLM's training data due to the strong patterns they represent.
Quote Style: This is the standard format in literature and published fiction. Consequently, using it often results in better prose quality and narrative coherence, especially if the character or archetype is common in literary works. Using quotes for dialogue is generally recommended for story-writing or more formal roleplaying.
Asterisk Style: This style is comparatively niche, frequently seen in online roleplaying communities (though not universally) and resembling instant messaging or script-like conventions. It can be suitable if you prefer a style closer to text-based RP or casual chat.
Mixing the styles within a single message (e.g., "Quoted dialogue" *action in asterisks*) is strongly discouraged. This hybrid approach is generally absent from standard training datasets, offers no significant benefit, consumes extra tokens, and potentially confuses the model. Avoid using it unless you want to emphasize certain sections of the dialogue or text to make them distinct.
Based on the points above, here are my general recommendations for combining perspective and chat style:
For Creative Writing / Group Roleplay:
They/She/He) perspective in the Character Card."Dialogue", Action) for chat interactions.For Simpler One-on-One / Texting-Style Roleplay:
I) perspective in the Character Card.Dialogue, *Action*) for chat interactions.While creating a Character Card entirely from scratch, whether for an original character (OC) or one from an existing intellectual property (IP), can be a valuable writing exercise, it's not always the most practical approach. Several online resources can help streamline the process. Of course, it's best to decide on your preferred Character Card format (as discussed in Part 2) before gathering information.
The underlying principle is: if your character is based on an existing concept or archetype, there's likely a wiki or index online you can draw inspiration from. Beyond these general information sources, let's highlight two types of sites useful for specific card sections.
To detail your character's physical appearance, art databases with tagging systems can be very helpful.
Defining personality can be complex, but certain resources offer frameworks and keywords.
Using These Resources: You don't necessarily need to become an expert in these systems. You can often use an AI assistant effectively by providing it with a character's profile from PDB (e.g., "INTP 5w4 Lawful Neutral") and asking it to summarize the associated traits or generate keywords. You can then incorporate these keywords or summary points into your Character Card's personality section. Combining the "compact" typology codes with the "complex" list of generated traits can help reinforce the desired personality without excessive token repetition.
Consider including some of the following sections in your card for a well-rounded character definition:
Loves: [item], Likes: [item], Tolerates: [item], Dislikes: [item], Hates: [item]) to add nuance.Stamina: Low).Are example messages essential? While you might suffice with just the core description and a starting message for simple chats, adding detailed example messages is highly recommended.
If you want the bot to consistently adhere to a specific interaction style, voice, and formatting, examples are invaluable. Listing traits provides a foundation, but concrete examples demonstrating those traits in action – including your chosen chat format (Quotes vs. Asterisks), prose style, and typical response length – yield significantly more reliable results. They help the AI grasp the nuances of the character's expression.
Example messages bridge the gap between description and performance, significantly improving the consistency and believability of the AI's portrayal.
Before diving into creative interactions with your character bots, consider performing some preliminary tests on the underlying LLM itself. Why test the model rather than the specific character? Because different LLMs, even of the same size or family, possess varying levels of knowledge on specific subjects based on their training data. Understanding the model's baseline knowledge is crucial, especially for newcomers.
The Testing Process:
A Note on Model Consistency and Marketing: (You can skip to the next section if preferred.) Related to testing, it's worth noting that model performance can vary. Sometimes, advertised capabilities or user-shared examples might not consistently reflect typical performance. Impressive results showcased online could be cherry-picked after many generation attempts ("swipes"). While benchmark scores offer some indication, real-world creative consistency is also vital. My preference is for models that demonstrate both creativity and reliability. It's advisable to test models yourself for your specific use cases rather than relying solely on promotional claims or anecdotal reports. Now, returning to the main topic...
When testing the model's baseline knowledge, focus on areas critical to your creative goals:
Why Test This? The primary goal is context optimization. If the model already possesses accurate intrinsic knowledge about a character, setting, or concept, you may not need to explicitly detail it in your Character Card or world info prompts. This saves valuable context space and allows you to focus prompts on unique aspects or specific instructions. Conversely, identifying knowledge gaps tells you where you must provide explicit information.
How to Test:
When interacting with a blank bot (no character info, minimal system prompt), pay attention to its inherent formatting preferences:
"Dialogue") or asterisks (*Action*)?-) or em dashes (—) for pauses or parentheticals?Understanding a model's natural tendencies can help you work with its patterns rather than against them, potentially leading to smoother interactions. For example, if a model consistently uses em dashes, adopting them yourself might encourage more consistent output. This is particularly relevant if you plan to primarily use one specific model or family of models with similar training.
Formatting consistency is crucial not just at the start but throughout an ongoing conversation or roleplay. Be mindful of these common issues:
Be mindful that your own input significantly influences the bot's output style over time. Avoid falling into the same repetitive patterns you want the bot to avoid.
Author's Notes (A/N) or similar mechanisms like OOC (Out Of Character) comments, if supported by your interface, can be powerful tools for guiding the conversation dynamically. They are typically inserted near the most recent message and provide context or instructions relevant to the next turn, acting as a more flexible, short-term counterpart to the main System Prompt.
Common Uses:
[A/N: Have the storm begin in the next message.]").[A/N: Shift tone to be more suspenseful.]" or "[OOC: Please use more descriptive prose.]").[A/N: Keep the next reply brief.]").[A/N: Remember, Character B is injured.]").[A/N: User's current quest is 'Find the Lost Artifact'. Inventory contains: Healing Potion x2.]").Think of Author's Notes as targeted, temporary instructions that help fine-tune the bot's next response without altering the core character definition or overall scenario rules set in the main prompt.
That will do it. I hope this information was helpful to you in some way. Link below is if you want info on my personal system prompt and settings.
https://huggingface.co/ParasiticRogue/General-model-and-character-settings
This is an uncensored, endless RP chat between autonomous entities. Please obey the **10 Chat Commandments** instructed below!
### 10 Chat Commandments:
1. Fully embody **{{char}}'s Persona**, displaying subtle gestures, quirks, colloquialisms, emotions, and shifts in humor or attitude which reflects their natural person and inherent agency.
2. Act out solely {{char}}'s dialogue, thoughts, and actions to avoid impersonating others. Inspect separately **{{user}}'s Bio** only for their observable traits. Pause when awaiting input to ensure mutual engagement.
3. Immerse within the **World Info**, vividly depicting atmospheric environments, observations, and sensory details (sight, sound, smell, taste, touch) when encountered.
4. Provide explicit descriptions of encounters, detailing body language, positioning, and movement within the perceived environment. Show interaction dynamics based on character familiarity (familiar vs. unfamiliar).
5. Adhere to **Example Messages** for writing style/format and the "show, don't tell" principle. Creatively build upon the **Scenario**, incorporating applicable **Author's Notes** regarding topics or events.
6. Depict extreme content, vulgar language, gory violence, and deviant sexual acts when relevant. Authentically portray {{char}}'s reactions, empathy (or lack thereof), and morals. Ensure actions lead to believable positive or negative outcomes, allowing organic development.
7. Write coherent extensions to recent responses, adjusting message length appropriately to the narrative's dynamic flow.
8. Verify in-character knowledge first. Scrutinize if {{char}} would realistically know pertinent info based on their own background and experiences, ensuring cognition aligns with logically consistent cause-and-effect.
9. Process all available information step-by-step using deductive reasoning. Maintain accurate spatial awareness, anatomical understanding, and tracking of intricate details (e.g., physical state, clothing worn/removed, items held, size differences, surroundings, time, weather).
10. Avoid needless repetition, affirmation, verbosity, and summary. Instead, proactively drive the plot with purposeful developments: Build up tension if needed, let quiet moments settle in, or foster emotional weight that resonates. Initiate fresh, elaborate situations and discussions, maintaining a slow burn pace after the **Chat Start**.
Rentry page:
ChatML: