Downloads · 30 days
20
35% of all-time downloads
xap/Llama3.1
Llama3.1 is a machine learning model from xap. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repository hosts an instruction fine-tuned large language model for semantic similarity prediction between a student explanation and an expert explanation of a code line.
Downloads · 30 days
20
35% of all-time downloads
All-time downloads
57
Public
Parameters
8B
16.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
This repository hosts an instruction fine-tuned large language model for semantic similarity prediction between a student explanation and an expert explanation of a code line.
The model is designed for educational assessment settings where the goal is to measure how closely a student's explanation matches the meaning of an expert reference explanation. Given a code snippet, a student explanation, and an expert explanation, the model returns a similarity score from 0 to 1, where:
This model was fine-tuned to support automated evaluation of programming explanations. It is especially useful for tasks such as:
The model takes three pieces of information:
It then predicts a single similarity score between 0 and 1.
The model expects input in Alpaca-style instruction format.
{
"instruction": "Assess the semantic similarity of two sentences using a similarity score on a scale from 0 to 1, with 0 indicating minimal or no semantic similarity and 1 representing maximal semantic similarity. Only provide the score without any other output. The input consists of a line of code followed by a student explanation of the line of code and an expert explanation of the line of code as given below.",
"input": "Code: System.out.println(Duplicate input for number: + num);\nStudent Explanation: If the two saved adjacent values from the sequence are duplicates, this line will run, printing out that a duplicate of the current number has been found.\nExpert Explanation: This statement prints the duplicate number to the default standard output stream.",
"output": 0.75,
"prompt": "Your task is to compute how semantically similar a student explanation of a line of code written in JAVA is compared to an expert explanation of the same line of code. Semantic similarity measures how close the meanings of the two explanations are. The goal is to assess whether the student explains the line of code as well as the expert. ### Instruction: Assess the semantic similarity of two sentences using a similarity score on a scale from 0 to 1, with 0 indicating minimal or no semantic similarity and 1 representing maximal semantic similarity. Only provide the score without any other output. The input consists of a line of code followed by a student explanation of the line of code and an expert explanation of the line of code as given below. ### Input: Code: System.out.println(Duplicate input for number: + num);\nStudent Explanation: If the two saved adjacent values from the sequence are duplicates, this line will run, printing out that a duplicate of the current number has been found.\nExpert Explanation: This statement prints the duplicate number to the default standard output stream. ### Response:"
}
Load the model from Hugging Face using transformers:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"xap/Mistral2,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(
"xap/Mistral2"
)
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"xap/CodeLlama-Instruction-Assessment-new1",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(
"xap/CodeLlama-Instruction-Assessment-new1"
)
prompt = """Below is an instruction that describes a task, paired with an input that provides further context. Write a response that appropriately completes the request.\n\n### Instruction:\nFor the given line of code, both the student and expert have provided the explanation for that line of code. Compute the semantic similarity between the student explanation and the expert explanation for the line of code.\n\n### Input:\nFor given line of code int[] values = {5, 8, 4, 78, 95, 12, 1, 0, 6, 35, 46};, the expert explanation is We declare an array of values to hold the numbers. and the student explanation is This line creates the integer array with the values. you need this to achieve the goal bc you need an array to look in\n\n### Response:\n"""
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=5)
result = tokenizer.decode(outputs, skip_special_tokens=True)
print(result)
For best results, format inference inputs as follows:
Below is an instruction that describes a task, paired with an input that provides further context. Write a response that appropriately completes the request.
### Instruction:
Assess the semantic similarity of two sentences using a similarity score on a scale from 0 to 1, with 0 indicating minimal or no semantic similarity and 1 representing maximal semantic similarity. Only provide the score without any other output.
### Input:
Code: <code line>
Student Explanation: <student explanation>
Expert Explanation: <expert explanation>
### Response:
The expected output is a single numeric score between 0 and 1.
Example:
0.75
This model is intended for:
If this model contributes to your research or application, cite the associated project, paper, or repository as appropriate.
This model was developed for semantic assessment of student explanations against expert explanations in programming education contexts.