Downloads · 30 days
0
richardyoung/llm-instruction-following-code
llm-instruction-following-code is a machine learning model from richardyoung. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
[](http://arxiv.org/abs/2510.18892) [](https://huggingface.co/datasets/richardyoung/llm-instruction-following-eval) [](https://www.python.org/)
Downloads · 30 days
0
Access
Public
Updated Oct 24, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.py49.1 KB · 65%
From the Hugging Face model README
This repository contains the complete evaluation framework used in our paper "When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs" (arXiv:2510.18892).
This code repository provides everything needed to:
# Clone the repository or download files
pip install pandas openpyxl requests matplotlib seaborn numpy
# Set your OpenRouter API key
export OPENROUTER_API_KEY="your_api_key_here"
# Run comprehensive evaluation (256 models × 20 tests)
python test_comprehensive_20_verified.py
# Generate analysis and visualizations
python analyze_comprehensive_final.py
test_comprehensive_20_verified.py - Main test runner
questions.json - Complete test bank (20 diagnostic prompts)
models_verified_working_v2_20251014_091649.py - Model configuration
analyze_comprehensive_final.py - Comprehensive analysis pipeline
requirements.txt - Python dependenciesREADME.md - This file (setup and usage instructions)Our 20 diagnostic tests cover five categories:
From our October 14, 2025 evaluation of 256 models:
Edit questions.json to add new diagnostic tests:
{
"id": 21,
"test_name": "Your New Test",
"category": "Custom Category",
"difficulty": "medium",
"prompt": "Your instruction prompt here",
"expected_output": "Exact expected response",
"exact_match": true,
"case_sensitive": false
}
Modify models_verified_working_v2_20251014_091649.py or create your own model list:
MODELS = [
{
"name": "provider/model-name",
"provider": "provider",
"verified": True
},
# Add more models...
]
Customize analyze_comprehensive_final.py to:
The evaluation produces:
Excel Workbook (comprehensive_20_tests_results_YYYYMMDD_HHMMSS.xlsx)
JSON Export (comprehensive_20_tests_results_YYYYMMDD_HHMMSS.json)
PDF Visualizations
fig1_heatmap.pdf - Performance matrixfig2_provider.pdf - Provider comparisonfig3_difficulty.pdf - Test difficultyfig4_category.pdf - Category performanceLaTeX Tables (paper_tables.tex)
To exactly reproduce our paper results:
# Use the frozen model list from October 14, 2025
python test_comprehensive_20_verified.py
# Use the frozen test bank
# (questions.json is already frozen at 20 tests)
# Generate analysis with same parameters
python analyze_comprehensive_final.py
Note: Model outputs may vary over time as providers update their models. For exact reproducibility, use the snapshot from our evaluation date.
# Edit test_comprehensive_20_verified.py
# Change MODELS to a subset:
MODELS = [
"openai/gpt-4o",
"anthropic/claude-3.7-sonnet",
"google/gemini-2.0-flash-exp:free",
"meta-llama/llama-3.3-70b-instruct",
"qwen/qwen-plus-2025-07-28:thinking"
]
import requests
import json
# Load questions
with open('questions.json', 'r') as f:
questions = json.load(f)
# Test a single model
model = "openai/gpt-4o"
for q in questions:
response = requests.post(
"https://openrouter.ai/api/v1/chat/completions",
headers={"Authorization": f"Bearer {OPENROUTER_API_KEY}"},
json={
"model": model,
"messages": [{"role": "user", "content": q["prompt"]}]
}
)
# Evaluate response...
import pandas as pd
# Load results
df = pd.read_excel('results.xlsx', sheet_name='All Results')
# Custom analysis
top_models = df.groupby('model')['passed'].mean().sort_values(ascending=False).head(10)
print(top_models)
# Category performance
category_perf = df.groupby('category')['passed'].mean()
print(category_perf)
1. API Rate Limiting
# OpenRouter may rate limit. Add delays between requests:
time.sleep(1) # Add to test_comprehensive_20_verified.py
2. JSON Serialization Errors
# Use export_json_from_excel.py to convert numpy types
python export_json_from_excel.py
3. Missing Packages
pip install pandas openpyxl requests matplotlib seaborn numpy
4. API Key Not Set
export OPENROUTER_API_KEY="your_key_here"
# Or set in Python: os.environ['OPENROUTER_API_KEY'] = "your_key"
If you use this code in your research, please cite:
@article{young2025instruction,
title={When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs},
author={Young, Richard J. and Gillins, Brandon and Matthews, Alice M.},
journal={arXiv preprint arXiv:2510.18892},
year={2025}
}
Research Team:
Affiliation: University of Nevada, Las Vegas
This code is released under the MIT License.
MIT License
Copyright (c) 2025 Richard J. Young, Brandon Gillins, Alice M. Matthews
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
Repository Version: 1.0 Last Updated: October 23, 2025 Evaluation Date: October 14, 2025