Skip to content
skillgrade-graders logo

Skillgrade Grader Authoring

skillgrade-graders

Authors deterministic and LLM rubric graders for skillgrade evaluations. Use when creating scoring scripts, writing evaluation rubrics, or combining multiple graders with weighted scoring. Don't use for setting up eval pipelines, configuring eval.yaml defaults, or general test writing.

SKILL.md

Full skill instructions

Skillgrade Grader Authoring

Procedures

Step 1: Identify the Grading Strategy

  1. Determine whether the task requires objective verification (deterministic) or qualitative assessment (LLM rubric).
  2. For most tasks, combine both: deterministic graders verify outcomes (weight 0.7), LLM rubrics assess approach quality (weight 0.3).

Step 2: Write a Deterministic Grader

  1. Create a script in the skill's graders/ directory (bash or TypeScript).
  2. The script must output a JSON object to stdout with the following structure:
    {"score": 0.67, "details": "2/​3 checks passed", "checks": [{"name": "check-name", "passed": true, "message": "Description"}]}
    
  3. score (0.0–1.0) and details are required. checks is optional but recommended.
  4. Read references/​grader-output-schema.md for the full output specification.
  5. Use awk for arithmetic in bash scripts — bc is not available in node:20-slim.
  6. Reference the grader in eval.yaml:
    - type: deterministic
      run: bash graders/​check.sh
      weight: 0.7
    

Step 3: Write an LLM Rubric Grader

  1. Draft a rubric with explicit scoring criteria and point allocations.
  2. Structure the rubric into weighted sections that sum to 1.0:
    Workflow Compliance (0-0.5):
    - Did the agent follow the mandatory workflow steps?
    Efficiency (0-0.5):
    - Completed in ≤5 commands without trial-and-error?
    
  3. Reference the rubric in eval.yaml:
    - type: llm_rubric
      rubric: |
        [rubric text or file path]
      weight: 0.3
      provider: gemini               # optional: gemini (default) | anthropic | openai
      model: gemini-3-flash-preview  # optional, each provider has a default model
    
  4. For long rubrics, store in a separate file and reference by path: rubric: rubrics/​quality.md.

Step 4: Combine Multiple Graders

  1. Assign weights to each grader based on importance. Weights are normalized automatically.
  2. Final reward is calculated as: Σ (grader_score × weight) / Σ weight.
  3. Example configuration:
    graders:
      - type: deterministic
        run: bash graders/​check.sh
        weight: 0.7
      - type: llm_rubric
        rubric: rubrics/​quality.md
        weight: 0.3
    

Step 5: Validate Graders

  1. Create a reference solution script that produces the expected output.
  2. Run skillgrade --validate to verify graders score the reference solution correctly.
  3. Test only deterministic graders: skillgrade --grader=deterministic (skips LLM calls, faster iteration).
  4. Test only LLM rubric graders: skillgrade --grader=llm_rubric.
  5. Run a specific eval with a specific grader type: skillgrade --eval=my-eval --grader=deterministic.
  6. If a grader returns unexpected scores, inspect the script output and adjust scoring logic.

Error Handling

  • If a deterministic grader outputs non-JSON, ensure all echo/console.log statements except the final JSON result are redirected to stderr.
  • If an LLM rubric grader returns 0.00 with a missing API key message, set the appropriate key for your provider: GEMINI_API_KEY (provider: gemini), ANTHROPIC_API_KEY (provider: anthropic), or OPENAI_API_KEY (provider: openai).
  • To use a custom/​self-hosted LLM endpoint, set ANTHROPIC_BASE_URL (for provider: anthropic) or OPENAI_BASE_URL (for provider: openai) — e.g. for Ollama or vLLM.
  • If scores are inconsistent across trials, reduce rubric ambiguity by adding concrete examples of passing and failing behavior.

More Productivity & Planning skills

brainstorming logo
Productivity & Planning

brainstorming

Structured design dialogue that validates ideas before implementation begins.

295.4K 385.7K
View
ui-ux-pro-max logo
Productivity & Planning

ui-ux-pro-max

Comprehensive design intelligence for web and mobile UI/UX across 10 technology stacks.

133.1K 383.9K
View
writing-plans logo
Productivity & Planning

writing-plans

Comprehensive implementation plans for multi-step tasks, breaking down specs into bite-sized, testable steps.

295.4K 268K
View
using-superpowers logo
Productivity & Planning

using-superpowers

Introduction to the obra skills system with mandatory skill invocation rules and best practices.

295.4K 259.8K
View
executing-plans logo
Productivity & Planning

executing-plans

Execute a written implementation plan with critical review and task checkpoints.

295.4K 229.1K
View
dispatching-parallel-agents logo
Productivity & Planning

dispatching-parallel-agents

Delegate independent tasks to specialized agents working concurrently with isolated context.

295.4K 206.3K
View
using-git-worktrees logo
Productivity & Planning

using-git-worktrees

Isolated git worktrees with smart directory selection and safety verification.

295.4K 205.2K
View
webapp-testing logo
Productivity & Planning

webapp-testing

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

179.7K 170.6K
View
content-strategy logo
Productivity & Planning

content-strategy

Plan searchable and shareable content that drives traffic, builds authority, and generates leads.

53.3K 150.9K
View
repo-intake-and-plan logo
Productivity & Planning

repo-intake-and-plan

README-first repository scanner that extracts commands and classifies reproduction candidates without executing them.

497 139.6K
View
marketing-ideas logo
Productivity & Planning

marketing-ideas

Brainstorm and prioritize marketing strategies tailored to your SaaS stage, budget, and goals.

53.3K 137.2K
View
site-architecture logo
Productivity & Planning

site-architecture

Plan and optimize your website's page hierarchy, navigation, URL structure, and internal linking.

53.3K 112.8K
View

Productivity AI tools

Vimcal logo
Productivity

Vimcal

The world's fastest calendar for remote work

Free
View
SaveDay logo
Productivity

SaveDay

Capture, organize, and utilize your knowledge effortlessly.

Free
View
A
Productivity

Any Summary

Instant Summaries of Audio & Video Interviews with AnySummary

Freemium
View
M
Productivity

Map This

Transform PDFs into engaging mind maps.

Freemium
View
ChatPDF logo
Productivity

ChatPDF

Chat with any PDF instantly

Free
View
I
Productivity

intellisay

Create an optimal daily plan using your voice

Paid
View
A
Productivity

Aurora AI

A productivity platform to centralize organizational knowledge and workflows with contextual AI assistance.

Paid
View
AskYourPDF logo
Productivity

AskYourPDF

AskYourPDF Pricing Plans: Tailored to Your Needs

Free
View