Skip to content
codex-readiness-integration-test logo

LLM Codex Readiness Integration Test

codex-readiness-integration-test

Run the Codex Readiness integration test. Use when you need an end-to-end agentic loop with build/test scoring.

SKILL.md

Full skill instructions

LLM Codex Readiness Integration Test

This skill runs a multi-stage integration test to validate agentic execution quality. It always runs in execute mode (no read-only mode).

Outputs

Each run writes to .codex-readiness-integration-test/<timestamp>/ and updates .codex-readiness-integration-test/​latest.json.

New outputs per run:

  • agentic_summary.json and logs/​agentic.log (agentic loop execution)
  • llm_results.json (automatic LLM evaluation)
  • summary.txt (human-readable summary)

Pre-conditions (Required)

  • Authenticate with the Codex CLI using the repo-local HOME before running the test. Run these in your own terminal (not via the integration test): HOME=$PWD/​.codex-home XDG_CACHE_HOME=$PWD/​.codex-home/​.cache codex login HOME=$PWD/​.codex-home XDG_CACHE_HOME=$PWD/​.codex-home/​.cache codex login status
  • The integration test creates {repo_root}/​.codex-home and {repo_root}/​.codex-home/​.cache/​codex as its first step.

Workflow

  1. Ask the user how to source the task.
    • Offer two explicit options: (a) user provides a custom task/​prompt, or (b) auto-generate a task.
    • Do not run the entry point until the user chooses one option.
  2. Generate or load {out_dir}/​prompt.pending.json.
    • Use the integration test's expected prompt path, not prompt.json at the repo root.
    • With the default out dir, this path is .codex-readiness-integration-test/​prompt.pending.json.
    • If --seed-task is provided, it is used as the starting task.
    • If not provided, generate a task with skills/​codex-readiness-integration-test/​references/​generate_prompt.md and save the JSON to {out_dir}/​prompt.pending.json.
    • The user must approve the prompt before execution (no auto-approve mode). Make sure to output a summary of the prompt when asking the user to approve.
  3. Execute the agentic loop via Codex CLI (uses AGENTS.md and change_prompt).
  4. Run build/​test commands from the prompt plan via skills/​codex-readiness-integration-test/​scripts/​run_plan.py.
  5. Collect evidence (evidence.json), deterministic checks, and run automatic LLM evals via Codex CLI.
  6. Score and write the report + summary output.

Configuration

Optional fields in {out_dir}/​prompt.pending.json:

  • agentic_loop: configure Codex CLI invocation for the agentic loop.
  • llm_eval: configure Codex CLI invocation for automatic evals.

If these fields are omitted, defaults are used.

Requirements

  • The LLM evaluator must fail if evidence mentions the phrase Context compaction enabled.
  • Use qualitative context-usage evaluation (no strict thresholds).

What this test covers well

  • Runs Codex CLI against the real repo root, producing real filesystem edits and git diffs.
  • Executes the approved change prompt and then runs the build/​test plan in-repo.
  • Captures evidence, deterministic checks, and LLM eval artifacts for review.

What this test does not represent

  • The agentic loop may use non-default flags (e.g., bypass approvals/​sandbox), so interactive guardrails differ.
  • Uses a dedicated HOME (.codex-home), which can change auth/​config/​cache vs normal CLI use.
  • Auto-generated prompts and one-shot execution do not simulate interactive guidance.
  • MCP servers/​tools are not exercised unless explicitly configured.

Notes

  • The prompts in skills/​codex-readiness-integration-test/​references/ expect strict JSON.
  • Use skills/​codex-readiness-integration-test/​references/​json_fix.md to repair invalid JSON output.
  • This skill calls the codex CLI. Ensure it is installed and available on PATH, or override the command in {out_dir}/​prompt.pending.json.
  • If the agentic loop detects sandbox-blocked tool access, it now writes requires_escalation: true to {run_dir}/​agentic_summary.json and exits with code 3. Re-run the integration test with escalated permissions in that case.

More Productivity & Planning skills

brainstorming logo
Productivity & Planning

brainstorming

Structured design dialogue that validates ideas before implementation begins.

295.4K 385.7K
View
ui-ux-pro-max logo
Productivity & Planning

ui-ux-pro-max

Comprehensive design intelligence for web and mobile UI/UX across 10 technology stacks.

133.1K 383.9K
View
writing-plans logo
Productivity & Planning

writing-plans

Comprehensive implementation plans for multi-step tasks, breaking down specs into bite-sized, testable steps.

295.4K 268K
View
using-superpowers logo
Productivity & Planning

using-superpowers

Introduction to the obra skills system with mandatory skill invocation rules and best practices.

295.4K 259.8K
View
executing-plans logo
Productivity & Planning

executing-plans

Execute a written implementation plan with critical review and task checkpoints.

295.4K 229.1K
View
dispatching-parallel-agents logo
Productivity & Planning

dispatching-parallel-agents

Delegate independent tasks to specialized agents working concurrently with isolated context.

295.4K 206.3K
View
using-git-worktrees logo
Productivity & Planning

using-git-worktrees

Isolated git worktrees with smart directory selection and safety verification.

295.4K 205.2K
View
webapp-testing logo
Productivity & Planning

webapp-testing

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

179.7K 170.6K
View
content-strategy logo
Productivity & Planning

content-strategy

Plan searchable and shareable content that drives traffic, builds authority, and generates leads.

53.3K 150.9K
View
repo-intake-and-plan logo
Productivity & Planning

repo-intake-and-plan

README-first repository scanner that extracts commands and classifies reproduction candidates without executing them.

497 139.6K
View
marketing-ideas logo
Productivity & Planning

marketing-ideas

Brainstorm and prioritize marketing strategies tailored to your SaaS stage, budget, and goals.

53.3K 137.2K
View
site-architecture logo
Productivity & Planning

site-architecture

Plan and optimize your website's page hierarchy, navigation, URL structure, and internal linking.

53.3K 112.8K
View

Productivity AI tools

Vimcal logo
Productivity

Vimcal

The world's fastest calendar for remote work

Free
View
SaveDay logo
Productivity

SaveDay

Capture, organize, and utilize your knowledge effortlessly.

Free
View
A
Productivity

Any Summary

Instant Summaries of Audio & Video Interviews with AnySummary

Freemium
View
M
Productivity

Map This

Transform PDFs into engaging mind maps.

Freemium
View
ChatPDF logo
Productivity

ChatPDF

Chat with any PDF instantly

Free
View
I
Productivity

intellisay

Create an optimal daily plan using your voice

Paid
View
A
Productivity

Aurora AI

A productivity platform to centralize organizational knowledge and workflows with contextual AI assistance.

Paid
View
AskYourPDF logo
Productivity

AskYourPDF

AskYourPDF Pricing Plans: Tailored to Your Needs

Free
View