google-agents-cli-scaffold
Project scaffolding, deployment configuration, and CI/CD setup for Google ADK agents.
fixing-tests
Use when tests are failing, test quality issues were identified, or user wants to fix/improve specific tests
Full skill instructions
Detect mode from user input, build work items accordingly.
| Mode | Detection | Action |
|---|---|---|
audit_report | Structured findings with patterns 1-8, "GREEN MIRAGE" verdicts, YAML block | Parse YAML, extract findings |
general_instructions | "Fix tests in X", "test_foo is broken", specific test references | Extract target tests/files |
run_and_fix | "Run tests and fix failures", "get suite green" | Run tests, parse failures |
If unclear: ask user to clarify target.
interface WorkItem {
id: string; // "finding-1", "failure-1", etc.
priority: "critical" | "important" | "minor" | "unknown";
test_file: string;
test_function?: string;
line_number?: number;
pattern?: number; // 1-8 from green mirage
pattern_name?: string;
current_code?: string; // Problematic test code
blind_spot?: string; // What broken code would pass
suggested_fix?: string; // From audit report
production_file?: string; // Related production code
error_type?: "assertion" | "exception" | "timeout" | "skip";
error_message?: string;
expected?: string;
actual?: string;
}
Parse YAML block between --- markers:
findings:
- id: "finding-1"
priority: critical
test_file: "tests/test_auth.py"
test_function: "test_login_success"
line_number: 45
pattern: 2
pattern_name: "Partial Assertions"
blind_spot: "Login could return malformed user object"
depends_on: []
remediation_plan:
phases:
- phase: 1
findings: ["finding-1"]
Use remediation_plan.phases for execution order. Honor depends_on dependencies.
Fallback parsing (if no YAML block):
**Finding #N:** headers**File:****Pattern:****Blind Spot:**A) Per-fix (recommended) - each fix separate commit B) Batch by file C) Single commit
Default to (A).
Skip for audit_report/general_instructions modes.
pytest --tb=short 2>&1 || npm test 2>&1 || cargo test 2>&1
Parse failures into WorkItems with error_type, message, stack trace, expected/actual.
Process by priority: critical > important > minor.
<RULE>Always read before fixing. Never guess at code structure.</RULE>
| Situation | Fix Type |
|---|---|
| Weak assertions (green mirage) | Strengthen assertions |
| Missing edge cases | Add test cases |
| Wrong expectations | Correct expectations |
| Broken setup | Fix setup, not weaken test |
| Flaky (timing/ordering) | Fix isolation/determinism |
| Tests implementation details | Rewrite to test behavior |
| Production code buggy | STOP and report |
PRODUCTION BUG DETECTED
Test: [test_function]
Expected behavior: [what test expects]
Actual behavior: [what code does]
This is not a test issue - production code has a bug.
Options:
A) Fix production bug (then test will pass)
B) Update test to match buggy behavior (not recommended)
C) Skip test, create issue for bug
Your choice: ___
Do NOT silently fix production bugs as "test fixes." </CRITICAL>
Green Mirage Fix (Pattern 2: Partial Assertions):
# BEFORE: Checks existence only
def test_generate_report():
report = generate_report(data)
assert report is not None
assert len(report) > 0
# AFTER: Validates actual content
def test_generate_report():
report = generate_report(data)
assert report == {
"title": "Expected Title",
"sections": [...expected sections...],
"generated_at": mock_timestamp
}
# OR at minimum:
assert report["title"] == "Expected Title"
assert len(report["sections"]) == 3
assert all(s["valid"] for s in report["sections"])
Edge Case Addition:
def test_generate_report_empty_data():
"""Edge case: empty input."""
with pytest.raises(ValueError, match="Data cannot be empty"):
generate_report([])
def test_generate_report_malformed_data():
"""Edge case: malformed input."""
result = generate_report({"invalid": "structure"})
assert result["error"] == "Invalid data format"
Flaky Test Fix:
# BEFORE: Sleep and hope
def test_async_operation():
start_operation()
time.sleep(1) # Hope it's done!
assert get_result() is not None
# AFTER: Deterministic waiting
def test_async_operation():
start_operation()
result = wait_for_result(timeout=5) # Polls with timeout
assert result == expected_value
Implementation-Coupling Fix:
# BEFORE: Tests implementation
def test_user_save():
user = User(name="test")
user.save()
assert user._db_connection.execute.called_with("INSERT...")
# AFTER: Tests behavior
def test_user_save():
user = User(name="test")
user.save()
loaded = User.find_by_name("test")
assert loaded is not None
assert loaded.name == "test"
# Run fixed test
pytest path/to/test.py::test_function -v
# Check file for side effects
pytest path/to/test.py -v
Verification checklist:
git add path/to/test.py
git commit -m "fix(tests): strengthen assertions in test_function
- [What was weak/broken]
- [What fix does]
- Pattern: N - [Pattern name] (if from audit)
"
FOR priority IN [critical, important, minor]:
FOR item IN work_items[priority]:
Execute Phase 2
IF stuck after 2 attempts:
Add to stuck_items[]
Continue to next item
## Stuck Items
### [item.id]: [test_function]
**Attempted:** [what was tried]
**Blocked by:** [why it didn't work]
**Recommendation:** [manual intervention / more context / etc.]
Run full test suite:
pytest -v # or appropriate test command
## Fix Tests Summary
### Input Mode
[audit_report / general_instructions / run_and_fix]
### Metrics
| Metric | Value |
|--------|-------|
| Total items | N |
| Fixed | X |
| Stuck | Y |
| Production bugs | Z |
### Fixes Applied
| Test | File | Issue | Fix | Commit |
|------|------|-------|-----|--------|
| test_foo | test_auth.py | Pattern 2 | Strengthened to full object match | abc123 |
### Test Suite Status
- Before: X passing, Y failing
- After: X passing, Y failing
### Stuck Items (if any)
[List with recommendations]
### Production Bugs Found (if any)
[List with recommended actions]
Fixes complete. Re-run audit-green-mirage to verify no new mirages?
A) Yes, audit fixed files
B) No, satisfied with fixes
Flaky tests: Identify non-determinism source (time, random, ordering, external state). Mock or control it. Use deterministic waits, not sleep-and-hope.
Implementation-coupled tests: Identify BEHAVIOR test should verify. Rewrite to test through public interface. Remove internal mocking.
Missing tests entirely: Read production code. Identify key behaviors. Write tests following codebase patterns. Ensure tests would catch real failures.
<FORBIDDEN> ## Anti-Patterns<RULE>Before completing, ALL boxes must be checked. If ANY unchecked: STOP and fix.</RULE>
<FINAL_EMPHASIS> Tests exist to catch bugs. Every fix you make must result in tests that actually catch failures, not tests that achieve green checkmarks.
Fix it. Prove it works. Move on. No over-engineering. No under-testing. </FINAL_EMPHASIS>
Project scaffolding, deployment configuration, and CI/CD setup for Google ADK agents.
Set up tracing, logging, and monitoring for deployed ADK agents across Cloud Trace, BigQuery, and third-party platforms.
Enterprise Azure infrastructure architect generating Bicep or Terraform from workload descriptions.
Plan and configure production-ready Azure Kubernetes Service clusters with Day-0 and Day-1 best practices.
Raw mechanical interfaces fusing Swiss typographic print with military terminal aesthetics. Rigid grids, extreme type scale contrast, utilitarian color, analog degradation effects. For data-heavy dashboards, portfolios, or editorial sites that need to feel like declassified blueprints.
Web search, scraping, extraction, crawling, and monitoring via ScrapeGraph AI CLI.
Skill for working with Firebase Hosting (Classic). Use this when you want to deploy static web apps, Single Page Apps (SPAs), or simple microservices. Do NOT use for Firebase App Hosting.
Deploy and manage web apps with Firebase App Hosting. Use this skill when deploying Next.js/Angular apps with backends.
Deploy applications and websites to Vercel. Use when the user requests deployment actions like "deploy my app", "deploy and give me the link", "push this live", or "create a preview deployment".
Build SEO-optimized pages at scale using templates, data, and proven playbook patterns.
Design and build isolated, reusable Convex backend components with clear boundaries and app-facing wrappers.
Deploy and manage projects on Vercel using token-based authentication. Use when working with Vercel CLI using access tokens rather than interactive login — e.g. "deploy to vercel", "set up vercel", "add environment variables to vercel".
Opus.ai: Revolutionize Your Web Experience
Build a no-code AI app in minutes.
An IDE for code migration from legacy to modern frameworks through coding agents.
Automate CGI animation in live-action scenes
Branded artistic QR-code concepts
Launch a website in seconds with AI.
Streamline Your Coding Experience with AI Code Helper
Convert any screenshot or design to clean code.