google-agents-cli-scaffold
Project scaffolding, deployment configuration, and CI/CD setup for Google ADK agents.
production-api-tester
Live testing and validation of production research API for strategy optimization loops
Full skill instructions
Skill for testing the production research API in live environments. Enables optimization loops: create strategy → test in production → analyze with Langfuse → iterate.
5 Core Functions:
# Production API Configuration
export PROD_API_URL="https://webresearchagent.replit.app" # Production URL
export PROD_API_KEY="your_api_key_here" # Your API secret key
export CALLBACK_URL="https://webhook.site/your-unique-url" # Webhook receiver URL
# Optional: Override for local testing
export PROD_API_URL="http://localhost:8000" # Test locally
Getting Your API Key:
--api-key flagSetting Up Webhook Receiver:
Goal: Execute a strategy once and check results
cd /home/user/web_research_agent/.claude/skills/production-api-tester/helpers
# Create test task
python3 create_test_task.py \
--api-key "$PROD_API_KEY" \
--email "[email protected]" \
--topic "AI developments in healthcare" \
--frequency "daily" \
--output /tmp/api_test/task.json
Output:
{
"id": "abc123",
"email": "[email protected]",
"research_topic": "AI developments in healthcare",
"frequency": "daily",
"schedule_time": "09:00",
"is_active": true,
"created_at": "2025-11-09T10:00:00Z",
"last_run_at": null
}
# Execute batch (all daily tasks)
python3 execute_batch.py \
--api-key "$PROD_API_KEY" \
--frequency "daily" \
--callback-url "$CALLBACK_URL" \
--output /tmp/api_test/execution.json
Output:
{
"status": "running",
"frequency": "daily",
"tasks_found": 1,
"started_at": "2025-11-09T10:01:00Z"
}
# Option A: Poll webhook endpoint
python3 check_webhook.py \
--webhook-url "$CALLBACK_URL" \
--wait-for-results \
--timeout 300
# Option B: Check task status
python3 get_task.py \
--api-key "$PROD_API_KEY" \
--task-id "abc123" \
--output /tmp/api_test/task_status.json
# Get Langfuse trace for this execution
python3 link_to_langfuse.py \
--task-id "abc123" \
--email "[email protected]" \
--output /tmp/api_test/langfuse_link.json
Output:
{
"task_id": "abc123",
"trace_query": {
"metadata_filter": {
"user_email": "[email protected]",
"research_topic": "AI developments in healthcare"
},
"time_range": "last_1_hour"
},
"langfuse_url": "https://cloud.langfuse.com/project/.../traces?filter=...",
"trace_count": 1,
"latest_trace_id": "xyz789"
}
# Delete test task
python3 delete_task.py \
--api-key "$PROD_API_KEY" \
--task-id "abc123"
Goal: Test new strategy → analyze performance → iterate
This combines strategy-builder skill with production-api-tester skill.
cd /home/user/web_research_agent/.claude/skills/production-api-tester/helpers
# 1. Generate/update strategy (using strategy-builder skill)
python3 ../../strategy-builder/helpers/generate_strategy.py \
--slug "legal/court_cases_de" \
--category "legal" \
--time-window "month" \
--depth "comprehensive" \
--output /tmp/optimization_loop/strategy_v1.yaml
# 2. Validate strategy
python3 ../../strategy-builder/helpers/validate_strategy.py \
--strategy /tmp/optimization_loop/strategy_v1.yaml
# 3. Deploy strategy to database
# (Manual: save to strategies/, add to index.yaml, migrate to DB)
# 4. Create test task for this strategy
python3 create_test_task.py \
--api-key "$PROD_API_KEY" \
--email "[email protected]" \
--topic "Datenschutz DSGVO Verstoß" \
--frequency "daily" \
--output /tmp/optimization_loop/task.json
# 5. Execute research
TASK_ID=$(jq -r '.id' /tmp/optimization_loop/task.json)
python3 execute_batch.py \
--api-key "$PROD_API_KEY" \
--frequency "daily" \
--callback-url "$CALLBACK_URL" \
--output /tmp/optimization_loop/execution.json
# 6. Wait for completion and get results
python3 wait_for_completion.py \
--task-id "$TASK_ID" \
--api-key "$PROD_API_KEY" \
--timeout 600 \
--output /tmp/optimization_loop/results.json
# 7. Link to Langfuse trace
python3 link_to_langfuse.py \
--task-id "$TASK_ID" \
--email "[email protected]" \
--output /tmp/optimization_loop/langfuse.json
# 8. Retrieve and analyze Langfuse trace
TRACE_ID=$(jq -r '.latest_trace_id' /tmp/optimization_loop/langfuse.json)
python3 ../../langfuse-optimization/helpers/retrieve_single_trace.py \
"$TRACE_ID" \
--filter-essential \
--output /tmp/optimization_loop/trace.json
# 9. Analyze performance
python3 ../../strategy-builder/helpers/analyze_strategy_performance.py \
--traces /tmp/optimization_loop/trace.json \
--strategy "legal/court_cases_de" \
--output /tmp/optimization_loop/performance.json
# 10. Review recommendations and iterate
cat /tmp/optimization_loop/performance.json
# Apply fixes to strategy YAML, repeat from step 1
Goal: Test multiple strategies in parallel
cd /home/user/web_research_agent/.claude/skills/production-api-tester/helpers
# Create tasks for different strategies
python3 batch_create_tasks.py \
--api-key "$PROD_API_KEY" \
--tasks-file /tmp/batch_test/tasks_config.json \
--output /tmp/batch_test/created_tasks.json
tasks_config.json:
[
{
"email": "[email protected]",
"research_topic": "AI regulation updates",
"frequency": "daily",
"strategy_hint": "daily_news_briefing"
},
{
"email": "[email protected]",
"research_topic": "Tesla stock analysis",
"frequency": "daily",
"strategy_hint": "financial_research"
},
{
"email": "[email protected]",
"research_topic": "GDPR compliance updates",
"frequency": "daily",
"strategy_hint": "legal/court_cases_de"
}
]
# Execute daily batch
python3 execute_batch.py \
--api-key "$PROD_API_KEY" \
--frequency "daily" \
--callback-url "$CALLBACK_URL"
# Wait for all tasks to complete
python3 monitor_batch.py \
--api-key "$PROD_API_KEY" \
--tasks-file /tmp/batch_test/created_tasks.json \
--timeout 900 \
--output /tmp/batch_test/batch_results.json
# Generate comparison report
python3 compare_strategies.py \
--results /tmp/batch_test/batch_results.json \
--output /tmp/batch_test/comparison.json
Output: Comparison of latency, success rate, error types across strategies
# Delete all test tasks
python3 batch_delete_tasks.py \
--api-key "$PROD_API_KEY" \
--tasks-file /tmp/batch_test/created_tasks.json
Goal: Verify production API is healthy
cd /home/user/web_research_agent/.claude/skills/production-api-tester/helpers
# Quick health check
python3 health_check.py \
--api-url "$PROD_API_URL"
# Extended monitoring
python3 health_check.py \
--api-url "$PROD_API_URL" \
--continuous \
--interval 60 \
--duration 3600
Output:
✓ API is healthy
Status: online
Database: connected
Langfuse: enabled
Response time: 234ms
Purpose: Create research task subscription
Usage:
python3 create_test_task.py \
--api-key "$PROD_API_KEY" \
[--api-url "$PROD_API_URL"] \
--email "[email protected]" \
--topic "Research topic" \
--frequency daily|weekly|monthly \
[--schedule-time "09:00"] \
--output /tmp/task.json
Output: Task object with ID for tracking
Purpose: Trigger batch research execution
Usage:
python3 execute_batch.py \
--api-key "$PROD_API_KEY" \
[--api-url "$PROD_API_URL"] \
--frequency daily|weekly|monthly \
--callback-url "https://webhook.site/..." \
--output /tmp/execution.json
Output: Execution status (running, tasks_found, started_at)
Purpose: Retrieve task details
Usage:
python3 get_task.py \
--api-key "$PROD_API_KEY" \
[--api-url "$PROD_API_URL"] \
--task-id "abc123" \
--output /tmp/task.json
Purpose: List all research tasks
Usage:
python3 list_tasks.py \
--api-key "$PROD_API_KEY" \
[--api-url "$PROD_API_URL"] \
[--email "[email protected]"] \
[--frequency daily] \
--output /tmp/tasks.json
Purpose: Delete research task
Usage:
python3 delete_task.py \
--api-key "$PROD_API_KEY" \
[--api-url "$PROD_API_URL"] \
--task-id "abc123"
Purpose: Poll task until completion
Usage:
python3 wait_for_completion.py \
--api-key "$PROD_API_KEY" \
--task-id "abc123" \
[--timeout 600] \
[--poll-interval 10] \
--output /tmp/results.json
Purpose: Find Langfuse trace for task execution
Usage:
python3 link_to_langfuse.py \
--task-id "abc123" \
--email "[email protected]" \
[--time-range "last_1_hour"] \
--output /tmp/langfuse_link.json
Output:
Purpose: Check API health
Usage:
python3 health_check.py \
[--api-url "$PROD_API_URL"] \
[--continuous] \
[--interval 60] \
[--duration 3600]
Purpose: Run local webhook receiver for testing
Usage:
# Start local webhook receiver
python3 webhook_receiver.py \
--port 8080 \
--output-dir /tmp/webhooks
# Use as callback URL
export CALLBACK_URL="http://localhost:8080/webhook"
Features:
Optimization Loop:
1. strategy-builder: analyze_research_query.py
→ Determine if new strategy needed
2. strategy-builder: generate_strategy.py
→ Create strategy YAML
3. strategy-builder: validate_strategy.py
→ Validate structure
4. [Manual: Deploy strategy to database]
5. production-api-tester: create_test_task.py
→ Create test subscription
6. production-api-tester: execute_batch.py
→ Run research
7. production-api-tester: wait_for_completion.py
→ Get results
8. production-api-tester: link_to_langfuse.py
→ Find trace
9. strategy-builder: analyze_strategy_performance.py
→ Analyze performance
10. [Iterate: Apply fixes and repeat]
Performance Deep Dive:
1. production-api-tester: execute_batch.py
→ Generate new traces
2. langfuse-optimization: retrieve_traces_and_observations.py
→ Get detailed trace data
3. langfuse-optimization: [analyze and fix configs]
→ Optimize style.yaml, template.yaml, tools.yaml
Targeted Analysis:
1. production-api-tester: Create multiple test tasks
2. production-api-tester: Execute batch
3. langfuse-advanced-filters: query_with_filters.py
→ Filter by specific criteria (e.g., latency > 10s)
4. langfuse-advanced-filters: analyze_filtered_results.py
→ Identify patterns
# Loop for quick iterations
for i in {1..5}; do
echo "Iteration $i"
# Modify strategy (manual or automated)
# Test
python3 create_test_task.py ... --output /tmp/test_$i/task.json
TASK_ID=$(jq -r '.id' /tmp/test_$i/task.json)
python3 execute_batch.py ...
python3 wait_for_completion.py --task-id "$TASK_ID" --output /tmp/test_$i/results.json
# Analyze
python3 link_to_langfuse.py --task-id "$TASK_ID" --output /tmp/test_$i/trace.json
# Cleanup
python3 delete_task.py --task-id "$TASK_ID"
echo "Iteration $i complete. Review /tmp/test_$i/"
sleep 5
done
# Test strategy A
python3 create_test_task.py --email "[email protected]" --topic "AI news" --output /tmp/ab_test/task_a.json
# Test strategy B (different strategy slug via topic classification)
python3 create_test_task.py --email "[email protected]" --topic "AI regulation detailed analysis" --output /tmp/ab_test/task_b.json
# Execute both
python3 execute_batch.py --frequency daily
# Compare results
python3 compare_tasks.py \
--task-a /tmp/ab_test/task_a.json \
--task-b /tmp/ab_test/task_b.json \
--output /tmp/ab_test/comparison.json
# Before deploying changes to production, test current vs new
# 1. Baseline (current production strategy)
python3 create_test_task.py --email "[email protected]" --topic "Test topic" --output /tmp/regression/baseline_task.json
python3 execute_batch.py ...
# Save results
# 2. Make changes to strategy
# 3. Test new version
python3 create_test_task.py --email "[email protected]" --topic "Test topic" --output /tmp/regression/new_task.json
python3 execute_batch.py ...
# Compare results
# 4. Validate no regressions
python3 validate_regression.py \
--baseline /tmp/regression/baseline_results.json \
--new /tmp/regression/new_results.json \
--output /tmp/regression/regression_report.json
Always use identifiable test email addresses:
test-strategy-{strategy_name}@example.comdev-{your_name}@example.comAlways delete test tasks after validation:
# List all test tasks
python3 list_tasks.py --api-key "$PROD_API_KEY" | grep "test-"
# Bulk delete
python3 batch_delete_tasks.py --pattern "test-*"
For quick tests without setting up infrastructure:
CALLBACK_URLWhen creating test tasks, use descriptive topics:
# Good
--topic "[TEST] AI news - strategy_v2_iteration_3"
# Bad
--topic "test"
This makes Langfuse traces easier to find and filter.
Create a script for the full optimization loop:
#!/bin/bash
# optimize_strategy.sh
STRATEGY_SLUG=$1
TEST_TOPIC=$2
# Generate → Validate → Deploy → Test → Analyze → Report
...
Production API may have rate limits:
--poll-interval to avoid overwhelming the API"Authentication failed":
PROD_API_KEY is set correctlyX-API-Key header is sent"Webhook not receiving results":
CALLBACK_URL is publicly accessible"Task execution times out":
--timeout parameter"Cannot find Langfuse trace":
--time-range "last_1_day" for safety"Health check fails":
PROD_API_URL is correctDO:
DON'T:
DO:
DON'T:
DO:
DON'T:
Good production testing should:
Remember: This skill is about safe production testing, not replacing proper staging environments. Use it for:
For high-risk changes, always test locally first using run_daily_briefing.py.
Project scaffolding, deployment configuration, and CI/CD setup for Google ADK agents.
Set up tracing, logging, and monitoring for deployed ADK agents across Cloud Trace, BigQuery, and third-party platforms.
Enterprise Azure infrastructure architect generating Bicep or Terraform from workload descriptions.
Plan and configure production-ready Azure Kubernetes Service clusters with Day-0 and Day-1 best practices.
Raw mechanical interfaces fusing Swiss typographic print with military terminal aesthetics. Rigid grids, extreme type scale contrast, utilitarian color, analog degradation effects. For data-heavy dashboards, portfolios, or editorial sites that need to feel like declassified blueprints.
Web search, scraping, extraction, crawling, and monitoring via ScrapeGraph AI CLI.
Skill for working with Firebase Hosting (Classic). Use this when you want to deploy static web apps, Single Page Apps (SPAs), or simple microservices. Do NOT use for Firebase App Hosting.
Deploy and manage web apps with Firebase App Hosting. Use this skill when deploying Next.js/Angular apps with backends.
Deploy applications and websites to Vercel. Use when the user requests deployment actions like "deploy my app", "deploy and give me the link", "push this live", or "create a preview deployment".
Build SEO-optimized pages at scale using templates, data, and proven playbook patterns.
Design and build isolated, reusable Convex backend components with clear boundaries and app-facing wrappers.
Deploy and manage projects on Vercel using token-based authentication. Use when working with Vercel CLI using access tokens rather than interactive login — e.g. "deploy to vercel", "set up vercel", "add environment variables to vercel".
Opus.ai: Revolutionize Your Web Experience
Build a no-code AI app in minutes.
An IDE for code migration from legacy to modern frameworks through coding agents.
Automate CGI animation in live-action scenes
Branded artistic QR-code concepts
Launch a website in seconds with AI.
Streamline Your Coding Experience with AI Code Helper
Convert any screenshot or design to clean code.