environment-setup-guide
Guide developers through setting up development environments with proper tools, dependencies, and configurations
databricks-prod-checklist
Execute Databricks production deployment checklist and rollback procedures. Use when deploying Databricks jobs to production, preparing for launch, or implementing go-live procedures. Trigger with phrases like "databricks production", "deploy databricks", "databricks go-live", "databricks launch ...
Full skill instructions
Complete checklist for deploying Databricks jobs and pipelines to production.
# Run tests via Asset Bundles
databricks bundle validate -t prod
databricks bundle run -t staging test-job
# Verify test results
databricks runs get --run-id $RUN_ID | jq '.state.result_state'
collect() on large datasets# resources/prod_job.yml
resources:
jobs:
etl_pipeline:
name: "prod-etl-pipeline"
tags:
environment: production
team: data-engineering
cost_center: analytics
schedule:
quartz_cron_expression: "0 0 6 * * ?"
timezone_id: "America/New_York"
email_notifications:
on_failure:
- "[email protected]"
on_success:
- "[email protected]"
webhook_notifications:
on_failure:
- id: "slack-webhook-id"
max_concurrent_runs: 1
timeout_seconds: 14400 # 4 hours
tasks:
- task_key: bronze_ingest
job_cluster_key: etl_cluster
notebook_task:
notebook_path: /Repos/prod/pipelines/bronze
timeout_seconds: 3600
- task_key: silver_transform
depends_on:
- task_key: bronze_ingest
job_cluster_key: etl_cluster
notebook_task:
notebook_path: /Repos/prod/pipelines/silver
job_clusters:
- job_cluster_key: etl_cluster
new_cluster:
spark_version: "14.3.x-scala2.12"
node_type_id: "Standard_DS3_v2"
num_workers: 4
autoscale:
min_workers: 2
max_workers: 8
spark_conf:
spark.sql.shuffle.partitions: "200"
spark.databricks.delta.optimizeWrite.enabled: "true"
instance_pool_id: "prod-pool-id"
# Pre-flight checks
echo "=== Pre-flight Checks ==="
databricks workspace list /Repos/prod/ # Verify repo exists
databricks clusters list | grep prod # Verify pools/clusters
databricks secrets list-scopes # Verify secrets
# Deploy with Asset Bundles
echo "=== Deploying ==="
databricks bundle deploy -t prod
# Verify deployment
databricks bundle summary -t prod
databricks jobs list | grep prod-etl
# Manual trigger to verify
echo "=== Verification Run ==="
RUN_ID=$(databricks jobs run-now --job-id $JOB_ID | jq -r '.run_id')
echo "Run ID: $RUN_ID"
# Monitor run
databricks runs get --run-id $RUN_ID --wait
# monitoring/health_check.py
from databricks.sdk import WorkspaceClient
from datetime import datetime, timedelta
def check_job_health(w: WorkspaceClient, job_id: int) -> dict:
"""Check job health metrics."""
# Get recent runs
runs = list(w.jobs.list_runs(
job_id=job_id,
completed_only=True,
limit=10,
))
if not runs:
return {"status": "NO_RUNS", "healthy": False}
# Calculate success rate
successful = sum(1 for r in runs if r.state.result_state == "SUCCESS")
success_rate = successful / len(runs)
# Calculate average duration
durations = [
(r.end_time - r.start_time) / 1000 / 60 # minutes
for r in runs if r.end_time
]
avg_duration = sum(durations) / len(durations) if durations else 0
# Check last run
last_run = runs[0]
last_state = last_run.state.result_state
return {
"status": "HEALTHY" if success_rate > 0.9 else "DEGRADED",
"healthy": success_rate > 0.9 and last_state == "SUCCESS",
"success_rate": success_rate,
"avg_duration_minutes": avg_duration,
"last_run_state": last_state,
"last_run_time": datetime.fromtimestamp(last_run.start_time / 1000),
}
#!/bin/bash
# rollback.sh - Emergency rollback procedure
JOB_ID=$1
PREVIOUS_VERSION=$2
echo "=== ROLLBACK INITIATED ==="
echo "Job: $JOB_ID"
echo "Target Version: $PREVIOUS_VERSION"
# 1. Pause the job
echo "Pausing job..."
databricks jobs update --job-id $JOB_ID --json '{"settings": {"schedule": null}}'
# 2. Cancel active runs
echo "Cancelling active runs..."
databricks runs list --job-id $JOB_ID --active-only | \
jq -r '.runs[].run_id' | \
xargs -I {} databricks runs cancel --run-id {}
# 3. Reset to previous version
echo "Rolling back to version $PREVIOUS_VERSION..."
databricks bundle deploy -t prod --force
# 4. Re-enable schedule
echo "Re-enabling schedule..."
# (restore from backup config)
# 5. Trigger verification run
echo "Triggering verification run..."
databricks jobs run-now --job-id $JOB_ID
echo "=== ROLLBACK COMPLETE ==="
| Alert | Condition | Severity |
|---|---|---|
| Job Failed | result_state = FAILED | P1 |
| Long Running | Duration > 2x average | P2 |
| Consecutive Failures | 3+ failures in a row | P1 |
| Data Quality | Expectations failed | P2 |
-- Job health metrics (Unity Catalog system tables)
SELECT
job_id,
job_name,
COUNT(*) as total_runs,
SUM(CASE WHEN result_state = 'SUCCESS' THEN 1 ELSE 0 END) as successes,
AVG(execution_duration) / 60000 as avg_minutes,
MAX(start_time) as last_run
FROM system.lakeflow.job_run_timeline
WHERE start_time > current_timestamp() - INTERVAL 7 DAYS
GROUP BY job_id, job_name
ORDER BY total_runs DESC
# Comprehensive pre-prod check
databricks bundle validate -t prod && \
databricks bundle deploy -t prod --dry-run && \
echo "Validation passed, ready to deploy"
For version upgrades, see databricks-upgrade-migration.
Guide developers through setting up development environments with proper tools, dependencies, and configurations
Implement Exa lint rules, policy enforcement, and automated guardrails. Use when setting up code quality rules for Exa integrations, implementing pre-commit hooks, or configuring CI policy checks for Exa best practices. Trigger with phrases like "exa policy", "exa lint", "exa guardrails", "exa be...
Configure CI/CD pipelines for Documenso integrations. Use when setting up automated testing, deployment pipelines, or continuous integration for Documenso projects. Trigger with phrases like "documenso CI", "documenso GitHub Actions", "documenso pipeline", "documenso automated testing".
Build check model availability and implement fallback chains. Use when building resilient systems or handling model outages. Trigger with phrases like 'openrouter availability', 'openrouter fallback', 'openrouter model down', 'openrouter health check'.
Configure Langfuse enterprise organization management and access control. Use when implementing team access controls, configuring organization settings, or setting up role-based permissions for Langfuse projects. Trigger with phrases like "langfuse RBAC", "langfuse teams", "langfuse organization"...
Execute compliance and security auditing for Cursor usage. Triggers on "cursor compliance", "cursor audit", "cursor security review", "cursor soc2", "cursor gdpr". Use when analyzing or auditing cursor compliance audit. Trigger with phrases like "cursor compliance audit", "cursor audit", "cursor".
Configure Apollo.io CI/CD integration. Use when setting up automated testing, continuous integration, or deployment pipelines for Apollo integrations. Trigger with phrases like "apollo ci", "apollo github actions", "apollo pipeline", "apollo ci/cd", "apollo automated tests".
Create robust Python automation with full logging and safety checks. Use when tasks need complex data processing, authenticated API work, conditional file operations, or error handling beyond simple shell commands.
Configure PostHog across development, staging, and production environments. Use when setting up multi-environment deployments, configuring per-environment secrets, or implementing environment-specific PostHog configurations. Trigger with phrases like "posthog environments", "posthog staging", "po...
Expert in designing and building autonomous AI agents. Masters tool use, memory systems, planning strategies, and multi-agent orchestration. Use when: build agent, AI agent, autonomous agent, tool use, function calling.
Validate insecure deserialization checker operations. Auto-activating skill for Security Fundamentals. Triggers on: insecure deserialization checker, insecure deserialization checker Part of the Security Fundamentals skill category. Use when working with insecure deserialization checker functiona...
Manage path traversal finder operations. Auto-activating skill for Security Fundamentals. Triggers on: path traversal finder, path traversal finder Part of the Security Fundamentals skill category. Use when working with path traversal finder functionality. Trigger with phrases like "path traversa...
Generate, edit, and beat-sync AI video with leading models in one workspace.
The world's fastest calendar for remote work
Transform Your Design with AI Designer by ImgCreator.ai
Revolutionizing Video Production with AI-Powered Creativity
Extend an image past the frame and let AI fill the new aspect ratio.
Discover your celebrity doppelgänger with StarByFace!
ChainClarity explains 700+ crypto whitepapers in plain English, with layered summaries, comparisons, research tools, alerts, and a $4.99 Pro plan.
Opus.ai: Revolutionize Your Web Experience