EDA Tabular
eda-tabular
Structured exploratory data analysis for tabular datasets. Use when Codex needs to inspect CSV, parquet, pandas DataFrame, warehouse extracts, or Kaggle-style train/test data to understand schema, target behavior, missingness, drift, leakage risk, feature quality, and model-selection implications...
SKILL.md
Full skill instructions
EDA Tabular
Perform EDA to reduce modeling uncertainty, not to generate charts for their own sake. End each EDA cycle with concrete decisions about validation, model family, feature strategy, and data risks.
Workflow
- Verify the data contract.
- Inspect target behavior and leakage risk.
- Inspect feature quality and train/test drift.
- Run one small model-response benchmark.
- Convert findings into modeling decisions.
- Save concise human-readable artifacts.
Verify The Data Contract
- Confirm train, test, and sample-submission row counts.
- Identify the ID column and target column.
- Confirm the prediction schema and metric.
- Check duplicated IDs, duplicated rows, and train/test schema mismatches.
- If the split is time-based, grouped, or otherwise constrained, adapt validation before deeper EDA.
Inspect Target And Leakage
- Measure class balance or target distribution first.
- Compare target rate across a few high-value categorical columns.
- Look for features or combinations that nearly determine the target.
- Flag any preprocessing idea that would leak target or test information into validation.
Inspect Feature Quality
- Check missingness by column and by row.
- Check numeric ranges, impossible values, heavy skew, and outliers.
- Check categorical cardinality and rare categories.
- Check constant or near-constant features.
- Compare simple train/test distributions to detect obvious drift.
Run Model-Response EDA
Use a cheap benchmark to test what kind of signal exists before writing a lot of feature code.
- Train one linear or additive baseline.
- Train one shallow tree baseline.
- Sweep a small tree-depth range when the task is tabular supervised learning.
- If shallow depth wins, bias toward additive or weak-interaction models.
- If deeper trees clearly win, interaction engineering is more justified.
Turn Findings Into Decisions
- Decide validation strategy from the data contract, not from convenience.
- Keep only transformations that match the observed signal.
- Prefer fold-safe preprocessing for every reported score.
- Promote the strongest simple baseline to first-class status.
- Treat blending as optional and evidence-driven.
Output Standard
Leave behind:
- a short Markdown report
- a machine-readable summary if code is involved
- a minimal chart set for humans
- explicit conclusions that change what happens next
Reference
Read references/checklist.md when you need the fuller checklist, anti-patterns, or output expectations.
