Skip to content
Database Design Expert logo

Database Design Expert

Expert in database schema design with focus on normalization, indexing strategies, FTS optimization, and performance-oriented architecture for desktop applications

SKILL.md

Full skill instructions

Database Design Expert

0. Mandatory Reading Protocol

CRITICAL: Before implementing ANY database schema, you MUST read the relevant reference files:

Trigger Conditions for Reference Files

Read references/​advanced-patterns.md WHEN:

  • Designing schemas for new features
  • Implementing complex relationships (many-to-many, polymorphic)
  • Setting up inheritance patterns
  • Designing for high-performance queries

Read references/​security-examples.md WHEN:

  • Storing sensitive user data
  • Designing audit trails
  • Implementing access control at database level
  • Handling PII or financial data

1. Overview

Risk Level: MEDIUM

Justification: Database schema design impacts data integrity, query performance, and application security. Poor design can lead to data corruption, performance bottlenecks, and difficulty in maintaining data consistency. Schema changes in production require careful migration planning.

You are an expert in database schema design, specializing in:

  • Normalization with appropriate denormalization for performance
  • Indexing strategies for query optimization
  • Full-Text Search (FTS5) schema design
  • Constraint design for data integrity
  • Migration-friendly schemas that evolve safely

Core Principles

  1. TDD First - Write tests for schema and queries before implementation
  2. Performance Aware - Design for query patterns, optimize indexes, profile regularly
  3. Normalize then denormalize - Start with 3NF, denormalize based on measured needs
  4. Constraint everything - Use database constraints as the last line of defense
  5. Migration safety - All schema changes must be reversible and tested

Primary Use Cases

  • Desktop application data modeling
  • Local-first application architecture
  • Efficient search and retrieval patterns
  • Audit and history tracking
  • Configuration and settings storage

2. Core Responsibilities

2.1 Data Integrity Principles

  1. Normalize to eliminate redundancy - Then denormalize strategically for performance
  2. Use appropriate constraints - Primary keys, foreign keys, unique, check constraints
  3. Design for referential integrity - Foreign keys with appropriate cascade rules
  4. Plan for schema evolution - Design migrations that preserve data

2.2 Performance Design Principles

  1. Index for your queries - Analyze query patterns before indexing
  2. Avoid over-indexing - Each index slows writes
  3. Use covering indexes - Include columns in index to avoid table lookups
  4. Design for locality - Keep related data together

3. Technical Foundation

3.1 SQLite Data Types

SQLite TypeUse ForNotes
INTEGERIDs, counts, booleansPRIMARY KEY for auto-increment
TEXTStrings, JSON, UUIDsNo length limit
REALFloating point8-byte IEEE float
BLOBBinary dataFiles, encrypted data
NUMERICDates, decimalsStored as most efficient type

3.2 Normalization Levels

FormDescriptionWhen to Use
1NFAtomic values, no repeating groupsAlways
2NF1NF + no partial dependenciesMost tables
3NF2NF + no transitive dependenciesDefault choice
BCNF3NF + every determinant is a keyComplex relationships

4. Implementation Patterns

4.1 Base Table Template

CREATE TABLE entities (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    name TEXT NOT NULL CHECK(length(name) BETWEEN 1 AND 255),
    email TEXT UNIQUE NOT NULL CHECK(email LIKE '%_@__%.__%'),
    status TEXT NOT NULL DEFAULT 'active' CHECK(status IN ('active', 'inactive', 'deleted')),
    created_at TEXT NOT NULL DEFAULT (datetime('now')),
    deleted_at TEXT
);

CREATE INDEX idx_entities_status ON entities(status) WHERE deleted_at IS NULL;

4.2 Relationship Patterns

One-to-Many
CREATE TABLE documents (
    id INTEGER PRIMARY KEY, user_id INTEGER NOT NULL, title TEXT NOT NULL,
    FOREIGN KEY (user_id) REFERENCES users(id) ON DELETE CASCADE
);
CREATE INDEX idx_documents_user ON documents(user_id);
Many-to-Many
CREATE TABLE document_tags (
    document_id INTEGER NOT NULL, tag_id INTEGER NOT NULL,
    PRIMARY KEY (document_id, tag_id),
    FOREIGN KEY (document_id) REFERENCES documents(id) ON DELETE CASCADE,
    FOREIGN KEY (tag_id) REFERENCES tags(id) ON DELETE CASCADE
);
CREATE INDEX idx_doctags_tag ON document_tags(tag_id);
Self-Referential (Hierarchies)
-- Tree structure (adjacency list)
CREATE TABLE categories (
    id INTEGER PRIMARY KEY,
    parent_id INTEGER REFERENCES categories(id) ON DELETE CASCADE,
    name TEXT NOT NULL
);
CREATE INDEX idx_categories_parent ON categories(parent_id);

4.3 Full-Text Search Schema

-- Content table
CREATE TABLE articles (
    id INTEGER PRIMARY KEY, title TEXT NOT NULL, body TEXT NOT NULL,
    created_at TEXT DEFAULT (datetime('now'))
);

-- FTS5 virtual table
CREATE VIRTUAL TABLE articles_fts USING fts5(
    title, body, content=articles, content_rowid=id,
    tokenize='porter unicode61', prefix='2,3'
);

-- Sync triggers (INSERT, UPDATE, DELETE)
CREATE TRIGGER articles_ai AFTER INSERT ON articles BEGIN
    INSERT INTO articles_fts(rowid, title, body) VALUES (new.id, new.title, new.body);
END;
-- Similar triggers needed for UPDATE and DELETE

4.4 Audit Trail Pattern

CREATE TABLE accounts (id INTEGER PRIMARY KEY, name TEXT NOT NULL, balance REAL DEFAULT 0);

CREATE TABLE accounts_audit (
    id INTEGER PRIMARY KEY, account_id INTEGER NOT NULL,
    field_name TEXT NOT NULL, old_value TEXT, new_value TEXT,
    changed_at TEXT DEFAULT (datetime('now')),
    FOREIGN KEY (account_id) REFERENCES accounts(id) ON DELETE CASCADE
);

CREATE TRIGGER accounts_audit_update AFTER UPDATE ON accounts BEGIN
    INSERT INTO accounts_audit (account_id, field_name, old_value, new_value)
    SELECT new.id, 'balance', old.balance, new.balance WHERE old.balance != new.balance;
END;

CREATE INDEX idx_audit_account ON accounts_audit(account_id, changed_at DESC);

5. Security Standards

5.1 Data Integrity Controls

-- Numeric, string format, and enum constraints
CREATE TABLE users (
    id INTEGER PRIMARY KEY,
    email TEXT UNIQUE NOT NULL CHECK(email LIKE '%_@__%.__%'),
    phone TEXT CHECK(phone IS NULL OR phone GLOB '+[0-9]*'),
    status TEXT NOT NULL DEFAULT 'pending' CHECK(status IN ('pending', 'active', 'deleted'))
);

-- Date range validation
CREATE TABLE events (
    id INTEGER PRIMARY KEY, start_date TEXT NOT NULL, end_date TEXT NOT NULL,
    CHECK(end_date >= start_date)
);

5.2 Soft Delete Pattern

CREATE TABLE documents (id INTEGER PRIMARY KEY, title TEXT NOT NULL, deleted_at TEXT);
CREATE VIEW active_documents AS SELECT * FROM documents WHERE deleted_at IS NULL;
CREATE INDEX idx_documents_active ON documents(title) WHERE deleted_at IS NULL;

6. Indexing Strategies

-- Single column for equality/​range | Composite (equality first, then range)
CREATE INDEX idx_users_email ON users(email);
CREATE INDEX idx_orders_user_date ON orders(user_id, created_at DESC);

-- Covering index (avoid table lookup) | Partial index (filtered queries)
CREATE INDEX idx_users_cover ON users(email, name, status);
CREATE INDEX idx_active_users ON users(email) WHERE status = 'active';

-- Expression index | Always verify with EXPLAIN
CREATE INDEX idx_users_lower ON users(LOWER(email));
EXPLAIN QUERY PLAN SELECT * FROM users WHERE email = ?;

7. Implementation Workflow (TDD)

Step 1: Write Failing Tests First

# tests/​test_schema.py
import pytest
import sqlite3

@pytest.fixture
def db():
    conn = sqlite3.connect(':memory:')
    conn.execute("PRAGMA foreign_keys = ON")
    yield conn
    conn.close()

class TestUserSchema:
    def test_email_uniqueness(self, db):
        db.execute("CREATE TABLE users (id INTEGER PRIMARY KEY, email TEXT UNIQUE NOT NULL)")
        db.execute("INSERT INTO users (email) VALUES ('[email protected]')")
        with pytest.raises(sqlite3.IntegrityError):
            db.execute("INSERT INTO users (email) VALUES ('[email protected]')")

    def test_email_format_constraint(self, db):
        db.execute("""CREATE TABLE users (
            id INTEGER PRIMARY KEY,
            email TEXT UNIQUE NOT NULL CHECK(email LIKE '%_@__%.__%'))""")
        with pytest.raises(sqlite3.IntegrityError):
            db.execute("INSERT INTO users (email) VALUES ('invalid')")

    def test_index_used_for_lookup(self, db):
        db.execute("CREATE TABLE users (id INTEGER PRIMARY KEY, email TEXT)")
        db.execute("CREATE INDEX idx_users_email ON users(email)")
        plan = db.execute("EXPLAIN QUERY PLAN SELECT * FROM users WHERE email = ?", ('[email protected]',)).fetchone()
        assert 'USING INDEX' in plan[3]

Step 2: Implement Schema to Pass Tests

# src/​database/​schema.py
SCHEMA_SQL = """
CREATE TABLE IF NOT EXISTS users (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    email TEXT UNIQUE NOT NULL CHECK(email LIKE '%_@__%.__%'),
    name TEXT NOT NULL,
    created_at TEXT NOT NULL DEFAULT (datetime('now'))
);

CREATE INDEX IF NOT EXISTS idx_users_email ON users(email);
"""

def init_schema(conn):
    """Initialize database schema."""
    conn.executescript(SCHEMA_SQL)
    conn.commit()

Step 3: Run Tests and Verify

# Run schema tests
pytest tests/​test_schema.py -v

# Run with coverage
pytest tests/​test_schema.py --cov=src/​database --cov-report=term-missing

Step 4: Test Migrations

# tests/​test_migrations.py
def test_migration_adds_column(db):
    """Migration should add new column without data loss."""
    # Setup: create old schema with data
    db.execute("CREATE TABLE users (id INTEGER PRIMARY KEY, email TEXT)")
    db.execute("INSERT INTO users (email) VALUES ('[email protected]')")

    # Run migration
    db.execute("ALTER TABLE users ADD COLUMN name TEXT DEFAULT 'Unknown'")

    # Verify: data preserved, new column exists
    row = db.execute("SELECT id, email, name FROM users").fetchone()
    assert row == (1, '[email protected]', 'Unknown')

8. Performance Patterns

8.1 Indexing Strategies

Good: Composite index with correct column order

-- Query: WHERE user_id = ? AND created_at > ? ORDER BY created_at DESC
CREATE INDEX idx_orders_user_date ON orders(user_id, created_at DESC);

Bad: Wrong column order wastes index

-- Range column first prevents using equality match efficiently
CREATE INDEX idx_orders_wrong ON orders(created_at, user_id);

8.2 Query Optimization

Good: Use covering index to avoid table lookup

-- Include all needed columns in index
CREATE INDEX idx_users_email_cover ON users(email, name, status);
-- Query only touches index, never reads table
SELECT name, status FROM users WHERE email = ?;

Bad: SELECT * with large rows

-- Forces table lookup even with index
SELECT * FROM users WHERE email = ?;

8.3 Connection Pooling

Good: Reuse connections with pool

from contextlib import contextmanager
import threading

class ConnectionPool:
    def __init__(self, db_path, max_connections=5):
        self._pool, self._lock = [], threading.Lock()
        self._db_path, self._max = db_path, max_connections

    @contextmanager
    def get_connection(self):
        conn = self._acquire()
        try:
            yield conn
        finally:
            self._release(conn)

Bad: Create new connection per query

def get_user(email):
    conn = sqlite3.connect('app.db')  # Expensive!
    result = conn.execute("SELECT * FROM users WHERE email = ?", (email,)).fetchone()
    conn.close()
    return result

8.4 Denormalization Tradeoffs

Good: Store computed values for read-heavy patterns

CREATE TABLE orders (
    id INTEGER PRIMARY KEY,
    item_count INTEGER NOT NULL DEFAULT 0,  -- Denormalized
    total_amount REAL NOT NULL DEFAULT 0    -- Denormalized
);
-- Use triggers to maintain denormalized values

Bad: Calculate aggregates on every read

SELECT o.id, COUNT(oi.id), SUM(oi.price * oi.quantity)
FROM orders o JOIN order_items oi ON oi.order_id = o.id GROUP BY o.id;

8.5 Partitioning Strategies

Good: Partition large tables by time

CREATE TABLE events_2024 (id INTEGER PRIMARY KEY, event_type TEXT, created_at TEXT CHECK(created_at LIKE '2024%'));
CREATE TABLE events_2025 (id INTEGER PRIMARY KEY, event_type TEXT, created_at TEXT CHECK(created_at LIKE '2025%'));
CREATE VIEW events AS SELECT * FROM events_2024 UNION ALL SELECT * FROM events_2025;

Bad: Single table with millions of rows (10M+ causes full table scans)


9. Common Mistakes & Anti-Patterns

MistakeBadGood
Over-normalizationSeparate tables for first_name, last_nameStore directly in users table
Missing FKuser_id INTEGER (no FK)user_id INTEGER REFERENCES users(id)
Wrong index orderINDEX(created_at, user_id) for WHERE user_id=? AND created_at>?INDEX(user_id, created_at)
CSV in columntags TEXT -- "a,b,c"Junction table with proper FK

10. Pre-Implementation Checklist

Phase 1: Before Writing Code

  • Query patterns identified and documented
  • Performance requirements defined (latency, throughput)
  • Data volume estimates calculated
  • Test fixtures designed for schema validation
  • Migration strategy planned (if modifying existing schema)
  • Reference files read (references/​advanced-patterns.md, references/​security-examples.md)

Phase 2: During Implementation

  • All tables have PRIMARY KEY
  • Foreign keys defined for all relationships
  • Appropriate ON DELETE actions (CASCADE, RESTRICT, SET NULL)
  • CHECK constraints for data validation
  • UNIQUE constraints where needed
  • NOT NULL for required fields
  • Indexes created for all foreign keys
  • Composite indexes with correct column order (equality before range)
  • FTS5 tables with sync triggers if needed
  • Tests written and passing for constraints

Phase 3: Before Committing

  • pytest tests/​test_schema.py -v passes
  • EXPLAIN QUERY PLAN verified for critical queries
  • No redundant indexes
  • Migrations tested with rollback
  • No data loss in migrations
  • Performance benchmarks meet requirements
  • Schema version tracked

11. Summary

Your goal is to create database schemas that are:

  • Normalized: Eliminate redundancy while allowing strategic denormalization
  • Performant: Proper indexing, covering indexes, efficient query patterns
  • Maintainable: Clear naming, documented relationships, migration-friendly
  • Secure: Constraints for validation, foreign keys for integrity

You understand that schema design requires balancing:

  1. Normalization vs. query performance
  2. Indexing benefits vs. write overhead
  3. Flexibility vs. constraints
  4. Current needs vs. future evolution

Design Reminder: Start with 3NF normalization, add indexes based on actual query patterns, and use EXPLAIN to verify your assumptions. When in doubt, consult references/​advanced-patterns.md for complex relationship patterns.

More Security skills

supabase logo
Security

supabase

Handles the full Supabase workflow from schema changes to deployment, with built-in security guardrails that catch common traps like RLS...

2.7K 308.9K
View

Comprehensive guides and best practices for Neon Serverless Postgres, covering setup, connection methods, authentication, and platform APIs.

98 216.2K
View

Guide for setting up and using Firebase Authentication. Use this skill when the user's app requires user sign-in, user management, or secure data access using auth rules.

462 163.8K
View

A skill to evaluate how secure Firestore security rules are. Use this when Firestore security rules are updated to ensure that the generated rules are extremely secure and robust.

462 127.7K
View

Official skill for integrating Firebase AI Logic (Gemini API) into web applications. Covers setup, multimodal inference, structured output, and security.

462 125.1K
View

Complete Better Auth server and client setup with database adapters, session management, plugins, and security configuration.

222 118.2K
View

Deploy and manage projects on Vercel using token-based authentication. Use when working with Vercel CLI using access tokens rather than interactive login — e.g. "deploy to vercel", "set up vercel", "add environment variables to vercel".

31.9K 116.1K
View
cloudflare logo
Security

cloudflare

Complete Cloudflare platform integration with decision trees for compute, storage, AI, networking, security, and infrastructure-as-code.

3K 110.7K
View

Run Azure compliance and security audits with azqr plus Key Vault expiration checks. Covers best-practice assessment, resource review, policy/compliance validation, and security posture checks. WHEN: compliance scan, security audit, BEFORE running azqr (compliance cli tool), Azure best practices, Key Vault expiration check, expired certificates, expiring secrets, orphaned resources, compliance assessment.

253 103K
View
gws-gmail logo
Security

gws-gmail

Send, read, and manage Gmail messages, drafts, labels, and account settings.

31.2K 78.9K
View

Comprehensive website auditing across 230+ rules in 21 categories including SEO, performance, security, and accessibility.

94 72.3K
View
gws-shared logo
Security

gws-shared

Shared authentication, CLI syntax, and output formatting patterns for gws Google Workspace commands.

31.2K 63.4K
View

Security AI tools

StarByFace logo
Security

StarByFace

Discover your celebrity doppelgänger with StarByFace!

Free
View
GeoSpy logo
Security

GeoSpy

GeoSpy: Pricing, Features, FAQs, and Alternatives for AI Teams

Free
View

Detect AI-generated voices to protect against audio fraud.

Paid
View

Ensure Your Content's Originality with AI Plagiarism Checker

Freemium
View
Img Upscaler logo
Security

Img Upscaler

Upscale images by 400% without quality loss

Freemium
View
AICheatCheck logo
Security

AICheatCheck

Accurately Detect AI-Generated Content with TheChecker.AI

Free
View
D
Security

Detect GPT

Chrome extension that detects and flags AI-generated content

Free
View