Skip to content
data-layer-architecture logo

Data Layer Architecture(Bronze Layer / Gold Layer)

data-layer-architecture

Bronze Layer(LLM抽出ログ層)とGold Layer(確定データ層)の2層アーキテクチャ設計。LLM抽出結果の履歴管理と人間修正の保護を実現。抽出処理の実装、ExtractionLogの使用、is_manually_verifiedフラグの扱いに関するガイダンスを提供。

SKILL.md

Full skill instructions

Data Layer Architecture(Bronze Layer / Gold Layer)

Purpose

LLM抽出結果と確定データを分離する2層アーキテクチャの設計ガイド。AIの抽出履歴を保持しながら、人間の修正を保護する。

When to Activate

このスキルは以下の場合にアクティベートされます:

  • LLM抽出処理を新規実装する時
  • ExtractionLogエンティティを使用する時
  • is_manually_verifiedフラグを扱う時
  • 抽出結果からGoldエンティティを更新する時
  • Bronze/​Gold Layerの設計について質問された時

Architecture Overview

┌─────────────────────────────────────────────────────────────┐
│                    Bronze Layer(抽出ログ層)                 │
│                                                              │
│  - LLM抽出結果を追記専用(Immutable)で保存                   │
│  - 精度分析・トレーサビリティのための履歴                      │
│  - テーブル: extraction_logs                                 │
└─────────────────────────────────────────────────────────────┘
                              │
                              │ is_manually_verified = false の場合のみ反映
                              ▼
┌─────────────────────────────────────────────────────────────┐
│                    Gold Layer(確定データ層)                  │
│                                                              │
│  - アプリケーションが参照する唯一の正解データ                   │
│  - 人間の修正が優先される                                     │
│  - テーブル: statements, politicians, speakers, etc.         │
└─────────────────────────────────────────────────────────────┘

Design Principles

1. Bronze Layer(抽出ログ層)

  • 目的: LLM抽出結果の履歴保持、精度分析
  • 特性: 追記専用(Immutable)、削除・更新なし
  • 用途:
    • AIモデル改善の検証データ
    • 抽出精度の時系列分析
    • デバッグ・トラブルシューティング

2. Gold Layer(確定データ層)

  • 目的: ユーザーに提供する正解データ
  • 特性: 人間の修正が最優先
  • 用途:
    • Streamlit UIでの表示
    • API経由でのデータ提供
    • レポート・分析の基礎データ

Data Flow

LLM抽出実行
    │
    ▼
┌────────────────────────────┐
│ ExtractionLog に必ず保存    │  ← Bronze Layer(常に履歴として残る)
│ (Immutable)               │
└────────────────────────────┘
    │
    ▼
┌────────────────────────────┐
│ is_manually_verified?      │
└────────────────────────────┘
    │                │
    │ false          │ true
    │(未検証)       │(検証済み)
    ▼                ▼
┌──────────┐    ┌──────────────────┐
│ Gold更新  │    │ Gold更新しない    │
│          │    │(人間の修正を保護)│
└──────────┘    └──────────────────┘

Target Entities

抽出オブジェクト(Bronze)→確定オブジェクト(Gold)
StatementExtraction→Statement(発言)
PoliticianExtraction→Politician(政治家)
SpeakerExtraction→Speaker(話者)
ConferenceMemberExtraction→ConferenceMember(会議体メンバー)
ParliamentaryGroupMemberExtraction→ParliamentaryGroupMember(議員団メンバー)

Key Components

ExtractionLog Entity

class EntityType(Enum):
    STATEMENT = "statement"
    POLITICIAN = "politician"
    SPEAKER = "speaker"
    CONFERENCE_MEMBER = "conference_member"
    PARLIAMENTARY_GROUP_MEMBER = "parliamentary_group_member"

@dataclass
class ExtractionLog:
    id: UUID
    entity_type: EntityType          # どのエンティティの抽出か
    entity_id: UUID                  # 対象GoldエンティティのID
    pipeline_version: str            # "gemini-2.0-flash-v1" など
    extracted_data: dict             # AIが出した生データ(JSON)
    confidence_score: Optional[float]
    extraction_metadata: dict        # モデル名、トークン数等
    created_at: datetime             # Immutable

Gold Entity Common Fields

# 全Goldエンティティ(Statement, Politician等)に追加
is_manually_verified: bool = False      # 人間が検証済みか
latest_extraction_log_id: Optional[UUID] # 最新の抽出ログへの参照

Behavior Rules

状態AI再抽出時の動作理由
is_manually_verified = falseGold Layerを更新最新AIの精度向上を反映
is_manually_verified = trueGold Layerは更新しない人間の判断を最優先
(両方)Bronze Layerには常に保存履歴・分析用

Implementation Guide

Adding New Extraction Process

  1. UpdateEntityFromExtractionUseCase を継承して実装
  2. 抽出結果は必ず ExtractionLog に保存
  3. is_manually_verified フラグをチェックしてから Gold を更新
class UpdateStatementFromExtractionUseCase(UpdateEntityFromExtractionUseCase):
    async def execute(
        self,
        entity_id: UUID,
        extraction_result: StatementExtractionResult,
        pipeline_version: str
    ) -> UpdateEntityResult:
        # 1. Bronze Layer: 抽出ログを必ず保存
        log = ExtractionLog(
            entity_type=EntityType.STATEMENT,
            entity_id=entity_id,
            pipeline_version=pipeline_version,
            extracted_data=extraction_result.to_dict(),
            ...
        )
        log_id = await self._extraction_log_repo.save(log)

        # 2. Gold Layer: 人間修正済みならスキップ
        entity = await self._statement_repo.find_by_id(entity_id)
        if entity.is_manually_verified:
            return UpdateEntityResult(updated=False, reason="manually_verified")

        # 3. Gold Layer: 未検証なら更新
        entity.update_from_extraction(extraction_result)
        entity.latest_extraction_log_id = log_id
        await self._statement_repo.save(entity)

        return UpdateEntityResult(updated=True)

Handling Human Modifications

  1. Gold エンティティを直接更新
  2. is_manually_verified = true をセット
  3. 以降のAI再抽出では上書きされない
async def mark_as_verified(entity_id: UUID) -> None:
    entity = await repo.find_by_id(entity_id)
    entity.is_manually_verified = True
    await repo.save(entity)

Quick Checklist

新しい抽出処理を実装する際のチェックリスト:

  • ExtractionLog にログを保存しているか
  • is_manually_verified をチェックしているか
  • latest_extraction_log_id を更新しているか
  • pipeline_version を適切に設定しているか
  • エラー時もログが保存されるか
  • 単体テストを書いたか

Related Issues

  • Product Goal: #813
  • Implementation PBIs: #861〜#872

References

More skills from majiayu000

xiaohongshu logo
majiayu000/claude-arsenal

xiaohongshu

xiaohongshu

286 148
View
agent-task-conductor logo
majiayu000/claude-skill-registry

agent-task-conductor

Conduct multi-agent task orchestration and workflow coordination.

663 1
View
conductor-setup logo
majiayu000/claude-skill-registry

conductor-setup

Initialize project with Conductor artifacts (product definition,

663 1
View
animation-designer logo
majiayu000/claude-skill-registry

animation-designer

Expert in web animations, transitions, and motion design using Framer Motion and CSS

663 1
View
diagramming logo
majiayu000/claude-skill-registry

diagramming

Creates Mermaid and ASCII diagrams for flowcharts, architecture, ERDs, state machines, mindmaps, and more. Use when user mentions diagram, flowchart, mermaid, ASCII diagram, text diagram, terminal diagram, visualize, C4, mindmap, architecture diagram, sequence diagram, ERD, or needs visual docume...

663 1
View
h3-pg logo
majiayu000/claude-skill-registry-data

h3-pg

PostgreSQL bindings for H3 hexagonal grid system. Use when working with H3 cells in Postgres, including spatial indexing, geometry/geography integration, and raster analysis.

23 1
View
conductor-development logo
majiayu000/claude-skill-registry

conductor-development

Context-Driven Development skill for projects using Conductor. Use this skill when you detect a `conductor/` directory in the project, when working on tasks defined in a `plan.md` file, or when the user asks about tracks, specs, or plans. Automatically applies TDD workflow, tracks task completion...

663 1
View
conductor-status logo
majiayu000/claude-skill-registry

conductor-status

Display project status, active tracks, and next actions

663 1
View
dockerization logo
majiayu000/claude-skill-registry

dockerization

Official Stakpak application containerization standard operating procedure, a step-by-step guidline to properly dockerize applications. This is a rule book curated by the Stakpak Team.

663 1
View

Popular AI tools

Kaiber logo
Video

Kaiber

Generate, edit, and beat-sync AI video with leading models in one workspace.

Paid
View
Vimcal logo
Productivity

Vimcal

The world's fastest calendar for remote work

Free
View

Transform Your Design with AI Designer by ImgCreator.ai

Freemium
View
Akool AI logo
Content & writing

Akool AI

Revolutionizing Video Production with AI-Powered Creativity

Paid
View

Extend an image past the frame and let AI fill the new aspect ratio.

Freemium
View
StarByFace logo
Security

StarByFace

Discover your celebrity doppelgänger with StarByFace!

Free
View
C

ChainClarity explains 700+ crypto whitepapers in plain English, with layered summaries, comparisons, research tools, alerts, and a $4.99 Pro plan.

Freemium
View
Opus Clip logo
Coding & apps

Opus Clip

Opus.ai: Revolutionize Your Web Experience

Free
View