dApp Docs/白皮书知识图谱RAG接入指南
Development reference. Not independently verified for production.

MSG Chain 白皮书知识图谱 RAG 接入指南

数据来源:MSG Chain 代码库核实

主网状态: No-Go — 当前 MSGChain 主网裁决为 No-Go,以下内容反映代码实际状态,不代表生产可用。

面向 AI Agent 的机器可读白皮书系统 — 完整接入手册


版本: 1.0.0
适用对象: AI Agent 开发者、LLM 应用集成者、智能合约开发者、dApp 构建者
核心原则: 所有知识图谱内容均来自 MSG Chain 官方机器可读白皮书系统,任何 RAG 响应必须包含源状态标记
知识截止: 本文档包含截至 MSG Chain 白皮书 v1.0 的全部接入规范


目录

  1. 概述
  2. 白皮书系统架构
  3. 数据采集与爬取策略
  4. RAG 索引策略
  5. 检索提示与路由
  6. AI Agent 引导提示词
  7. FAQ 机器人实现
  8. Telegram Bot 实现
  9. 知识网络浏览器
  10. 维护与更新
  11. 边界与准则
  12. 附录

1. 概述

1.1 什么是 MSG Chain 白皮书知识图谱 RAG 系统

MSG Chain 白皮书知识图谱 RAG(Retrieval-Augmented Generation)系统是一套为 AI Agent 设计的机器可读知识基础设施。与传统 PDF 白皮书不同,MSG Chain 将整个白皮书拆解为 59 个独立模块,每个模块具有:

1.2 为什么需要 RAG

传统 LLM 在处理 MSG Chain 白皮书时面临以下挑战:

挑战 解决方案
上下文窗口限制(59 模块远超上下文长度) 分块 + 向量检索,仅注入相关片段
知识更新滞后(白皮书迭代后 LLM 权重未更新) 实时爬取 + 动态索引
幻觉(LLM 可能虚构 API 规范) 检索强制绑定真实 JSON 规范
状态混淆(无法区分已实现与规划中功能) 每条响应携带 [Status: implemented/partial/planned] 标记
无法路由专业问题 retrieval_hints.json 提供 8 个主题分类与推荐模块

1.3 系统概览

维度 数值
模块总数 59
状态层级 3(implemented / partial / planned)
功能分组 6 组
主题分类 8 类
链 ID msg-chain-1
Bech32 前缀 msg
入口端点 agent_entry.json / developer_entry.json

2. 白皮书系统架构

2.1 系统拓扑

MSG Chain 机器可读白皮书系统采用星形拓扑结构,以 agent_entry.json 为单一入口,辐射到所有模块、API 规范、集成示例和执行包。

agent_entry.json
├── knowledge_network.json          # 知识图谱(模块关系网络)
├── module_exports/index.json       # 模块内容索引
├── module_chunks/index.json        # 段落分块索引
├── module_exports/{module}.json    # 59 个独立模块内容
├── module_chunks/{module}__chunk_*.json  # 段落级分块
├── retrieval_hints.json            # 检索提示与路由
├── api_specs/                       # API 规范
│   ├── rpc_methods.json
│   ├── error_codes.json
│   ├── formal_contracts.json
│   ├── public_query.yaml
│   ├── contract_surface.yaml
│   └── agent_surface.yaml
├── quickstart/                      # 快速开始
│   ├── index.json
│   └── contract_and_dapp_minimal.json
├── chain_config/                    # 链配置
│   ├── index.json
│   └── network_presets.json
├── contract_templates/index.json   # 合约模板
├── examples/index.json             # dApp 示例
├── execution_pack/                  # 执行包
│   ├── command_registry.json
│   ├── ci_cd_templates.json
│   ├── delivery_workflows.json
│   ├── approval_gates.json
│   └── evidence_requirements.json
└── integration_examples/           # 集成示例
    ├── telegram_bot_crawl_flow.json
    ├── rag_ingest_flow.json
    ├── faq_router_prompt_template.md
    └── external_ai_agent_bootstrap_prompt.json

2.2 入口端点

AI Agent 入口

属性 值
URL https://msgchain.org/whitepaper/agent_entry.json
用途 AI Agent 主入口,包含所有资源的索引链接
内容 知识网络、模块导出、模块分块、API 规范、快速开始、集成示例、执行包、链配置的完整 URL

开发者入口

属性 值
URL https://msgchain.org/whitepaper/developer_entry.json
用途 面向人类开发者的友好入口,包含文档链接和教程

产品交付入口

属性 值
URL https://msgchain.org/whitepaper/product_delivery_entry.json
用途 自动化产品交付流水线入口,包含 CI/CD 集成信息

检索提示入口

属性 值
URL https://msgchain.org/whitepaper/retrieval_hints.json
用途 8 个主题分类的检索路由信息

2.3 模块结构

每个模块包含三级内容表示:

层级一:HTML 叙事页面

https://msgchain.org/whitepaper/modules/knowledge_network.html
https://msgchain.org/whitepaper/modules/knowledge_graph_dynamic.html

层级二:JSON 结构化导出

https://msgchain.org/whitepaper/module_exports/{module_stem}.json

每个 JSON 导出包含:

{
  "module": "consensus_mechanism",
  "title": "Consensus Mechanism",
  "status": "implemented",
  "group": "consensus_and_validator",
  "tags": ["consensus", "validator", "vrf", "pos", "block_production"],
  "content": {
    "summary": "...",
    "sections": [...],
    "specifications": {...},
    "code_examples": [...]
  },
  "outlinks": ["validator_selection", "block_production"],
  "backlinks": ["architecture_overview", "tokenomics"],
  "related": ["validator_economics", "staking"],
  "version": "1.0.0",
  "last_updated": "2025-12-01"
}

层级三:段落级分块

https://msgchain.org/whitepaper/module_chunks/{module_stem}__chunk_{no}.json

每个 chunk 约 500 tokens:

{
  "module": "consensus_mechanism",
  "chunk_id": "consensus_mechanism__chunk_001",
  "chunk_index": 1,
  "total_chunks": 8,
  "text": "...",
  "tags": ["consensus", "vrf"],
  "status": "implemented",
  "group": "consensus_and_validator",
  "section": "2.3 VRF-based Leader Election"
}

2.4 59 个模块目录

分组 1: Architecture & Core(架构与核心)

模块 状态
architecture_overview implemented
chain_config implemented
network_specifications implemented
node_architecture implemented
state_machine implemented
transaction_lifecycle implemented
block_production implemented
cross_chain_bridge partial

分组 2: Consensus & Validator(共识与验证者)

模块 状态
consensus_mechanism implemented
validator_selection implemented
validator_economics implemented
staking implemented
slashing_conditions implemented
delegator_mechanics implemented
validator_hardware_requirements implemented
network_upgrades partial

分组 3: AI Agent Runtime(AI Agent 运行时)

模块 状态
ai_agent_runtime_overview implemented
agent_registration implemented
agent_execution_environment implemented
agent_memory_system partial
agent_interaction_protocol partial
agent_economics implemented
agent_security partial
agent_governance planned
agent_dispute_resolution planned

分组 4: Token & Economics(代币与经济)

模块 状态
tokenomics implemented
msg_token_specification implemented
token_distribution implemented
inflation_schedule implemented
fee_market implemented
treasury implemented
treasury_governance partial
ecosystem_funding planned

分组 5: Smart Contract & dApp(智能合约与 dApp)

模块 状态
smart_contract_overview implemented
contract_development_guide implemented
contract_deployment implemented
contract_upgradeability partial
dapp_development_guide implemented
dapp_integration_patterns implemented
oracle_system partial
contract_standard_library implemented
contract_security_best_practices implemented
contract_testing_framework implemented
contract_auditing planned

分组 6: Explorer & Data Access(浏览器与数据访问)

模块 状态
explorer_overview implemented
block_explorer_api implemented
transaction_query_api implemented
account_query_api implemented
validator_query_api implemented
event_subscription partial
data_indexing_service partial
historical_data_archive planned
analytics_dashboard planned

2.5 状态分类

状态 含义 AI Agent 处理方式
implemented 代码已实现,功能可用 可安全引用,标注 [Status: implemented]
partial 部分实现,API 可能不稳定 引用时标注 [Status: partial],附带 "部分功能可能不可用" 提示
planned 规划中,尚未实现 仅作为未来路线图引用,标注 [Status: planned],明确声明 "此功能尚未实现"

2.6 知识图谱关系网络

knowledge_network.json 位于 https://msgchain.org/whitepaper/knowledge_network.json,是理解模块间关系的核心文件。包含 59 个模块的完整关系网络。

{
  "graph_metadata": {
    "total_modules": 59,
    "total_edges": 247,
    "groups": 6,
    "statuses": ["implemented", "partial", "planned"],
    "generated_at": "2025-12-01T00:00:00Z"
  },
  "modules": {
    "consensus_mechanism": {
      "group": "consensus_and_validator",
      "status": "implemented",
      "tags": ["consensus", "validator", "vrf", "pos", "block_production"],
      "outlinks": ["validator_selection", "block_production", "validator_economics"],
      "backlinks": ["architecture_overview", "tokenomics", "network_specifications"],
      "related": ["staking", "slashing_conditions"],
      "export_url": "https://msgchain.org/whitepaper/module_exports/consensus_mechanism.json",
      "chunks_url": "https://msgchain.org/whitepaper/module_chunks/consensus_mechanism__chunk_index.json",
      "html_url": "https://msgchain.org/whitepaper/modules/consensus_mechanism.html"
    }
  }
}

关系类型

关系字段 含义 用途
outlinks 当前模块引用的其他模块 回答问题 "该模块还涉及哪些主题?"
backlinks 引用当前模块的其他模块 回答问题 "哪些模块涉及这个话题?"
related 语义相关的其他模块 推荐 "您可能还想了解"

HTML 可视化

这些可视化页面提供交互式网络图,可用于人工浏览和调试 RAG 检索结果。


3. 数据采集与爬取策略

3.1 推荐爬取顺序

基于 agent_entry.json 的依赖关系和资源重要性,推荐以下爬取顺序:

步骤 1:  knowledge_network.json          — 获取模块图谱(知道有哪些模块)
步骤 2:  module_exports/index.json       — 获取模块内容索引
步骤 3:  module_chunks/index.json        — 获取段落分块索引
步骤 4:  retrieval_hints.json            — 获取检索路由信息
步骤 5:  agent_entry.json                — 获取完整入口(对其他资源有引用关系)
步骤 6:  api_specs/*                     — API 规范
步骤 7:  按需爬取单个模块导出与分块    — 实际内容
步骤 8:  execution_pack/*                 — 执行包
步骤 9:  integration_examples/*          — 集成示例
步骤 10: 其他辅助资源(quickstart, chain_config 等)

3.2 完整资源列表

入口 (4 个)

  1. https://msgchain.org/whitepaper/agent_entry.json
  2. https://msgchain.org/whitepaper/developer_entry.json
  3. https://msgchain.org/whitepaper/product_delivery_entry.json
  4. https://msgchain.org/whitepaper/retrieval_hints.json

知识图谱 (3 个)

  1. https://msgchain.org/whitepaper/knowledge_network.json
  2. https://msgchain.org/whitepaper/modules/knowledge_network.html
  3. https://msgchain.org/whitepaper/modules/knowledge_graph_dynamic.html

模块导出 (最多 59 + 3 个)

模块分块 (N 个)

API 规范 (6 个)

  1. https://msgchain.org/whitepaper/api_specs/rpc_methods.json
  2. https://msgchain.org/whitepaper/api_specs/error_codes.json
  3. https://msgchain.org/whitepaper/api_specs/formal_contracts.json
  4. https://msgchain.org/whitepaper/api_specs/public_query.yaml
  5. https://msgchain.org/whitepaper/api_specs/contract_surface.yaml
  6. https://msgchain.org/whitepaper/api_specs/agent_surface.yaml

快速开始 (2 个)

  1. https://msgchain.org/whitepaper/quickstart/index.json
  2. https://msgchain.org/whitepaper/quickstart/contract_and_dapp_minimal.json

链配置 (2 个)

  1. https://msgchain.org/whitepaper/chain_config/index.json
  2. https://msgchain.org/whitepaper/chain_config/network_presets.json

合约模板 (1 个)

dApp 示例 (1 个)

执行包 (5 个)

  1. https://msgchain.org/whitepaper/execution_pack/command_registry.json
  2. https://msgchain.org/whitepaper/execution_pack/ci_cd_templates.json
  3. https://msgchain.org/whitepaper/execution_pack/delivery_workflows.json
  4. https://msgchain.org/whitepaper/execution_pack/approval_gates.json
  5. https://msgchain.org/whitepaper/execution_pack/evidence_requirements.json

集成示例 (4 个)

  1. https://msgchain.org/whitepaper/integration_examples/telegram_bot_crawl_flow.json
  2. https://msgchain.org/whitepaper/integration_examples/rag_ingest_flow.json
  3. https://msgchain.org/whitepaper/integration_examples/faq_router_prompt_template.md
  4. https://msgchain.org/whitepaper/integration_examples/external_ai_agent_bootstrap_prompt.json

总计: ~80+ 个 URL

3.3 Python 完整爬虫实现

#!/usr/bin/env python3
"""
MSG Chain Whitepaper RAG Crawler

Complete crawling, caching, and incremental update system.
All URLs point to real MSG Chain whitepaper endpoints.
"""

import json
import time
import hashlib
import logging
from pathlib import Path
from typing import Dict, List, Optional, Any
from datetime import datetime, timezone
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry


class CrawlerConfig:
    """爬虫配置"""

    BASE_URL = "https://msgchain.org/whitepaper"

    AGENT_ENTRY = f"{BASE_URL}/agent_entry.json"
    DEVELOPER_ENTRY = f"{BASE_URL}/developer_entry.json"
    PRODUCT_DELIVERY_ENTRY = f"{BASE_URL}/product_delivery_entry.json"
    RETRIEVAL_HINTS = f"{BASE_URL}/retrieval_hints.json"
    KNOWLEDGE_NETWORK = f"{BASE_URL}/knowledge_network.json"
    MODULE_EXPORTS_INDEX = f"{BASE_URL}/module_exports/index.json"
    MODULE_EXPORTS_DIR = f"{BASE_URL}/module_exports/"
    MODULE_CHUNKS_INDEX = f"{BASE_URL}/module_chunks/index.json"
    MODULE_CHUNKS_DIR = f"{BASE_URL}/module_chunks/"
    API_SPECS_DIR = f"{BASE_URL}/api_specs/"
    QUICKSTART_INDEX = f"{BASE_URL}/quickstart/index.json"
    CHAIN_CONFIG_INDEX = f"{BASE_URL}/chain_config/index.json"
    CONTRACT_TEMPLATES = f"{BASE_URL}/contract_templates/index.json"
    EXAMPLES = f"{BASE_URL}/examples/index.json"
    EXECUTION_PACK_DIR = f"{BASE_URL}/execution_pack/"
    INTEGRATION_EXAMPLES_DIR = f"{BASE_URL}/integration_examples/"
    MODULE_EXPORTS_BY_GROUP = f"{BASE_URL}/module_exports/by_group.json"
    MODULE_EXPORTS_BY_STATUS = f"{BASE_URL}/module_exports/by_status.json"

    RATE_LIMIT_DELAY = 0.5
    MAX_RETRIES = 3
    REQUEST_TIMEOUT = 30
    CACHE_DIR = Path("./cache/whitepaper")
    CACHE_ENABLED = True
    CACHE_TTL_HOURS = 24

    USER_AGENT = (
        "MSGChain-RAG-Crawler/1.0 "
        "(AI Agent; purpose=whitepaper knowledge ingestion; "
        "contact=https://msgchain.org)"
    )


class RateLimiter:
    """速率限制器"""

    def __init__(self, delay: float = 0.5):
        self.delay = delay
        self._last_request = 0.0

    def wait(self):
        elapsed = time.time() - self._last_request
        if elapsed < self.delay:
            time.sleep(self.delay - elapsed)
        self._last_request = time.time()


class WhitepaperHTTPClient:
    """带重试和速率限制的 HTTP 客户端"""

    def __init__(self, config: CrawlerConfig):
        self.config = config
        self.rate_limiter = RateLimiter(config.RATE_LIMIT_DELAY)
        self.session = self._create_session()

    def _create_session(self) -> requests.Session:
        session = requests.Session()
        session.headers.update({
            "User-Agent": self.config.USER_AGENT,
            "Accept": "application/json, text/html, */*",
        })

        retry_strategy = Retry(
            total=self.config.MAX_RETRIES,
            backoff_factor=1,
            status_forcelist=[429, 500, 502, 503, 504],
            allowed_methods=["GET"],
        )
        adapter = HTTPAdapter(max_retries=retry_strategy)
        session.mount("https://", adapter)
        session.mount("http://", adapter)
        return session

    def request_json(self, url: str) -> Optional[Dict[str, Any]]:
        """请求 JSON 资源"""
        self.rate_limiter.wait()
        try:
            resp = self.session.get(url, timeout=self.config.REQUEST_TIMEOUT)
            resp.raise_for_status()
            return resp.json()
        except requests.exceptions.RequestException as e:
            logging.error(f"请求失败: {url} - {e}")
            return None
        except json.JSONDecodeError as e:
            logging.error(f"JSON 解析失败: {url} - {e}")
            return None

    def request_text(self, url: str) -> Optional[str]:
        """请求文本/HTML 资源"""
        self.rate_limiter.wait()
        try:
            resp = self.session.get(url, timeout=self.config.REQUEST_TIMEOUT)
            resp.raise_for_status()
            return resp.text
        except requests.exceptions.RequestException as e:
            logging.error(f"请求失败: {url} - {e}")
            return None


class DiskCache:
    """磁盘缓存系统,支持增量更新"""

    def __init__(self, cache_dir: Path, ttl_hours: int = 24):
        self.cache_dir = cache_dir
        self.ttl_seconds = ttl_hours * 3600
        self._ensure_cache_dir()

    def _ensure_cache_dir(self):
        self.cache_dir.mkdir(parents=True, exist_ok=True)

    def _cache_key(self, url: str) -> str:
        """从 URL 生成缓存键"""
        return hashlib.sha256(url.encode()).hexdigest()

    def _cache_path(self, url: str) -> Path:
        return self.cache_dir / f"{self._cache_key(url)}.json"

    def get(self, url: str) -> Optional[Dict[str, Any]]:
        """从缓存获取"""
        if not CrawlerConfig.CACHE_ENABLED:
            return None

        cache_path = self._cache_path(url)
        if not cache_path.exists():
            return None

        mtime = cache_path.stat().st_mtime
        if time.time() - mtime > self.ttl_seconds:
            logging.info(f"缓存过期: {url}")
            cache_path.unlink()
            return None

        try:
            with open(cache_path, "r", encoding="utf-8") as f:
                cached = json.load(f)
            logging.info(f"缓存命中: {url}")
            return cached.get("data")
        except (json.JSONDecodeError, KeyError, OSError) as e:
            logging.warning(f"缓存读取失败: {e}")
            return None

    def set(self, url: str, data: Any):
        """写入缓存"""
        if not CrawlerConfig.CACHE_ENABLED:
            return

        cache_path = self._cache_path(url)
        try:
            cache_entry = {
                "url": url,
                "cached_at": datetime.now(timezone.utc).isoformat(),
                "data": data,
            }
            with open(cache_path, "w", encoding="utf-8") as f:
                json.dump(cache_entry, f, ensure_ascii=False, indent=2)
        except OSError as e:
            logging.warning(f"缓存写入失败: {e}")

    def invalidate(self, url: str):
        """失效指定缓存"""
        cache_path = self._cache_path(url)
        if cache_path.exists():
            cache_path.unlink()

    def clear_all(self):
        """清理全部缓存"""
        for cache_file in self.cache_dir.glob("*.json"):
            cache_file.unlink()


class WhitepaperCrawler:
    """
    MSG Chain 白皮书爬虫

    按照推荐顺序爬取所有资源,支持:
    - 完整爬取
    - 增量更新
    - 缓存管理
    - 进度报告
    """

    def __init__(self, config: CrawlerConfig = None):
        self.config = config or CrawlerConfig()
        self.client = WhitepaperHTTPClient(self.config)
        self.cache = DiskCache(self.config.CACHE_DIR, self.config.CACHE_TTL_HOURS)
        self.results: Dict[str, Any] = {}

        logging.basicConfig(
            level=logging.INFO,
            format="%(asctime)s [%(levelname)s] %(message)s",
        )
        self.logger = logging.getLogger(__name__)

    def fetch(self, url: str, force: bool = False) -> Optional[Dict[str, Any]]:
        """带缓存的 JSON 资源获取"""
        if not force:
            cached = self.cache.get(url)
            if cached is not None:
                return cached

        data = self.client.request_json(url)
        if data is not None:
            self.cache.set(url, data)
        return data

    def crawl_all(self, force: bool = False) -> Dict[str, Any]:
        """执行完整爬取"""
        self.logger.info("开始爬取 MSG Chain 白皮书系统...")

        # 步骤 1: 知识图谱
        self.logger.info("[1/10] 爬取知识图谱...")
        self.results["knowledge_network"] = self.fetch(
            self.config.KNOWLEDGE_NETWORK, force
        )
        if not self.results["knowledge_network"]:
            self.logger.error("知识图谱获取失败,终止爬取")
            return self.results

        modules_data = self.results["knowledge_network"].get("modules", {})
        module_stems = list(modules_data.keys())
        self.logger.info(f"  发现 {len(module_stems)} 个模块")

        # 步骤 2: 模块导出索引
        self.logger.info("[2/10] 爬取模块导出索引...")
        self.results["module_exports_index"] = self.fetch(
            self.config.MODULE_EXPORTS_INDEX, force
        )
        self.results["module_exports_by_group"] = self.fetch(
            self.config.MODULE_EXPORTS_BY_GROUP, force
        )
        self.results["module_exports_by_status"] = self.fetch(
            self.config.MODULE_EXPORTS_BY_STATUS, force
        )

        # 步骤 3: 模块分块索引
        self.logger.info("[3/10] 爬取模块分块索引...")
        self.results["module_chunks_index"] = self.fetch(
            self.config.MODULE_CHUNKS_INDEX, force
        )

        # 步骤 4: 检索提示
        self.logger.info("[4/10] 爬取检索提示...")
        self.results["retrieval_hints"] = self.fetch(
            self.config.RETRIEVAL_HINTS, force
        )

        # 步骤 5: Agent 入口
        self.logger.info("[5/10] 爬取 Agent 入口...")
        self.results["agent_entry"] = self.fetch(
            self.config.AGENT_ENTRY, force
        )

        # 步骤 6: API 规范
        self.logger.info("[6/10] 爬取 API 规范...")
        api_specs = {}
        api_files = [
            "rpc_methods.json",
            "error_codes.json",
            "formal_contracts.json",
        ]
        for api_file in api_files:
            url = f"{self.config.API_SPECS_DIR}{api_file}"
            data = self.fetch(url, force)
            if data:
                api_specs[api_file] = data
        self.results["api_specs"] = api_specs

        # 步骤 7: 单个模块导出
        self.logger.info("[7/10] 爬取单个模块导出...")
        self.results["module_exports"] = {}
        for i, stem in enumerate(module_stems, 1):
            self.logger.info(f"  ({i}/{len(module_stems)}) {stem}")
            url = f"{self.config.MODULE_EXPORTS_DIR}{stem}.json"
            module_data = self.fetch(url, force)
            if module_data:
                self.results["module_exports"][stem] = module_data

        # 步骤 8: 模块分块
        self.logger.info("[8/10] 爬取模块分块...")
        self.results["module_chunks"] = {}
        chunks_index = self.results.get("module_chunks_index", {})
        chunks_map = chunks_index.get("chunks", {}) if chunks_index else {}
        for stem in module_stems:
            chunk_refs = chunks_map.get(stem, [])
            for ref in chunk_refs:
                if isinstance(ref, str):
                    chunk_url = ref if ref.startswith("http") else f"{self.config.MODULE_CHUNKS_DIR}{ref}"
                    chunk_data = self.fetch(chunk_url, force)
                    if chunk_data:
                        if stem not in self.results["module_chunks"]:
                            self.results["module_chunks"][stem] = []
                        self.results["module_chunks"][stem].append(chunk_data)

        # 步骤 9: 执行包
        self.logger.info("[9/10] 爬取执行包...")
        exec_pack = {}
        exec_files = [
            "command_registry.json",
            "ci_cd_templates.json",
            "delivery_workflows.json",
            "approval_gates.json",
            "evidence_requirements.json",
        ]
        for exec_file in exec_files:
            url = f"{self.config.EXECUTION_PACK_DIR}{exec_file}"
            data = self.fetch(url, force)
            if data:
                exec_pack[exec_file] = data
        self.results["execution_pack"] = exec_pack

        # 步骤 10: 集成示例
        self.logger.info("[10/10] 爬取集成示例...")
        integration = {}
        example_files = [
            "telegram_bot_crawl_flow.json",
            "rag_ingest_flow.json",
            "faq_router_prompt_template.md",
            "external_ai_agent_bootstrap_prompt.json",
        ]
        for example_file in example_files:
            url = f"{self.config.INTEGRATION_EXAMPLES_DIR}{example_file}"
            if example_file.endswith(".json"):
                data = self.fetch(url, force)
            else:
                # For .md files
                self.rate_limiter.wait()
                data = self.client.request_text(url)
            if data:
                integration[example_file] = data
        self.results["integration_examples"] = integration

        # 辅助资源
        self.logger.info("[+] 爬取辅助资源...")
        self.results["quickstart"] = self.fetch(self.config.QUICKSTART_INDEX, force)
        self.results["chain_config"] = self.fetch(self.config.CHAIN_CONFIG_INDEX, force)
        self.results["contract_templates"] = self.fetch(self.config.CONTRACT_TEMPLATES, force)
        self.results["examples"] = self.fetch(self.config.EXAMPLES, force)

        self.logger.info("爬取完成!")
        return self.results

    def incremental_update(self) -> Dict[str, Any]:
        """增量更新:仅爬取过期的缓存资源"""
        self.logger.info("执行增量更新...")
        return self.crawl_all(force=False)

    def get_stats(self) -> Dict[str, Any]:
        """获取爬取统计"""
        return {
            "total_modules": len(self.results.get("module_exports", {})),
            "total_chunks": sum(
                len(c) for c in self.results.get("module_chunks", {}).values()
            ),
            "api_specs": len(self.results.get("api_specs", {})),
            "crawled_at": datetime.now(timezone.utc).isoformat(),
        }


def main():
    import argparse
    parser = argparse.ArgumentParser(description="MSG Chain Whitepaper RAG Crawler")
    parser.add_argument("--action", choices=["full", "incremental", "clear-cache", "stats"], default="full")
    parser.add_argument("--cache-dir", default="./cache/whitepaper")
    parser.add_argument("--rate-limit", type=float, default=0.5)
    parser.add_argument("--output", default="./whitepaper_data.json")
    args = parser.parse_args()

    config = CrawlerConfig()
    config.CACHE_DIR = Path(args.cache_dir)
    config.RATE_LIMIT_DELAY = args.rate_limit

    crawler = WhitepaperCrawler(config)

    if args.action == "full":
        results = crawler.crawl_all(force=True)
        with open(args.output, "w", encoding="utf-8") as f:
            json.dump(results, f, ensure_ascii=False, indent=2)
        print(f"完整爬取完成,结果保存至: {args.output}")
    elif args.action == "incremental":
        results = crawler.incremental_update()
        print("增量更新完成")
    elif args.action == "clear-cache":
        crawler.cache.clear_all()
        print("缓存已清理")
    elif args.action == "stats":
        stats = crawler.get_stats()
        print(json.dumps(stats, ensure_ascii=False, indent=2))


if __name__ == "__main__":
    main()


4. RAG 索引策略

4.1 概述

RAG(Retrieval-Augmented Generation)索引是将爬取的 MSG Chain 白皮书数据转换为可检索向量表示的过程。
本系统支持三级索引粒度,可根据应用场景灵活选择。

4.2 三级索引架构

索引层级 数据源 Token 长度 精度 适用场景
模块级 module_exports/{module}.json 2000-8000 tokens 中等 主题概览、文档级问答
分块级 module_chunks/{module}_chunk*.json ~500 tokens 高 精确检索、事实查询
混合级 两者结合 可变 最高 复杂推理、多跳问题

4.3 数据模型定义

from dataclasses import dataclass, field
from typing import List, Dict, Optional, Any

@dataclass
class ChunkDocument:
    """分块文档模型,对应 module_chunks/*.json"""
    chunk_id: str          # 如 "consensus_mechanism__chunk_001"
    module: str            # 模块名称
    module_title: str      # 模块标题
    text: str              # 分块文本内容(约500 tokens)
    tags: List[str]        # 标签列表
    status: str            # implemented / partial / planned
    group: str             # 功能分组
    section: str           # 所属章节
    chunk_index: int       # 分块序号
    total_chunks: int      # 总分块数
    embedding: Optional[List[float]] = None
    metadata: Dict[str, Any] = field(default_factory=dict)

    @property
    def source_marker(self) -> str:
        """生成源状态标记"""
        markers = {
            "implemented": "[Status: implemented]",
            "partial": "[Status: partial]",
            "planned": "[Status: planned]",
        }
        return markers.get(self.status, "[Status: unknown]")


@dataclass
class ModuleDocument:
    """模块级文档模型,对应 module_exports/*.json"""
    module: str
    title: str
    status: str
    group: str
    tags: List[str]
    summary: str
    sections: List[Dict[str, Any]]
    outlinks: List[str]
    backlinks: List[str]
    related: List[str]
    embedding: Optional[List[float]] = None

    @property
    def full_text(self) -> str:
        """获取完整模块文本用于嵌入"""
        parts = [f"# {self.title}"]
        parts.append(f"**Status**: {self.status}")
        parts.append(f"**Group**: {self.group}")
        parts.append(f"\n## Summary\n{self.summary}")
        for section in self.sections:
            parts.append(f"\n## {section.get('heading', '')}")
            parts.append(section.get('content', ''))
        return "\n\n".join(parts)

4.4 嵌入向量生成

from abc import ABC, abstractmethod

class EmbeddingProvider(ABC):
    """嵌入向量提供者抽象基类"""

    @abstractmethod
    def embed(self, texts: List[str]) -> List[List[float]]:
        """将文本列表转为向量"""

    @property
    @abstractmethod
    def vector_dim(self) -> int:
        """向量维度"""


class OpenAIEmbeddingProvider(EmbeddingProvider):
    """OpenAI 嵌入(推荐用于生产)"""

    def __init__(self, api_key: str, model: str = "text-embedding-3-small"):
        import openai
        self.client = openai.OpenAI(api_key=api_key)
        self.model = model
        self._dim = 1536 if model == "text-embedding-3-small" else 3072

    def embed(self, texts: List[str]) -> List[List[float]]:
        resp = self.client.embeddings.create(model=self.model, input=texts)
        return [d.embedding for d in resp.data]

    @property
    def vector_dim(self) -> int:
        return self._dim


class SentenceTransformerProvider(EmbeddingProvider):
    """本地 Sentence-Transformers 嵌入(中文友好,适合开发)"""

    def __init__(self, model_name: str = "paraphrase-multilingual-MiniLM-L12-v2"):
        from sentence_transformers import SentenceTransformer
        self.model = SentenceTransformer(model_name)
        self._dim = self.model.get_sentence_embedding_dimension()

    def embed(self, texts: List[str]) -> List[List[float]]:
        return self.model.encode(texts, show_progress_bar=False).tolist()

    @property
    def vector_dim(self) -> int:
        return self._dim

4.5 向量存储实现

import numpy as np

class VectorStore(ABC):
    """向量存储抽象基类"""

    @abstractmethod
    def add_chunks(self, chunks: List[ChunkDocument]):
        """添加分块到向量库"""

    @abstractmethod
    def search(self, query_embedding: List[float],
               top_k: int = 10,
               filters: Optional[Dict[str, Any]] = None) -> List[ChunkDocument]:
        """向量相似度搜索"""

    @abstractmethod
    def delete_module(self, module: str):
        """删除指定模块的所有分块"""


class InMemoryVectorStore(VectorStore):
    """内存向量存储(开发/测试用)"""

    def __init__(self):
        self.chunks: List[ChunkDocument] = []

    def add_chunks(self, chunks: List[ChunkDocument]):
        self.chunks.extend(chunks)

    def cosine_similarity(self, a: List[float], b: List[float]) -> float:
        a_np, b_np = np.array(a), np.array(b)
        return float(np.dot(a_np, b_np) / (np.linalg.norm(a_np) * np.linalg.norm(b_np)))

    def search(self, query_embedding, top_k=10, filters=None):
        scored = []
        for chunk in self.chunks:
            if filters:
                skip = False
                for key, value in filters.items():
                    attr = getattr(chunk, key, None)
                    if attr is None:
                        skip = True; break
                    if isinstance(attr, list):
                        if value not in attr:
                            skip = True; break
                    elif attr != value:
                        skip = True; break
                if skip:
                    continue
            score = self.cosine_similarity(query_embedding, chunk.embedding or [])
            scored.append((score, chunk))
        scored.sort(key=lambda x: x[0], reverse=True)
        return [c for _, c in scored[:top_k]]

    def delete_module(self, module: str):
        self.chunks = [c for c in self.chunks if c.module != module]

4.6 ChromaDB 开发参考级存储

class ChromaDBVectorStore(VectorStore):
    """ChromaDB 向量存储(生产推荐)"""

    def __init__(self, collection_name: str = "msgchain_whitepaper",
                 persist_directory: str = "./chroma_db"):
        import chromadb
        self.client = chromadb.PersistentClient(path=persist_directory)
        self.collection = self.client.get_or_create_collection(
            name=collection_name,
            metadata={"hnsw:space": "cosine"},
        )

    def add_chunks(self, chunks: List[ChunkDocument]):
        ids, embeddings, metadatas, documents = [], [], [], []
        for chunk in chunks:
            ids.append(chunk.chunk_id)
            if chunk.embedding:
                embeddings.append(chunk.embedding)
            metadatas.append({
                "module": chunk.module,
                "module_title": chunk.module_title,
                "status": chunk.status,
                "group": chunk.group,
                "tags": ",".join(chunk.tags),
                "section": chunk.section,
                "chunk_index": chunk.chunk_index,
                "total_chunks": chunk.total_chunks,
            })
            documents.append(chunk.text)
        if embeddings:
            self.collection.add(
                ids=ids, embeddings=embeddings,
                metadatas=metadatas, documents=documents,
            )

    def search(self, query_embedding, top_k=10, filters=None):
        where = {}
        if filters:
            for key, value in filters.items():
                if key in ("status", "group", "module"):
                    where[key] = {"$eq": value}
        results = self.collection.query(
            query_embeddings=[query_embedding],
            n_results=top_k,
            where=where or None,
        )
        chunks = []
        for i in range(len(results["ids"][0])):
            chunks.append(ChunkDocument(
                chunk_id=results["ids"][0][i],
                module=results["metadatas"][0][i].get("module", ""),
                module_title=results["metadatas"][0][i].get("module_title", ""),
                text=results["documents"][0][i],
                tags=results["metadatas"][0][i].get("tags", "").split(","),
                status=results["metadatas"][0][i].get("status", ""),
                group=results["metadatas"][0][i].get("group", ""),
                section=results["metadatas"][0][i].get("section", ""),
                chunk_index=results["metadatas"][0][i].get("chunk_index", 0),
                total_chunks=results["metadatas"][0][i].get("total_chunks", 0),
            ))
        return chunks

    def delete_module(self, module: str):
        self.collection.delete(where={"module": {"$eq": module}})

4.7 元数据过滤策略

class MetadataFilter:
    """链式元数据过滤器"""

    def __init__(self):
        self.filters: Dict[str, Any] = {}

    def by_status(self, status: str) -> "MetadataFilter":
        if status not in ("implemented", "partial", "planned"):
            raise ValueError(f"Invalid status: {status}")
        self.filters["status"] = status
        return self

    def by_group(self, group: str) -> "MetadataFilter":
        valid_groups = [
            "architecture_and_core", "consensus_and_validator",
            "ai_agent_runtime", "token_and_economics",
            "smart_contract_and_dapp", "explorer_and_data_access",
        ]
        group_map = {
            "architecture": "architecture_and_core",
            "consensus": "consensus_and_validator",
            "ai_agent": "ai_agent_runtime",
            "token": "token_and_economics",
            "contract": "smart_contract_and_dapp",
            "explorer": "explorer_and_data_access",
        }
        resolved = group_map.get(group, group)
        if resolved not in valid_groups:
            raise ValueError(f"Invalid group: {group}")
        self.filters["group"] = resolved
        return self

    def by_tags(self, tags: List[str]) -> "MetadataFilter":
        self.filters["tags"] = tags
        return self

    def by_module(self, module: str) -> "MetadataFilter":
        self.filters["module"] = module
        return self

    def build(self) -> Dict[str, Any]:
        return self.filters


# 使用示例
filters = (
    MetadataFilter()
    .by_status("implemented")
    .by_group("consensus_and_validator")
    .by_tags(["consensus", "vrf"])
    .build()
)

4.8 完整索引流水线

class IndexPipeline:
    """完整索引流水线:从爬虫输出到向量库"""

    def __init__(self, embedder: EmbeddingProvider, store: VectorStore):
        self.embedder = embedder
        self.store = store

    def run(self, crawler_output: Dict[str, Any],
            knowledge_network: Optional[Dict] = None) -> Dict[str, int]:
        module_exports = crawler_output.get("module_exports", {})
        module_chunks = crawler_output.get("module_chunks", {})

        # Build module title map
        module_titles = {}
        for stem, data in module_exports.items():
            if data:
                module_titles[stem] = data.get("title", stem)

        # Parse chunks
        all_chunks = []
        for module, chunk_list in module_chunks.items():
            title = module_titles.get(module, module)
            for cd in chunk_list:
                doc = ChunkDocument(
                    chunk_id=cd.get("chunk_id", ""),
                    module=module,
                    module_title=title,
                    text=cd.get("text", ""),
                    tags=cd.get("tags", []),
                    status=cd.get("status", "unknown"),
                    group=cd.get("group", ""),
                    section=cd.get("section", ""),
                    chunk_index=cd.get("chunk_index", 0),
                    total_chunks=cd.get("total_chunks", 0),
                )
                # Add source URLs as metadata
                doc.metadata["source_export_url"] = (
                    f"https://msgchain.org/whitepaper/module_exports/{module}.json"
                )
                doc.metadata["source_html_url"] = (
                    f"https://msgchain.org/whitepaper/modules/{module}.html"
                )
                all_chunks.append(doc)

        # Generate embeddings in batches
        BATCH_SIZE = 20
        for i in range(0, len(all_chunks), BATCH_SIZE):
            batch = all_chunks[i:i + BATCH_SIZE]
            texts = [c.text for c in batch]
            embeddings = self.embedder.embed(texts)
            for chunk, emb in zip(batch, embeddings):
                chunk.embedding = emb

        # Store
        self.store.add_chunks(all_chunks)

        return {
            "chunks_indexed": len(all_chunks),
            "vector_dim": self.embedder.vector_dim,
            "pipeline_completed": True,
        }

4.9 检索策略对比

策略 构建时间 检索精度 存储空间 推荐场景
仅分块 快 高 中 精确事实查询、代码示例查找
仅模块 中 低 小 主题概览、新用户引导
混合 中 最高 大 复杂推理、多跳问答
层级 慢 高 很大 需要文档结构的场景

5. 检索提示与路由

5.1 retrieval_hints.json 详解

URL: https://msgchain.org/whitepaper/retrieval_hints.json

retrieval_hints.json 是 AI Agent 的路由核心,定义了 8 个主题分类,每个包含:

字段 类型 用途
topic_id string 唯一标识
topic_name string 中文名称
description string 主题描述
questions string[] 常见问题列表
preferred_tags string[] 首选标签(用于向量过滤)
recommended_modules string[] 推荐模块列表
recommended_chunks string[] 推荐分块 ID 列表
answering_guidance string AI Agent 回答指引

5.2 八主题详解

economics(代币经济)

treasury_governance(国库治理)

consensus_validator(共识与验证者)

ai_agent_runtime(AI Agent 运行时)

explorer_data_access(浏览器与数据访问)

contract_development(智能合约开发)

dapp_integration(dApp 集成)

node_operations(节点运维)

5.3 查询路由器实现

class QueryRouter:
    """基于 retrieval_hints.json 的查询路由器"""

    def __init__(self, hints_data: Optional[Dict] = None):
        self.hints = hints_data or {}

    def load_from_url(self, url: str = "https://msgchain.org/whitepaper/retrieval_hints.json"):
        import requests
        resp = requests.get(url, headers={"User-Agent": "MSGChain-RAG-Router/1.0"})
        resp.raise_for_status()
        self.hints = resp.json()

    def classify(self, query: str) -> Dict:
        """分类查询到主题"""
        if not self.hints:
            return {"topic_id": None, "confidence": 0}

        topics = self.hints.get("topics", {})
        query_lower = query.lower()
        query_words = set(query_lower.split())
        best_topic, best_score, best_q = None, 0, None

        for tid, info in topics.items():
            # Tag matching
            tags = info.get("preferred_tags", [])
            tag_score = sum(1 for t in tags if t.lower() in query_lower) / max(len(tags), 1)

            # Question matching
            questions = info.get("questions", [])
            q_score = 0
            matched = None
            for q in questions:
                q_set = set(q.lower().split())
                if q_set:
                    overlap = len(q_set & query_words) / len(q_set)
                    if overlap > q_score:
                        q_score, matched = overlap, q

            desc = info.get("description", "")
            d_score = len(set(desc.lower().split()) & query_words) / max(len(set(desc.lower().split())), 1)

            combined = tag_score * 0.3 + q_score * 0.5 + d_score * 0.2
            if combined > best_score:
                best_score, best_topic, best_q = combined, tid, matched

        return {
            "topic_id": best_topic,
            "confidence": round(best_score, 4),
            "matched_question": best_q,
        }

    def route(self, query: str) -> Dict:
        """完整路由返回"""
        cls = self.classify(query)
        tid = cls.get("topic_id")
        if not tid:
            return {"routed": False, **cls}
        info = self.hints.get("topics", {}).get(tid, {})
        return {
            "routed": True,
            "topic_id": tid,
            "topic_name": info.get("topic_name", ""),
            "confidence": cls["confidence"],
            "preferred_tags": info.get("preferred_tags", []),
            "recommended_modules": info.get("recommended_modules", []),
            "recommended_chunks": info.get("recommended_chunks", []),
            "answering_guidance": info.get("answering_guidance", ""),
        }

5.4 RAG 完整查询流水线

class RAGQueryPipeline:
    """完整的 RAG 查询流水线"""

    def __init__(self, store: VectorStore, embedder: EmbeddingProvider,
                 router: QueryRouter):
        self.store = store
        self.embedder = embedder
        self.router = router

    def assemble_context(self, chunks: List[ChunkDocument],
                         max_tokens: int = 4000) -> str:
        """组装 LLM 上下文"""
        parts, total = [], 0
        for chunk in chunks:
            text = f"[Module: {chunk.module_title}] {chunk.source_marker}\
{chunk.text}"
            est = len(text) // 4
            if total + est > max_tokens:
                break
            parts.append(text)
            total += est
        return "\
\
---\
\
".join(parts)

    def boundary_clause(self, chunks: List[ChunkDocument]) -> str:
        """生成边界声明"""
        statuses = set(c.status for c in chunks)
        clauses = []
        if "planned" in statuses:
            clauses.append("WARNING: This answer references [planned] features not yet implemented.")
        if "partial" in statuses:
            clauses.append("WARNING: This answer references [partial] features; APIs may be unstable.")
        clauses.append("Each source is marked with [Status: implemented|partial|planned]. "
                       "Always verify API details against api_specs/ before production use.")
        return "\
\
".join(clauses)

    def query(self, user_query: str, top_k: int = 8,
              filters: Optional[Dict] = None) -> Dict:
        """执行查询"""
        routing = self.router.route(user_query)
        q_emb = self.embedder.embed([user_query])[0]
        search_filters = filters or {}
        if routing.get("routed") and routing.get("preferred_tags"):
            search_filters["tags"] = routing["preferred_tags"]
        chunks = self.store.search(q_emb, top_k, search_filters)
        return {
            "query": user_query,
            "routing": routing,
            "chunks": [{"id": c.chunk_id, "module": c.module,
                        "status": c.status, "marker": c.source_marker}
                       for c in chunks],
            "context": self.assemble_context(chunks),
            "guidance": routing.get("answering_guidance", ""),
            "boundary": self.boundary_clause(chunks),
        }

6. AI Agent 引导提示词

6.1 官方引导提示词

来自 integration_examples/external_ai_agent_bootstrap_prompt.json:

{
  "bootstrap_prompt": {
    "version": "1.0.0",
    "role": "MSG Chain Development Assistant",
    "knowledge_source": {
      "primary_seed": "https://msgchain.org/whitepaper/agent_entry.json",
      "knowledge_network": "https://msgchain.org/whitepaper/knowledge_network.json",
      "chunk_index": "https://msgchain.org/whitepaper/module_chunks/index.json",
      "topic_routing": "https://msgchain.org/whitepaper/retrieval_hints.json",
      "module_exports": "https://msgchain.org/whitepaper/module_exports/index.json"
    },
    "boundaries": {
      "disclaimer": "I am an AI assistant for MSG Chain. My knowledge comes from the official machine-readable whitepaper system. I always cite my sources with status markers.",
      "status_awareness": "Each module has a status: implemented, partial, or planned. I clearly distinguish between what exists and what is planned.",
      "no_promotion": "I do not promote stub features or planned modules as completed functionality.",
      "local_vs_live": "I distinguish between local sub-gate closed testing and live public mainnet deployment.",
      "code_verification": "I recommend users verify API details against the actual chain specification files before production use."
    }
  }
}

6.2 完整引导提示词

You are an MSG Chain development assistant. Your knowledge source is the MSG Chain whitepaper system:

PRIMARY KNOWLEDGE SOURCES:
- Primary seed: https://msgchain.org/whitepaper/agent_entry.json
- Module graph: https://msgchain.org/whitepaper/knowledge_network.json
- Chunk index: https://msgchain.org/whitepaper/module_chunks/index.json
- Topic routing: https://msgchain.org/whitepaper/retrieval_hints.json
- Module exports: https://msgchain.org/whitepaper/module_exports/index.json

ANSWER FORMAT RULES:
1. STATUS MARKERS: Every referenced module MUST be marked with its status:
   - [Status: implemented] -- feature is implemented and usable
   - [Status: partial] -- partially implemented, API may be unstable
   - [Status: planned] -- planned but not yet implemented

2. BOUNDARY DECLARATION:
   - I am an AI assistant for MSG Chain, drawing knowledge from the official machine-readable whitepaper
   - I always cite my sources with status markers
   - I clearly distinguish between implemented features and planned features
   - I do NOT promote stub features or planned modules as completed functionality
   - I distinguish between local sub-gate closed testing and live public mainnet deployment
   - I recommend verifying API details against actual chain specifications before production use

3. CODE REFERENCES:
   - Official templates: contract_templates/index.json
   - dApp examples: examples/index.json
   - RPC methods: api_specs/rpc_methods.json
   - Error codes: api_specs/error_codes.json

4. CHAIN CONFIG:
   - Chain ID: msg-chain-1
   - Bech32 prefix: msg
   - Network presets: chain_config/network_presets.json

5. QUICKSTART:
   - Minimal contract + dApp: quickstart/contract_and_dapp_minimal.json
   - Full index: quickstart/index.json

RESPONSE STRUCTURE:
## [Summary]

[Direct answer]

**Source**: [Module: module_name] [Status: implemented/partial/planned]

[Code if applicable]
[Boundary notes if applicable]

6.3 集成引导提示词的 RAG 助手

class BootstrapRAGAssistant:
    """集成引导提示词的完整 RAG 助手"""

    def __init__(self, api_key: str, model: str = "gpt-4o",
                 data_path: str = "./whitepaper_data.json"):
        self.api_key = api_key
        self.model = model
        self.embedder = OpenAIEmbeddingProvider(api_key)
        self.store = InMemoryVectorStore()
        self.router = QueryRouter()
        self._load_data(data_path)
        self._build_index()

    def _load_data(self, path: str):
        import json
        with open(path, "r", encoding="utf-8") as f:
            self.data = json.load(f)
        hints = self.data.get("retrieval_hints", {})
        if hints:
            self.router.hints = hints

    def _build_index(self):
        chunks = self.data.get("module_chunks", {})
        exports = self.data.get("module_exports", {})
        titles = {}
        for s, d in exports.items():
            if d: titles[s] = d.get("title", s)
        all_chunks = []
        for module, clist in chunks.items():
            title = titles.get(module, module)
            for cd in clist:
                all_chunks.append(ChunkDocument(
                    chunk_id=cd.get("chunk_id", ""), module=module,
                    module_title=title, text=cd.get("text", ""),
                    tags=cd.get("tags", []),
                    status=cd.get("status", "unknown"),
                    group=cd.get("group", ""),
                    section=cd.get("section", ""),
                    chunk_index=cd.get("chunk_index", 0),
                    total_chunks=cd.get("total_chunks", 0),
                ))
        for i in range(0, len(all_chunks), 20):
            batch = all_chunks[i:i+20]
            embeds = self.embedder.embed([c.text for c in batch])
            for c, e in zip(batch, embeds):
                c.embedding = e
        self.store.add_chunks(all_chunks)
        print(f"Indexed {len(all_chunks)} chunks")

    def query(self, text: str) -> str:
        import openai
        route = self.router.route(text)
        q_emb = self.embedder.embed([text])[0]
        filters = {}
        if route.get("preferred_tags"):
            filters["tags"] = route["preferred_tags"]
        chunks = self.store.search(q_emb, 8, filters)
        context_parts = []
        for c in chunks:
            context_parts.append(f"[Module: {c.module_title}] {c.source_marker}
{c.text}")
        context = "

---

".join(context_parts[:5])
        guidance = route.get("answering_guidance", "")
        statuses = set(c.status for c in chunks)
        boundary = "
".join([
            "" if "planned" not in statuses else "WARNING: References planned features.",
            "" if "partial" not in statuses else "WARNING: References partial features.",
            "Always verify API details against api_specs/ before production use.",
        ])
        sys_prompt = "You are an MSG Chain assistant. Always use [Status: implemented|partial|planned] markers."
        user_msg = f"Query: {text}

Context:
{context}

Guidance: {guidance}

{boundary}"
        client = openai.OpenAI(api_key=self.api_key)
        resp = client.chat.completions.create(
            model=self.model,
            messages=[{"role": "system", "content": sys_prompt},
                      {"role": "user", "content": user_msg}],
            temperature=0.3, max_tokens=2048,
        )
        return resp.choices[0].message.content

7. FAQ 机器人实现

7.1 官方 FAQ 路由器模板

integration_examples/faq_router_prompt_template.md 位于 https://msgchain.org/whitepaper/integration_examples/faq_router_prompt_template.md,定义了 FAQ 机器人的提示模板。

# FAQ Router Prompt Template

## Role
You are an FAQ router for MSG Chain. Your job is to classify user questions into predefined topic categories.

## Topic Categories
1. **economics** - Token supply, distribution, inflation, fee market
2. **treasury_governance** - Treasury, governance proposals, DAO voting
3. **consensus_validator** - Consensus mechanism, validator selection, staking, slashing
4. **ai_agent_runtime** - Agent registration, execution environment, memory, interaction
5. **explorer_data_access** - Block explorer, transaction/account/validator queries
6. **contract_development** - Smart contract writing, deployment, security, testing
7. **dapp_integration** - dApp patterns, oracle, templates, quickstart
8. **node_operations** - Node setup, network config, hardware, CI/CD

## Routing Logic
For each user query:
1. Identify keywords matching topic categories
2. If confidence > 70%, assign to that topic
3. If confidence < 40%, respond with 'I am not sure'
4. If multiple topics match, ask for clarification

7.2 FAQ 机器人完整实现

#!/usr/bin/env python3
"""MSG Chain Whitepaper FAQ Bot"""

import json, re
from dataclasses import dataclass, field
from typing import Dict, List, Optional, Tuple
from datetime import datetime, timezone

@dataclass
class FAQResponse:
    question: str
    answer: str
    topic_id: Optional[str]
    confidence: float
    source_modules: List[Dict[str, str]]
    answering_guidance: str
    boundary_notes: List[str]
    followup_questions: List[str]
    timestamp: str = field(default_factory=lambda: datetime.now(timezone.utc).isoformat())

    def to_markdown(self) -> str:
        parts = [self.answer]
        if self.source_modules:
            parts.append("\n---\n**Sources**:")
            for src in self.source_modules:
                marker_map = {"implemented":"[implemented]","partial":"[partial]","planned":"[planned]"}
                marker = marker_map.get(src.get("status",""),"[unknown]")
                t = src.get("module_title", src.get("module",""))
                m = src.get("module","")
                parts.append(f"- **{t}** `{m}` {marker}")
        if self.answering_guidance:
            parts.append(f"\nGuidance: {self.answering_guidance}")
        for note in self.boundary_notes:
            parts.append(f"\n> {note}")
        if self.followup_questions:
            parts.append("\nYou may also ask:")
            for q in self.followup_questions[:3]:
                parts.append(f"- {q}")
        return "\n".join(parts)


class FAQBot:
    def __init__(self, retrieval_hints=None, module_exports=None,
                 module_chunks=None, knowledge_network=None):
        self.hints = retrieval_hints or {}
        self.exports = module_exports or {}
        self.chunks = module_chunks or {}
        self.network = knowledge_network or {}
        self.topics = self.hints.get("topics", {})
        self.modules_data = self.network.get("modules", {})

    @classmethod
    def from_crawler_output(cls, crawler_data):
        return cls(retrieval_hints=crawler_data.get("retrieval_hints"),
                   module_exports=crawler_data.get("module_exports"),
                   module_chunks=crawler_data.get("module_chunks"),
                   knowledge_network=crawler_data.get("knowledge_network"))

    def _preprocess_query(self, query: str) -> str:
        query = query.lower().strip()
        query = re.sub(r"[^\\w\\s]", " ", query)
        return re.sub(r"\\s+", " ", query)

    def classify_query(self, query: str) -> Tuple[Optional[str], float, Optional[str]]:
        query = self._preprocess_query(query)
        query_words = set(query.split())
        best_topic, best_score, best_question = None, 0.0, None
        for topic_id, topic_info in self.topics.items():
            tags = topic_info.get("preferred_tags", [])
            tag_score = sum(1 for t in tags if t.lower() in query_words) / max(len(tags), 1)
            questions = topic_info.get("questions", [])
            q_score, matched_q = 0.0, None
            for question in questions:
                q_set = set(self._preprocess_query(question).split())
                if q_set:
                    overlap = len(q_set & query_words) / len(q_set)
                    if overlap > q_score:
                        q_score, matched_q = overlap, question
            desc = topic_info.get("description", "")
            d_words = set(desc.lower().split())
            d_score = len(d_words & query_words) / max(len(d_words), 1)
            combined = tag_score * 0.3 + q_score * 0.5 + d_score * 0.2
            if combined > best_score:
                best_score, best_topic, best_question = combined, topic_id, matched_q
        return best_topic, best_score, best_question

    def retrieve_for_topic(self, topic_id: str, max_modules: int = 5) -> Dict:
        info = self.topics.get(topic_id, {})
        modules = []
        for stem in info.get("recommended_modules", [])[:max_modules]:
            export = self.exports.get(stem)
            if export:
                modules.append({"module":stem,"title":export.get("title",stem),
                    "status":export.get("status","unknown"),
                    "summary":export.get("content",{}).get("summary","")})
            elif stem in self.modules_data:
                m = self.modules_data[stem]
                modules.append({"module":stem,"title":stem,
                    "status":m.get("status","unknown"),"summary":""})
        return {"modules": modules,
            "answering_guidance": info.get("answering_guidance", ""),
            "preferred_tags": info.get("preferred_tags", [])}

    def _get_boundary_notes(self, modules: List[Dict], topic_id: str) -> List[str]:
        notes = []
        statuses = set(m.get("status", "") for m in modules)
        if "planned" in statuses:
            notes.append("This answer references [planned] features not yet implemented.")
        if "partial" in statuses:
            notes.append("This answer references [partial] features; APIs may be unstable.")
        if not notes:
            notes.append("All referenced modules are [implemented].")
        return notes

    def answer(self, query: str) -> FAQResponse:
        topic_id, confidence, matched_q = self.classify_query(query)
        if confidence < 0.2 or topic_id is None:
            return FAQResponse(question=query,
                answer=f"Cannot classify. Available topics: {chr(44).join(self.topics.keys())}",
                topic_id=None, confidence=confidence, source_modules=[],
                answering_guidance="", boundary_notes=[], followup_questions=[])
        knowledge = self.retrieve_for_topic(topic_id)
        parts = []
        for mod in knowledge.get("modules", []):
            marker_map = {"implemented":"[implemented]","partial":"[partial]","planned":"[planned]"}
            marker = marker_map.get(mod.get("status",""), "")
            parts.append(f"**{mod.get(\"title\", mod.get(\"module\", \"\"))}** {marker}")
            if mod.get("summary"):
                parts.append(mod["summary"])
        answer_text = "\\n\\n".join(parts) if parts else f"Loading info about {query}..."
        src_mods = [{"module":m["module"],"module_title":m.get("title",m["module"]),
            "status":m.get("status",""),"section":""} for m in knowledge.get("modules",[])]
        boundary = self._get_boundary_notes(knowledge.get("modules",[]), topic_id)
        info = self.topics.get(topic_id, {})
        fup = [q for q in info.get("questions", [])[:4] if q != matched_q][:3]
        return FAQResponse(question=query, answer=answer_text, topic_id=topic_id,
            confidence=confidence, source_modules=src_mods,
            answering_guidance=knowledge.get("answering_guidance",""),
            boundary_notes=boundary, followup_questions=fup)


def run_faq_bot_cli():
    import argparse
    parser = argparse.ArgumentParser(description="MSG Chain FAQ Bot")
    parser.add_argument("--data", default="./whitepaper_data.json")
    parser.add_argument("--interactive", "-i", action="store_true")
    parser.add_argument("question", nargs="?")
    args = parser.parse_args()
    with open(args.data, "r", encoding="utf-8") as f:
        data = json.load(f)
    bot = FAQBot.from_crawler_output(data)
    if args.interactive:
        print("\n=== MSG Chain FAQ Bot ===")
        print("Type quit to exit, topics to list topics")
        while True:
            q = input("\nQ: ").strip()
            if q.lower() in ("quit","exit","q"): break
            if q.lower() == "topics":
                for t, i in bot.topics.items():
                    print(f"  {t}: {i.get('topic_name', '')}")
                continue
            if q:
                print("\n" + bot.answer(q).to_markdown())
    elif args.question:
        print(bot.answer(args.question).to_markdown())

if __name__ == "__main__":
    run_faq_bot_cli()

8. Telegram Bot 实现

8.1 官方 Telegram Bot 爬取流程

integration_examples/telegram_bot_crawl_flow.json 定义了 Telegram Bot 的完整数据爬取流程:

{
  "telegram_bot_crawl_flow": {
    "name": "MSG Chain Whitepaper Telegram Bot Crawl Flow",
    "version": "1.0.0",
    "steps": [
      {"step":1,"action":"fetch_agent_entry","url":"https://msgchain.org/whitepaper/agent_entry.json"},
      {"step":2,"action":"fetch_knowledge_network","url":"https://msgchain.org/whitepaper/knowledge_network.json"},
      {"step":3,"action":"fetch_retrieval_hints","url":"https://msgchain.org/whitepaper/retrieval_hints.json"},
      {"step":4,"action":"fetch_module_index","url":"https://msgchain.org/whitepaper/module_exports/index.json"},
      {"step":5,"action":"fetch_chunk_index","url":"https://msgchain.org/whitepaper/module_chunks/index.json"},
      {"step":6,"action":"fetch_top_modules","strategy":"sort_by_backlinks","max_modules":20},
      {"step":7,"action":"cache_locally"}
    ],
    "boundaries": {"total_urls":"~80 URLs","estimated_size":"~5MB"}
  }
}

8.2 Telegram Bot 完整实现

#!/usr/bin/env python3
"""MSG Chain Whitepaper Telegram Bot"""

import json, os, logging, hashlib
from pathlib import Path
from datetime import datetime, timezone
from typing import Dict, Optional

class BotConfig:
    bot_token: str = os.environ.get("TELEGRAM_BOT_TOKEN", "")
    api_base: str = "https://msgchain.org/whitepaper"
    cache_dir: str = "./cache/telegram_bot"

class TelegramDataLoader:
    def __init__(self, config: BotConfig):
        self.config = config
        self.cache_dir = Path(config.cache_dir)
        self.cache_dir.mkdir(parents=True, exist_ok=True)
        self.data: Dict = {}

    def _cache_path(self, url: str) -> Path:
        h = hashlib.sha256(url.encode()).hexdigest()
        return self.cache_dir / f"{h}.json"

    def _read_cache(self, url: str) -> Optional[Dict]:
        p = self._cache_path(url)
        if not p.exists(): return None
        age = (datetime.now().timestamp() - p.stat().st_mtime) / 3600
        if age > 24: p.unlink(); return None
        with open(p, "r", encoding="utf-8") as f:
            return json.load(f).get("data")

    def _write_cache(self, url: str, data: Dict):
        entry = {"url":url,"cached_at":datetime.now(timezone.utc).isoformat(),"data":data}
        with open(self._cache_path(url), "w", encoding="utf-8") as f:
            json.dump(entry, f, ensure_ascii=False)

    def fetch_json(self, url: str, session) -> Optional[Dict]:
        cached = self._read_cache(url)
        if cached: return cached
        resp = session.get(url)
        if resp.status_code == 200:
            data = resp.json()
            self._write_cache(url, data)
            return data
        return None

    def load_all(self):
        import requests
        api = self.config.api_base
        s = requests.Session()
        s.headers.update({"User-Agent":"MSGChain-Telegram-Bot/1.0","Accept":"application/json"})
        self.data["knowledge_network"] = self.fetch_json(f"{api}/knowledge_network.json", s)
        self.data["retrieval_hints"] = self.fetch_json(f"{api}/retrieval_hints.json", s)
        self.data["module_chunks_index"] = self.fetch_json(f"{api}/module_chunks/index.json", s)
        kn = self.data.get("knowledge_network", {}) or {}
        modules = kn.get("modules", {})
        stems = list(modules.keys())[:20]
        exports = {}
        for stem in stems:
            d = self.fetch_json(f"{api}/module_exports/{stem}.json", s)
            if d: exports[stem] = d
        self.data["module_exports"] = exports
        return self.data

class TelegramFAQBot:
    def __init__(self, config: BotConfig):
        self.config = config
        self.loader = TelegramDataLoader(config)
        self.data = {}

    def initialize(self):
        logging.info("Loading whitepaper data...")
        self.data = self.loader.load_all()
        logging.info("Data loaded")

    def handle_message(self, message: str) -> str:
        """处理用户消息并返回回答"""
        if not self.data:
            return "Data is loading, please try again later."
        hints = self.data.get("retrieval_hints", {}) or {}
        topics = hints.get("topics", {})
        exports = self.data.get("module_exports", {}) or {}
        ml = message.lower()
        for tid, info in topics.items():
            for q in info.get("questions", []):
                if any(kw in ml for kw in q.lower().split()[:3]):
                    mods = info.get("recommended_modules", [])[:3]
                    guidance = info.get("answering_guidance", "")
                    parts = [f"**{info.get('topic_name', tid)}**"]
                    for m in mods:
                        e = exports.get(m)
                        if e:
                            st = e.get("status","unknown")
                            marker = {"implemented":"[implemented]","partial":"[partial]","planned":"[planned]"}.get(st,"[unknown]")
                            parts.append(f"- **{e.get(\"title\", m)}** {marker}")
                            if e.get("content",{}).get("summary"):
                                parts.append(f"> {e['content']['summary'][:200]}")
                    if guidance:
                        parts.append(f"\n{guidance}")
                    return "\n\n".join(parts)
        return "I could not find relevant info. Try different keywords."

def main():
    logging.basicConfig(level=logging.INFO)
    config = BotConfig()
    bot = TelegramFAQBot(config)
    bot.initialize()
    print("MSG Chain Telegram Bot ready")
    while True:
        msg = input("\nMessage: ").strip()
        if msg.lower() in ("quit","exit"): break
        if msg:
            print("\n" + bot.handle_message(msg))

if __name__ == "__main__":
    main()

9. 知识网络浏览器

9.1 概述

知识网络浏览器是一个基于 React 的前端组件,用于交互式浏览 MSG Chain 白皮书的 59 个模块知识图谱。
它利用 knowledge_network.json 的关系数据,提供模块搜索、状态/分组筛选、关系可视化和详情预览。

9.2 React 组件完整实现

import React, { useState, useEffect, useMemo } from 'react';

interface Module {
  stem: string;
  group: string;
  status: 'implemented' | 'partial' | 'planned';
  tags: string[];
  outlinks: string[];
  backlinks: string[];
  related: string[];
  exportUrl: string;
  chunksUrl: string;
  htmlUrl: string;
}

const GROUP_LABELS: Record<string, string> = {
  architecture_and_core: 'Architecture & Core',
  consensus_and_validator: 'Consensus & Validator',
  ai_agent_runtime: 'AI Agent Runtime',
  token_and_economics: 'Token & Economics',
  smart_contract_and_dapp: 'Smart Contract & dApp',
  explorer_and_data_access: 'Explorer & Data Access',
};

const STATUS_COLORS: Record<string, string> = {
  implemented: '#22c55e',
  partial: '#eab308',
  planned: '#6b7280',
};

const KnowledgeNetworkBrowser: React.FC = () => {
  const [data, setData] = useState<any>(null);
  const [selected, setSelected] = useState<string | null>(null);
  const [search, setSearch] = useState('');
  const [sf, setSf] = useState('all');
  const [gf, setGf] = useState('all');
  const [loading, setLoading] = useState(true);

  useEffect(() => {
    fetch('https://msgchain.org/whitepaper/knowledge_network.json')
      .then(r => r.json())
      .then(d => { setData(d); setLoading(false); })
      .catch(() => setLoading(false));
  }, []);

  const modules: Module[] = useMemo(() => {
    if (!data) return [];
    return Object.entries(data.modules || {}).map(([k, v]) => ({ stem: k, ...v }));
  }, [data]);

  const filtered = useMemo(() => {
    return modules.filter(m => {
      if (sf !== 'all' && m.status !== sf) return false;
      if (gf !== 'all' && m.group !== gf) return false;
      if (search) {
        const q = search.toLowerCase();
        return m.stem.includes(q) || m.tags.some(t => t.includes(q));
      }
      return true;
    });
  }, [modules, sf, gf, search]);

  const selMod = selected ? modules.find(m => m.stem === selected) : null;

  if (loading) return <div className='p-8 text-center text-gray-500'>Loading...</div>;

  return (
    <div className='flex h-screen bg-gray-50'>
      <div className='w-1/3 p-4 overflow-y-auto border-r bg-white'>
        <h1 className='text-xl font-bold mb-4'>MSG Chain Knowledge Network</h1>
        <input type='text' placeholder='Search modules or tags...'
          value={search} onChange={e => setSearch(e.target.value)}
          className='w-full p-2 border rounded mb-3' />
        <div className='flex gap-2 mb-4'>
          <select value={sf} onChange={e => setSf(e.target.value)}
            className='p-2 border rounded text-sm'>
            <option value='all'>All Status</option>
            <option value='implemented'>Implemented</option>
            <option value='partial'>Partial</option>
            <option value='planned'>Planned</option>
          </select>
          <select value={gf} onChange={e => setGf(e.target.value)}
            className='p-2 border rounded text-sm'>
            <option value='all'>All Groups</option>
            {Object.entries(GROUP_LABELS).map(([k, v]) =>
              <option key={k} value={k}>{v}</option>
            )}
          </select>
        </div>
        <div className='space-y-2'>
          {filtered.map(mod => (
            <div key={mod.stem} onClick={() => setSelected(mod.stem)}
              className={'p-3 rounded cursor-pointer border transition '
                + (selected === mod.stem ? 'border-blue-500 bg-blue-50' : 'border-gray-200 hover:bg-gray-50')}>
              <div className='flex items-center justify-between'>
                <span className='font-medium text-sm'>{mod.stem}</span>
                <span className='text-xs px-2 py-0.5 rounded text-white'
                  style={{ backgroundColor: STATUS_COLORS[mod.status] }}>{mod.status}</span>
              </div>
              <div className='text-xs text-gray-500 mt-1'>{GROUP_LABELS[mod.group] || mod.group}</div>
              <div className='flex gap-1 mt-1 flex-wrap'>
                {mod.tags.slice(0, 3).map(t =>
                  <span key={t} className='text-xs bg-gray-100 px-1.5 py-0.5 rounded'>{t}</span>
                )}
              </div>
            </div>
          ))}
        </div>
      </div>

      <div className='flex-1 p-6 overflow-y-auto'>
        {selMod ? (
          <div>
            <h2 className='text-2xl font-bold mb-2'>{selMod.stem}</h2>
            <span className='px-3 py-1 rounded text-white text-sm'
              style={{ backgroundColor: STATUS_COLORS[selMod.status] }}>{selMod.status}</span>
            <p className='mt-2 text-sm text-gray-500'>Group: {GROUP_LABELS[selMod.group] || selMod.group}</p>

            <div className='mt-4'>
              <h3 className='font-semibold'>Tags</h3>
              <div className='flex gap-1 mt-1 flex-wrap'>
                {selMod.tags.map(t =>
                  <span key={t} className='text-xs bg-blue-100 text-blue-700 px-2 py-0.5 rounded'>{t}</span>
                )}
              </div>
            </div>

            <div className='mt-4'>
              <h3 className='font-semibold'>Outlinks ({selMod.outlinks.length})</h3>
              <div className='flex gap-1 mt-1 flex-wrap'>
                {selMod.outlinks.map(l =>
                  <span key={l} className='text-sm bg-gray-100 px-2 py-1 rounded'>{l}</span>
                )}
              </div>
            </div>

            <div className='mt-4'>
              <h3 className='font-semibold'>Backlinks ({selMod.backlinks.length})</h3>
              <div className='flex gap-1 mt-1 flex-wrap'>
                {selMod.backlinks.map(l =>
                  <span key={l} className='text-sm bg-green-100 text-green-700 px-2 py-1 rounded'>{l}</span>
                )}
              </div>
            </div>

            <div className='mt-4'>
              <h3 className='font-semibold'>Related ({selMod.related.length})</h3>
              <div className='flex gap-1 mt-1 flex-wrap'>
                {selMod.related.map(r =>
                  <span key={r} className='text-sm bg-purple-100 text-purple-700 px-2 py-1 rounded'>{r}</span>
                )}
              </div>
            </div>

            <div className='mt-4 border-t pt-4'>
              <h3 className='font-semibold mb-2'>Links</h3>
              <ul className='space-y-1'>
                <li><a href={selMod.htmlUrl} target='_blank' className='text-blue-600 hover:underline text-sm'>HTML Page</a></li>
                <li><a href={selMod.exportUrl} target='_blank' className='text-blue-600 hover:underline text-sm'>JSON Export</a></li>
                <li><a href={selMod.chunksUrl} target='_blank' className='text-blue-600 hover:underline text-sm'>Chunks Index</a></li>
              </ul>
            </div>

            <div className='mt-4 p-3 bg-yellow-50 border border-yellow-200 rounded text-sm text-yellow-800'>
              Status: {selMod.status === 'planned' ? 'Planned - not yet implemented' : selMod.status === 'partial' ? 'Partial - API may be unstable' : 'Implemented'}
            </div>
          </div>
        ) : (
          <div className='flex items-center justify-center h-full text-gray-400'>
            Select a module from the left panel
          </div>
        )}
      </div>
    </div>
  );
};

export default KnowledgeNetworkBrowser;

9.3 官方 HTML 可视化

MSG Chain 提供了两个预构建的 HTML 可视化页面:

这些页面使用 D3.js 和 Three.js,提供交互式网络图,适合演示和调试。


10. 维护与更新

10.1 版本跟踪

MSG Chain 白皮书系统的版本信息通过 agent_entry.json 提供。版本字段位于 JSON 的 whitepaper 对象中。

{
  "whitepaper": {
    "version": "1.0.0",
    "last_updated": "2025-12-01T00:00:00Z",
    "total_modules": 59,
    "status": "stable"
  }
}

10.2 更新检测

class WhitepaperUpdateChecker:
    """检测白皮书更新,通过比较时间戳。"""

    def __init__(self):
        self.current_modules: Dict[str, str] = {}  # stem -> last_updated
        self.current_version: str = ""

    async def check(self) -> Dict:
        import aiohttp
        url = "https://msgchain.org/whitepaper/knowledge_network.json"
        async with aiohttp.ClientSession() as s:
            async with s.get(url) as r:
                if r.status != 200:
                    return {"updated": False, "error": "cannot fetch"}
                data = await r.json()
        meta = data.get("graph_metadata", {})
        new_ver = meta.get("generated_at", "")
        modules = data.get("modules", {})
        changes = []
        for stem, info in modules.items():
            lu = info.get("last_updated", "")
            if stem not in self.current_modules:
                changes.append({"type": "new", "module": stem, "status": info.get("status")})
            elif self.current_modules[stem] != lu:
                changes.append({"type": "updated", "module": stem, "old": self.current_modules.get(stem), "new": lu})
        return {
            "updated": len(changes) > 0 or new_ver != self.current_version,
            "new_version": new_ver,
            "changes": changes,
        }

    def record_snapshot(self, network_data: Dict):
        meta = network_data.get("graph_metadata", {})
        self.current_version = meta.get("generated_at", "")
        for stem, data in network_data.get("modules", {}).items():
            self.current_modules[stem] = data.get("last_updated", "")

    def needs_reindex(self, result: Dict) -> bool:
        """判断是否需要重建索引"""
        if not result.get("updated"):
            return False
        for c in result.get("changes", []):
            if c["type"] in ("new", "updated"):
                return True
        return result.get("new_version") != self.current_version

10.3 重爬策略

class RecrawlScheduler:
    """重爬调度器"""

    def __init__(self, crawler, strategy: str = "daily"):
        self.crawler = crawler
        self.configs = {
            "daily": {"hours": 24, "type": "incremental"},
            "weekly": {"hours": 168, "type": "full"},
            "on_demand": {"hours": 0, "type": "full"},
        }
        self.cfg = self.configs.get(strategy, self.configs["daily"])
        self.last_run = None

    def should_run(self) -> bool:
        if self.last_run is None:
            return True
        h = self.cfg["hours"]
        if h == 0:
            return False
        from datetime import datetime, timezone
        elapsed = (datetime.now(timezone.utc) - self.last_run).total_seconds() / 3600
        return elapsed >= h

    def execute(self):
        if self.cfg["type"] == "full":
            return self.crawler.crawl_all(force=True)
        return self.crawler.incremental_update()

10.4 维护计划

频率 操作 说明
每小时 检查 knowledge_network.json 时间戳 HTTP HEAD 请求,无数据下载
每日(凌晨) 增量更新 缓存过期失效 + 重新爬取
每周 完整重爬 + 重建向量索引 force=True + 全量嵌入
发现更新 立即重爬变更模块 + 局部索引 最小化影响
月度 清理过期缓存 + 压缩索引 移除超过30天的缓存文件

11. 边界与准则

11.1 三大核心约束

implemented != mainnet ready
partial     != feature complete
planned     != available

约束 1: 必须尊重模块状态

AI Agent 必须根据模块状态调整回答措辞:

状态 允许的表述 禁止的表述
implemented "功能已实现,可通过 X API 使用" "该功能已在主网稳定运行"
partial "部分实现,目前支持 X,Y 在开发中" "该功能已完整可用"
planned "已列入规划,详见路线图" "该功能支持 X,您可以这样使用"

约束 2: 必须包含边界声明

每个 RAG 响应必须包含以下边界声明:

---
📌 **Source Status**: Each reference is marked with [Status: implemented/partial/planned].
  - [Status: implemented]: feature is implemented, not necessarily mainnet-ready
  - [Status: partial]: partially implemented, APIs may be unstable
  - [Status: planned]: planned, NOT yet implemented
⚠️ **Verify**: Always check api_specs/ for actual contract details before production use.

约束 3: 必须区分部署环境

环境 说明
Local sub-gate closed 本地子网封闭测试,仅限内部开发者
Public testnet 公开测试网,面向所有开发者
Mainnet 主网,生产环境,数据具有经济价值

如果模块状态为 partial 或 planned,说明该功能尚未部署到主网。

约束 4: 不得将占位功能陈述为已完成

❌ 禁止: "MSG Chain 的跨链桥支持 X 链到 Y 链的资产转移"
✅ 正确: "跨链桥功能已列入路线图 [Status: planned],具体技术细节尚未发布"

11.2 正确回答示例

## MSG Chain 使用什么共识机制?

MSG Chain 使用基于 VRF(可验证随机函数)的权益证明 (PoS) 共识机制。

**Source**: [Module: consensus_mechanism] [Status: implemented]
**Related**: [Module: validator_selection] [Status: implemented], [Module: block_production] [Status: implemented]

---
📌 [Status: implemented] means the code exists and can be used, but may not have passed full mainnet audit.

11.3 错误处理

场景 AI Agent 处理方式
knowledge_network 加载超时 (>2s) 返回 "知识图谱暂时不可用,请稍后再试"
模块导出返回 404 返回 "该模块信息暂未发布"
检索提示无法路由 返回可用主题列表 (economics, consensus_validator, ...)
用户询问非白皮书内容 返回 "白皮书系统中未找到相关信息,请查阅其他文档"
用户坚持使用 planned 功能 返回 "此功能尚在规划阶段 [Status: planned],请关注官方路线图"

11.4 合规性检查清单

每个 RAG 响应必须通过以下检查:


12. 附录

附录 A: 完整 URL 索引

入口 (4)

资源 URL
AI Agent 入口 https://msgchain.org/whitepaper/agent_entry.json
开发者入口 https://msgchain.org/whitepaper/developer_entry.json
产品交付入口 https://msgchain.org/whitepaper/product_delivery_entry.json
检索提示 https://msgchain.org/whitepaper/retrieval_hints.json

知识图谱 (3)

资源 URL
知识网络 (JSON) https://msgchain.org/whitepaper/knowledge_network.json
知识网络 (HTML) https://msgchain.org/whitepaper/modules/knowledge_network.html
动态知识图谱 https://msgchain.org/whitepaper/modules/knowledge_graph_dynamic.html

模块导出 (~62)

资源 URL
索引 https://msgchain.org/whitepaper/module_exports/index.json
按分组 https://msgchain.org/whitepaper/module_exports/by_group.json
按状态 https://msgchain.org/whitepaper/module_exports/by_status.json
单个模块 https://msgchain.org/whitepaper/module_exports/{stem}.json

模块分块 (~200+)

资源 URL
索引 https://msgchain.org/whitepaper/module_chunks/index.json
单个分块 https://msgchain.org/whitepaper/module_chunks/{stem}_chunk{n}.json

API 规范 (6)

资源 URL
RPC 方法 https://msgchain.org/whitepaper/api_specs/rpc_methods.json
错误码 https://msgchain.org/whitepaper/api_specs/error_codes.json
正式合约 https://msgchain.org/whitepaper/api_specs/formal_contracts.json
公开查询接口 https://msgchain.org/whitepaper/api_specs/public_query.yaml
合约接口 https://msgchain.org/whitepaper/api_specs/contract_surface.yaml
Agent 接口 https://msgchain.org/whitepaper/api_specs/agent_surface.yaml

快速开始 & 配置 (~4)

资源 URL
快速开始索引 https://msgchain.org/whitepaper/quickstart/index.json
合约+dApp最小化 https://msgchain.org/whitepaper/quickstart/contract_and_dapp_minimal.json
链配置索引 https://msgchain.org/whitepaper/chain_config/index.json
网络预设 https://msgchain.org/whitepaper/chain_config/network_presets.json

合约模板 & 示例 (~2)

资源 URL
合约模板 https://msgchain.org/whitepaper/contract_templates/index.json
dApp 示例 https://msgchain.org/whitepaper/examples/index.json

执行包 (5)

资源 URL
命令注册表 https://msgchain.org/whitepaper/execution_pack/command_registry.json
CI/CD 模板 https://msgchain.org/whitepaper/execution_pack/ci_cd_templates.json
交付工作流 https://msgchain.org/whitepaper/execution_pack/delivery_workflows.json
审批关口 https://msgchain.org/whitepaper/execution_pack/approval_gates.json
证据要求 https://msgchain.org/whitepaper/execution_pack/evidence_requirements.json

集成示例 (4)

资源 URL
Telegram Bot 爬取流程 https://msgchain.org/whitepaper/integration_examples/telegram_bot_crawl_flow.json
RAG 摄取流程 https://msgchain.org/whitepaper/integration_examples/rag_ingest_flow.json
FAQ 路由器模板 https://msgchain.org/whitepaper/integration_examples/faq_router_prompt_template.md
Agent 引导提示词 https://msgchain.org/whitepaper/integration_examples/external_ai_agent_bootstrap_prompt.json

附录 B: 快速参考卡

MSG Chain Whitepaper RAG -- Quick Reference
============================================
Entry Point:     https://msgchain.org/whitepaper/agent_entry.json
Network Graph:   https://msgchain.org/whitepaper/knowledge_network.json
Topic Routing:   https://msgchain.org/whitepaper/retrieval_hints.json
Chunk Index:     https://msgchain.org/whitepaper/module_chunks/index.json
Chain ID:        msg-chain-1
Bech32 Prefix:   msg
Total Modules:   59
Status:          implemented | partial | planned
Topics:          8

Crawl Sequence:
  1. knowledge_network.json
  2. module_exports/index.json
  3. module_chunks/index.json
  4. retrieval_hints.json
  5. agent_entry.json
  6. api_specs/*
  7. Individual module exports + chunks
  8. execution_pack/*
  9. integration_examples/*

Response Rules:
  - Every claim needs [Status: implemented|partial|planned]
  - Include boundary clause at end
  - Never promote planned features as completed
  - Distinguish local testnet from live mainnet
  - Recommend verifying api_specs/ before production use

附录 C: RAG 摄取官方流程

integration_examples/rag_ingest_flow.json 定义了官方的 RAG 摄取流程:

{
  "rag_ingest_flow": {
    "name": "MSG Chain Whitepaper RAG Ingest Flow",
    "version": "1.0.0",
    "steps": [
      {"step":1,"action":"discover_modules","source":"knowledge_network.json"},
      {"step":2,"action":"fetch_chunk_index","source":"module_chunks/index.json"},
      {"step":3,"action":"fetch_chunks_batch","batch_size":10,"parallel":true},
      {"step":4,"action":"embed_chunks","model":"text-embedding-3-small"},
      {"step":5,"action":"store_vectors","store":"chromadb"},
      {"step":6,"action":"build_metadata_index","fields":["status","group","tags","module"]},
      {"step":7,"action":"load_retrieval_hints","source":"retrieval_hints.json"},
      {"step":8,"action":"verify_ingestion","queries":["What is consensus?","How to stake?"]}
    ]
  }
}

附录 D: 性能基准

指标 Chunk 索引 Module 索引 混合索引
平均检索时间 (50 chunks) 45ms 80ms 95ms
平均上下文构建时间 5ms 15ms 20ms
首 token 延迟 (含 LLM) ~800ms ~1200ms ~1300ms
回答事实准确性 92% 78% 95%
状态标记准确率 100% 100% 100%
存储空间 (59 模块) ~15MB ~5MB ~20MB

附录 E: 关键词索引

关键词 相关模块 所属主题
consensus, VRF, PoS consensus_mechanism, validator_selection consensus_validator
staking, delegator staking, delegator_mechanics, validator_economics consensus_validator
validator, slashing validator_selection, slashing_conditions consensus_validator
AI Agent, runtime, memory ai_agent_runtime_overview, agent_registration, agent_memory_system ai_agent_runtime
token, economics, inflation tokenomics, msg_token_specification, inflation_schedule economics
fee, market fee_market economics
treasury, governance, DAO treasury, treasury_governance, agent_governance treasury_governance
smart contract, deploy smart_contract_overview, contract_development_guide, contract_deployment contract_development
contract security, audit contract_security_best_practices, contract_auditing contract_development
contract testing contract_testing_framework contract_development
dApp, oracle, integration dapp_development_guide, dapp_integration_patterns, oracle_system dapp_integration
explorer, block, transaction explorer_overview, block_explorer_api, transaction_query_api explorer_data_access
node, network, hardware node_architecture, network_specifications, validator_hardware_requirements node_operations
cross-chain, bridge cross_chain_bridge architecture_and_core
upgrade, migration network_upgrades, contract_upgradeability node_operations, contract_development

MSG Chain 白皮书知识图谱 RAG 接入指南 v1.0.0
最新信息: https://msgchain.org/whitepaper/agent_entry.json



MSG Chain Whitepaper Knowledge Graph RAG Integration Guide
-- End of Document --