dApp Docs/IBC 协议升级与链恢复流程指南
Development reference. Not independently verified for production.

IBC 协议升级与链恢复流程指南

数据来源:MSG Chain 代码库核实

主网状态: No-Go — 当前 MSGChain 主网裁决为 No-Go,以下内容反映代码实际状态,不代表生产可用。


目录

  1. IBC 协议升级机制
  2. ibc-go v7 → v8+ 升级指南
  3. 链升级与 IBC 连续性
  4. 轻客户端冻结与恢复
  5. 通道关闭与重启
  6. 数据迁移与状态证明
  7. MSG Chain 的 IBC 升级历史与兼容性策略
  8. 跨链协调
  9. 回滚策略
  10. 升级自动化
  11. 案例:Cosmos 生态 IBC 升级事件回顾
  12. 总结

1. IBC 协议升级机制

1.1 升级概述

IBC(Inter-Blockchain Communication)协议作为 Cosmos 生态的核心互操作协议,其升级涉及多个层次。对于运行 msg-chain-1 的网络而言,理解 IBC 升级机制是保障跨链通信连续性的基础。

IBC 升级可以从以下维度进行分类:

升级类型 影响范围 协调难度 典型场景
ibc-go 版本升级 协议层 高 v7 → v8/v9
Light Client 升级 客户端层 中 02-client 升级
通道升级 通道层 中 通道版本协商
链自身升级 共识层 高 SDK 版本升级

1.2 ibc-go 版本升级机制

ibc-go 是 IBC 协议的 Go 语言参考实现。每次主版本升级都可能包含以下变更:

核心模块变更:

升级路径:

v7.x → v7.y(补丁升级)→ 兼容
v7.x → v8.x(主版本升级)→ 需迁移代码
v8.x → v9.x(主版本升级)→ 需迁移代码

对于 msg-chain-1,升级 ibc-go 版本时需同步升级以下依赖:

github.com/cosmos/ibc-go/v7 → github.com/cosmos/ibc-go/v8

1.3 Light Client 升级

轻客户端是 IBC 安全的核心组件。升级轻客户端通常涉及:

ICS-02 Client 升级:

升级触发条件:

  1. 治理提案通过新客户端参数
  2. 链升级后客户端状态不再兼容
  3. 安全漏洞修复需要更新客户端逻辑

MSG Chain 的 Light Client 配置:

{
  "@type": "/ibc.lightclients.tendermint.v1.ClientState",
  "chain_id": "msg-chain-1",
  "trust_level": {
    "numerator": 1,
    "denominator": 3
  },
  "trusting_period": "336h",
  "unbonding_period": "504h",
  "max_clock_drift": "10s",
  "frozen_height": {
    "revision_number": "0",
    "revision_height": "0"
  },
  "latest_height": {
    "revision_number": "1",
    "revision_height": "1000000"
  },
  "proof_specs": [],
  "upgrade_path": []
}

1.4 通道升级机制

IBC 通道升级允许在不关闭通道的情况下更新通道参数。这在 ibc-go v8 中得到原生支持。

可升级的通道参数:

通道升级流程:

阶段一:通道升级初始化
  ┌─────────┐          ┌─────────┐
  │ Chain A │          │ Chain B │
  │(发起方)│          │(接收方)│
  └────┬────┘          └────┬────┘
       │                     │
       │ ChanUpgradeOpen     │
       │────────────────────>│
       │                     │
       │                     │  ChanUpgradeOpen
       │<────────────────────│
       │                     │
       │ ChanUpgradeAck      │
       │────────────────────>│
       │                     │
       │                     │  ChanUpgradeConfirm
       │<────────────────────│
       │                     │
       │ ChanUpgradeOpen     │
       │────────────────────>│
       │                     │
       ┌─┴─┐               ┌─┴─┐
       │完成│               │完成│
       └───┘               └───┘

通道升级的治理恢复路径:

如果通道升级过程中出现异常(如对端链未及时响应),可通过以下治理提案恢复:

// governance proposal for channel upgrade recovery
type ChannelUpgradeRecoveryProposal struct {
    Title       string
    Description string
    Channel     string
    Port        string
}

1.5 升级兼容性矩阵

组件 v7.x v8.0 v8.1+ v9.0
ICS-02 Client ✓ ✓ ✓ ✓
ICS-03 Connection ✓ ✓ ✓ ✓
ICS-04 Channel ✓ ✓ ✓ ✓
ICS-20 Transfer ✓ ✓ ✓ ✓
ICS-27 Interchain Accounts 需迁移 ✓ ✓ ✓
ICS-721 NFT Transfer ✓ ✓ ✓ ✓
ICS-28 Fee Middleware ✓ ✓ ✓ ✓
Packet Forward Middleware 需迁移 ✓ ✓ ✓

2. ibc-go v7 → v8+ 升级指南

2.1 迁移概览

从 ibc-go v7 升级到 v8 是 msg-chain-1 需要重点关注的迁移路径。v8 引入了多项重大变更,包括通道升级原生支持、ICS-27 增强、以及 API 重构。

前置条件检查清单:

□ 当前 ibc-go 版本:v7.3.x 或更高
□ Cosmos SDK 版本:v0.47.x
□ CometBFT 版本:v0.37.x
□ CosmWasm 版本:v1.5.x
□ Go 版本:1.21+

2.2 模块迁移

2.2.1 导入路径变更

// v7 导入路径
import (
    "github.com/cosmos/ibc-go/v7/modules/apps/transfer"
    "github.com/cosmos/ibc-go/v7/modules/core"
    "github.com/cosmos/ibc-go/v7/modules/core/02-client"
)

// v8 导入路径
import (
    "github.com/cosmos/ibc-go/v8/modules/apps/transfer"
    "github.com/cosmos/ibc-go/v8/modules/core"
    "github.com/cosmos/ibc-go/v8/modules/core/02-client"
)

2.2.2 应用模块注册变更

v7 方式:

app.IBCKeeper = ibckeeper.NewKeeper(
    appCodec,
    keys[ibcexported.StoreKey],
    app.GetSubspace(ibcexported.ModuleName),
    app.StakingKeeper,
    app.UpgradeKeeper,
    app.ScopedIBCKeeper,
)

v8 方式:

app.IBCKeeper = ibckeeper.NewKeeper(
    appCodec,
    keys[ibcexported.StoreKey],
    app.GetSubspace(ibcexported.ModuleName),
    app.StakingKeeper,
    app.UpgradeKeeper,
    app.ScopedIBCKeeper,
    "", // authority(通常留空或设置治理模块地址)
)

关键变更:v8 的 NewKeeper 新增了 authority 参数,用于指定可以执行治理操作的模块地址。

2.2.3 ICS-20 Transfer 变更

v7 Transfer Keeper:

transferKeeper := ibctransferkeeper.NewKeeper(
    appCodec,
    keys[ibctransfertypes.StoreKey],
    app.GetSubspace(ibctransfertypes.ModuleName),
    app.IBCKeeper.ChannelKeeper,
    app.IBCKeeper.PortKeeper,
    app.AccountKeeper,
    app.BankKeeper,
    app.ScopedTransferKeeper,
)

v8 Transfer Keeper:

transferKeeper := ibctransferkeeper.NewKeeper(
    appCodec,
    keys[ibctransfertypes.StoreKey],
    app.GetSubspace(ibctransfertypes.ModuleName),
    app.IBCKeeper.ChannelKeeper,
    app.IBCKeeper.ChannelKeeper,
    app.IBCKeeper.PortKeeper,
    app.AccountKeeper,
    app.BankKeeper,
    app.ScopedTransferKeeper,
    "", // authority
)

注意:v8 中的 ICS4Wrapper 和 ChannelKeeper 参数分离,以及新增的 authority 参数。

2.3 API 变更详解

2.3.1 核心 API 变更

MsgServer 变更:

// v7 - MsgServer 接口
type MsgServer interface {
    Transfer(context.Context, *MsgTransfer) (*MsgTransferResponse, error)
}

// v8 - MsgServer 接口(新增 UpdateParams)
type MsgServer interface {
    Transfer(context.Context, *MsgTransfer) (*MsgTransferResponse, error)
    UpdateParams(context.Context, *MsgUpdateParams) (*MsgUpdateParamsResponse, error)
}

查询接口变更:

// v7 查询路径
/cosmos.ibc.core.channel.v1.Query/Channel
/cosmos.ibc.core.channel.v1.Query/Channels

// v8 新增查询路径
/cosmos.ibc.core.channel.v1.Query/ChannelUpgrade
/cosmos.ibc.core.channel.v1.Query/ChannelUpgrades
/cosmos.ibc.core.channel.v1.Query/NextChannelUpgradeSequence

CLI 命令变更:

# v7 CLI
msgd tx ibc transfer transfer [src-port] [src-channel] [receiver] [amount]

# v8 CLI(新增通道升级相关命令)
msgd tx ibc channel upgrade [port-id] [channel-id] [flags]
msgd query ibc channel upgrade [port-id] [channel-id]
msgd tx ibc channel upgrade-recovery [port-id] [channel-id] [flags]

2.4 存储迁移

ibc-go v8 引入了存储布局变更,需要执行状态迁移:

升级处理函数示例:

func CreateV8UpgradeHandler(
    mm *module.Manager,
    configurator module.Configurator,
    ibcKeeper *ibckeeper.Keeper,
) upgradetypes.UpgradeHandler {
    return func(ctx sdk.Context, plan upgradetypes.Plan, fromVM module.VersionMap) (module.VersionMap, error) {
        // 执行 IBC 存储迁移
        ibcKeeper.UpgradeKeeper.SetUpgradeVersion(ctx, "v8")

        // 迁移所有轻客户端
        clients := ibcKeeper.ClientKeeper.GetAllClients(ctx)
        for _, client := range clients {
            if err := ibcKeeper.ClientKeeper.MigrateClient(ctx, client.ClientId); err != nil {
                return nil, err
            }
        }

        // 迁移通道存储
        channels := ibcKeeper.ChannelKeeper.GetAllChannels(ctx)
        for _, channel := range channels {
            if err := ibcKeeper.ChannelKeeper.MigrateChannel(ctx, channel.PortId, channel.ChannelId); err != nil {
                return nil, err
            }
        }

        // 运行模块迁移
        return mm.RunMigrations(ctx, configurator, fromVM)
    }
}

2.5 兼容性测试

2.5.1 跨版本兼容性矩阵

测试 ibc-go v7 与 v8 之间的互通性:

测试场景 v7 → v7 v7 → v8 v8 → v8
ICS-20 Transfer ✓ ✓ ✓
ICS-27 ICA ✓ ⚠️ 需协调 ✓
ICS-721 NFT ✓ ✓ ✓
ICS-28 Fee ✓ ⚠️ 需协调 ✓
Packet Forward ✓ ✓ ✓

2.5.2 推荐的测试方案

# 本地启动两个不同版本的链进行测试
# 启动 v7 链
msgd start --home /tmp/msg-v7 --chain-id msg-chain-1

# 启动 v8 链(不同数据目录)
msgd start --home /tmp/msg-v8 --chain-id msg-chain-2

# 建立 IBC 连接
msgd tx ibc connection open-init --client-id 07-tendermint-0 \
  --counterparty-client-id 07-tendermint-0

# 测试代币转账
msgd tx ibc transfer transfer channel-0 msg1xyz... 1000umsg \
  --from key1 --chain-id msg-chain-1

# 验证跨链交易成功
msgd query ibc transfer denom-trace <hash>

2.6 v8 → v9 迁移要点

ibc-go v9 进一步增强了协议能力:

主要变更:

  1. 通道升级 GA:通道升级功能从 Beta 进入 GA 阶段
  2. ICS-27 增强:Interchain Accounts 支持更丰富的控制
  3. 性能优化:状态证明验证性能提升 30%+
  4. 新的消息类型:MsgRecvPacketResponse 和 MsgTimeoutResponse 支持异步确认

迁移注意事项:

// v9 的新模块参数
type Params struct {
    // 启用通道升级
    ChannelUpgradeEnabled bool `json:"channel_upgrade_enabled"`
    // 最大并发通道升级数
    MaxConcurrentChannelUpgrades uint64 `json:"max_concurrent_channel_upgrades"`
}

3. 链升级与 IBC 连续性

3.1 升级类型与 IBC 影响

链升级可能对 IBC 连接产生不同程度的影响:

升级类型 IBC 影响 恢复方式 用户影响
补丁升级(v1.0.1 → v1.0.2) 无影响 自动 无
SDK 补丁升级 无影响 自动 无
SDK 主版本升级 轻客户端暂停 需更新客户端 短暂暂停
ibc-go 主版本升级 高 需迁移 需协调恢复
共识参数变更 轻客户端可能过期 客户端更新 需手动操作
分叉/回滚 严重 通道关闭 需治理恢复

3.2 Cosmovisor 自动升级

Cosmovisor 是 Cosmos 生态的标准升级工具,对于最小化 IBC 中断时间至关重要。

3.2.1 配置 Cosmovisor

目录结构:

~/.msgd/
├── cosmovisor/
│   ├── current -> genesis/bin/msgd
│   ├── genesis/
│   │   └── bin/
│   │       └── msgd
│   ├── upgrades/
│   │   ├── v2.0.0/
│   │   │   └── bin/
│   │   │       └── msgd
│   │   ├── v3.0.0/
│   │   │   └── bin/
│   │   │       └── msgd
│   │   └── v4.0.0/
│   │       └── bin/
│   │           └── msgd
│   └── config.toml
└── data/

Cosmovisor 配置:

# ~/.msgd/cosmovisor/config.toml
[min_msgs]
minimum_gas_prices = "1000000000umsg"

[upgrade]
# 自动下载升级二进制(需配置可信源)
allow_download_binaries = false
# 升级前备份数据
backup_before_upgrade = true
# 升级超时时间(秒)
upgrade_timeout = "300s"
# 是否在升级失败后自动回滚
auto_rollback = false

系统服务配置:

[Unit]
Description=MSG Chain Node with Cosmovisor
After=network.target

[Service]
Type=simple
User=msg
ExecStart=/usr/local/bin/cosmovisor run start \
  --home /var/lib/msgd \
  --x-crisis-keeper.skip-invariants=true \
  --iavl-disable-fastnode=false
Restart=always
RestartSec=10
LimitNOFILE=65535
Environment="DAEMON_NAME=msgd"
Environment="DAEMON_HOME=/var/lib/msgd"
Environment="DAEMON_RESTART_AFTER_UPGRADE=true"
Environment="DAEMON_ALLOW_DOWNLOAD_BINARIES=false"
Environment="DAEMON_DATA_BACKUP_DIR=/var/backups/msgd"

[Install]
WantedBy=multi-user.target

3.2.2 Cosmovisor 升级生命周期

阶段一:检测升级提案
  ↓
治理提案通过 UpgradeHeight=5000000
  ↓
阶段二:升级前准备(UpgradeHeight - 100 区块)
  同步状态、通知 IBC 对端
  ↓
阶段三:达到升级高度
  Cosmovisor 停止旧二进制 → 备份数据 → 启动新二进制
  ↓
阶段四:升级后验证
  验证 app_hash 一致性 → 重启 IBC 连接
  ↓
阶段五:IBC 恢复
  更新轻客户端 → 恢复通道 → 恢复交易

3.3 升级高度协调

3.3.1 治理提案中的升级高度

{
  "messages": [
    {
      "@type": "/cosmos.upgrade.v1beta1.MsgSoftwareUpgrade",
      "authority": "msg10d07y265gmmuvt4z0w9aw880jnsr700jm9d9zw",
      "plan": {
        "name": "v2.0.0",
        "height": "5000000",
        "info": "{\"binaries\":{\"linux/amd64\":\"https://github.com/msgchain/msgd/releases/download/v2.0.0/msgd-linux-amd64.zip\"}}"
      }
    }
  ],
  "metadata": "https://ipfs.msgchain.org/ipfs/QmXYZ...",
  "deposit": "1000000000umsg"
}

3.3.2 升级窗口选择

理想的升级高度选择需要考虑以下因素:

IBC 协调窗口:

[当前高度]────[通知窗口(1-3天)]────[升级高度]────[恢复窗口(2-4h)]────[正常运行]

升读高度选择的推荐原则:
1. 避免亚洲时区凌晨(UTC 06:00-14:00 最佳)
2. 避开主流链的升级窗口
3. 保证至少 48 小时通知期
4. 预留 6 小时以上的版本验证窗口

推荐升级时间窗口:

时区 推荐时间 说明
UTC 06:00-10:00 亚洲下午,欧美凌晨
CST(Asia/Shanghai) 14:00-18:00 下午工作时间
EST(America) 01:00-05:00 北美凌晨
CET(Europe) 07:00-11:00 欧洲上午

3.3.3 升级前 IBC 暂停策略

在关键升级前,建议主动暂停 IBC 活动以减少风险:

// 升级前暂停 IBC 的逻辑
func PreUpgradeIBCHalt(ctx sdk.Context, ibcKeeper *ibckeeper.Keeper) error {
    // 暂停所有待处理的数据包
    pendingPackets := ibcKeeper.ChannelKeeper.GetAllPacketIds(ctx)
    for _, packet := range pendingPackets {
        if packet.Sequence > 0 {
            // 标记待处理数据包
            log.Printf("Pending packet: channel %s, sequence %d",
                packet.ChannelId, packet.Sequence)
        }
    }

    // 记录升级前的 IBC 状态快照
    ibcKeeper.ClientKeeper.IterateClients(ctx, func(clientID string, state *ibcclient.ClientState) bool {
        log.Printf("Pre-upgrade client state: %s height=%d",
            clientID, state.GetLatestHeight())
        return false
    })

    return nil
}

3.4 IBC 暂停与恢复

3.4.1 IBC 暂停的触发条件

IBC 模块可能在以下情况下自动暂停:

  1. 链达到升级高度:SDK 升级模块暂停所有模块
  2. 轻客户端过期:TrustingPeriod 超时
  3. 连接断开:对端链不可达
  4. 共识故障:检测到无效区块

3.4.2 手动暂停 IBC

# 通过治理提案暂停特定端口
msgd tx gov submit-proposal \
  --title "Pause IBC Transfer Module" \
  --description "Temporarily pause IBC transfers for upgrade" \
  --type TextProposal \
  --deposit 100000000umsg

# 暂停后验证状态
msgd query ibc channel channels
msgd query ibc connection connections

3.4.3 IBC 恢复流程

自动恢复(升级完成后):

func PostUpgradeIBCHandler(ctx sdk.Context, ibcKeeper *ibckeeper.Keeper) error {
    // 验证所有轻客户端状态
    clients := ibcKeeper.ClientKeeper.GetAllClients(ctx)
    for _, client := range clients {
        clientState, ok := ibcKeeper.ClientKeeper.GetClientState(ctx, client.ClientId)
        if !ok {
            return fmt.Errorf("client state not found: %s", client.ClientId)
        }

        // 检查客户端是否过期
        if clientState.IsFrozen() || clientState.IsExpired() {
            log.Printf("Client %s needs recovery", client.ClientId)
            // 触发客户端恢复
        }
    }

    // 恢复所有通道
    channels := ibcKeeper.ChannelKeeper.GetAllChannels(ctx)
    for _, channel := range channels {
        if channel.State != ibcchannel.OPEN {
            log.Printf("Channel %s is in state %s, attempting recovery",
                channel.ChannelId, channel.State.String())
        }
    }

    return nil
}

手动恢复步骤:

# 步骤 1:验证链已成功升级
msgd status | jq '.SyncInfo.latest_block_height'
msgd query upgrade applied v2.0.0

# 步骤 2:更新 IBC 轻客户端
msgd tx ibc client update 07-tendermint-0 \
  --from validator --gas auto --fees 2500umsg

# 步骤 3:验证连接状态
msgd query ibc connection end connection-0

# 步骤 4:恢复通道
msgd tx ibc channel open-try transfer channel-0 \
  --from validator --gas auto --fees 2500umsg

# 步骤 5:验证代币转账恢复正常
msgd query bank balances msg1xyz...

3.5 升级后的 IBC 状态验证

3.5.1 一致性检查

// 升级后状态一致性验证
func VerifyIBCStateConsistency(ctx sdk.Context, ibcKeeper *ibckeeper.Keeper) error {
    checks := []struct {
        name string
        fn   func() error
    }{
        {"client-state", func() error {
            return verifyAllClients(ctx, ibcKeeper)
        }},
        {"connection-state", func() error {
            return verifyAllConnections(ctx, ibcKeeper)
        }},
        {"channel-state", func() error {
            return verifyAllChannels(ctx, ibcKeeper)
        }},
        {"packet-commitments", func() error {
            return verifyPacketCommitments(ctx, ibcKeeper)
        }},
    }

    for _, check := range checks {
        if err := check.fn(); err != nil {
            return fmt.Errorf("%s check failed: %w", check.name, err)
        }
    }

    return nil
}

3.5.2 IBC 状态校验指标

指标 正常值 警告阈值 告警阈值
活跃客户端数 > 0 0 0
OPEN 通道比例 100% < 95% < 80%
待处理数据包 0 < 100 > 1000
客户端过期数 0 > 1 > 5
连接平均延迟 < 5s < 30s > 60s

4. 轻客户端冻结与恢复

4.1 轻客户端冻结原因

轻客户端(Light Client)冻结是 IBC 协议中最常见的故障模式之一。对于 msg-chain-1,可能的冻结原因包括:

4.1.1 TrustingPeriod 过期

这是最常见的冻结原因。轻客户端的 TrustingPeriod 设置为 336 小时(14 天),如果在客户端更新时间窗口内未更新,客户端将进入过期状态。

过期条件:

当前时间 - 最新客户端更新时间 > TrustingPeriod(336小时)

触发场景:

4.1.2 检测到双重签名(Misbehaviour)

如果轻客户端检测到验证人集在相同高度签署了不同的区块,客户端将立即冻结。

误行为证据结构:

type Evidence struct {
    ClientId       string
    Misbehaviour   *Misbehaviour
    // v8 新增
    ClientState    *ClientState
    ConsensusState *ConsensusState
}

4.1.3 链升级引起的客户端不兼容

当 msg-chain-1 升级到新的 SDK 版本时,旧版本的轻客户端可能无法理解新版本的状态证明。

4.2 客户端冻结检测

4.2.1 自动检测

# 查询所有客户端状态
msgd query ibc client states --count 100

# 查询特定客户端详细信息
msgd query ibc client state 07-tendermint-0 --verbose

# 检查客户端过期状态
msgd query ibc client status 07-tendermint-0

4.2.2 监控告警

# IBC 客户端监控脚本示例
import requests
import json
import time

def check_client_health():
    endpoint = "http://localhost:26657"

    # 查询所有 IBC 客户端
    response = requests.post(
        f"{endpoint}/ibc/core/client/v1/client_states",
        json={}
    )

    clients = response.json().get("client_states", [])

    for client in clients:
        client_id = client["client_id"]
        status = client["status"]

        if status != "Active":
            alert_msg = f"ALERT: IBC Client {client_id} is {status}"
            print(alert_msg)
            # 发送告警
            send_alert(alert_msg)

        # 检查剩余存活时间
        trusting_period = int(client["client_state"]["trusting_period"][:-1])
        last_update = int(client["client_state"]["last_update"])
        current_time = int(time.time())

        remaining = trusting_period - (current_time - last_update)
        if remaining < 86400:  # 少于 24 小时
            warn_msg = f"WARN: Client {client_id} will expire in {remaining}s"
            print(warn_msg)
            send_warning(warn_msg)

def send_alert(message):
    # 集成告警系统(Slack/PagerDuty/Telegram)
    webhook_url = "https://hooks.alert.msgchain.org/ibc"
    requests.post(webhook_url, json={"text": message})

while True:
    check_client_health()
    time.sleep(300)  # 每 5 分钟检查一次

4.3 轻客户端恢复方法

4.3.1 标准恢复流程

当客户端过期时,可以通过提交最新的共识状态来恢复:

// 客户端恢复逻辑
func RestoreClient(
    ctx sdk.Context,
    clientKeeper ibcclient.Keeper,
    clientID string,
    height int64,
) error {
    // 1. 验证当前链状态
    consensusState, err := clientKeeper.GetSelfConsensusState(ctx, height)
    if err != nil {
        return fmt.Errorf("failed to get consensus state: %w", err)
    }

    // 2. 创建新的客户端状态(继承旧客户端的参数)
    oldState, ok := clientKeeper.GetClientState(ctx, clientID)
    if !ok {
        return fmt.Errorf("client state not found: %s", clientID)
    }

    newState := oldState
    newState.UpdateHeight(height)
    newState.ResetFrozen()

    // 3. 更新客户端
    clientKeeper.SetClientState(ctx, clientID, &newState)
    clientKeeper.SetClientConsensusState(ctx, clientID, height, consensusState)

    // 4. 验证恢复后的客户端
    status := clientKeeper.GetClientStatus(ctx, clientID)
    if status != ibcclient.Active {
        return fmt.Errorf("client recovery failed, status: %s", status)
    }

    return nil
}

4.3.2 通过治理提案恢复

// 治理提案:恢复轻客户端
type ClientUpdateProposal struct {
    Title       string
    Description string
    SubjectClientId string
    SubstituteClientId string
}

func HandleClientUpdateProposal(ctx sdk.Context,
    clientKeeper ibcclient.Keeper,
    proposal *ClientUpdateProposal) error {

    // 使用替代客户端恢复主客户端
    err := clientKeeper.UpdateClient(ctx,
        proposal.SubjectClientId,
        proposal.SubstituteClientId)
    if err != nil {
        return err
    }

    // 验证恢复结果
    status := clientKeeper.GetClientStatus(ctx, proposal.SubjectClientId)
    ctx.EventManager().EmitEvent(sdk.NewEvent(
        "client_recovered",
        sdk.NewAttribute("client_id", proposal.SubjectClientId),
        sdk.NewAttribute("status", status.String()),
    ))

    return nil
}

4.3.3 命令行恢复

# 方法一:使用替代客户端恢复
msgd tx ibc client update 07-tendermint-0 \
  --substitute 07-tendermint-1 \
  --from validator \
  --gas auto \
  --gas-adjustment 1.5 \
  --fees 5000umsg

# 方法二:通过治理提案恢复
msgd tx gov submit-proposal update-client 07-tendermint-0 07-tendermint-1 \
  --title "Recover IBC Client 07-tendermint-0" \
  --description "Client expired due to chain upgrade, recovering with substitute" \
  --deposit 100000000umsg \
  --from validator

# 方法三:直接更新客户端(如果链运行正常)
msgd tx ibc client update 07-tendermint-0 \
  --from relayer \
  --gas auto \
  --fees 2500umsg

4.4 Submit Misbehaviour

4.4.1 提交误行为证据

当检测到验证人的不当行为时,需及时提交证据以保护 IBC 连接安全:

# 准备证据文件
cat > misbehaviour.json << EOF
{
  "client_id": "07-tendermint-0",
  "misbehaviour": {
    "client_id": "07-tendermint-0",
    "height": {
      "revision_number": 1,
      "revision_height": 500000
    },
    "header_1": {
      "signed_header": { ... },
      "validator_set": { ... },
      "trusted_height": {
        "revision_number": 1,
        "revision_height": 499999
      },
      "trusted_validators": { ... }
    },
    "header_2": {
      "signed_header": { ... },
      "validator_set": { ... },
      "trusted_height": {
        "revision_number": 1,
        "revision_height": 499999
      },
      "trusted_validators": { ... }
    }
  }
}
EOF

# 提交证据
msgd tx ibc client submit-misbehaviour 07-tendermint-0 misbehaviour.json \
  --from validator \
  --gas auto \
  --gas-adjustment 1.5 \
  --fees 10000umsg

4.4.2 误行为处理后的恢复

误行为提交后,客户端会进入冻结状态(Frozen),需要通过治理提案解冻:

# 1. 提交误行为证据
msgd tx ibc client submit-misbehaviour 07-tendermint-0 evidence.json ...

# 2. 验证客户端已冻结
msgd query ibc client status 07-tendermint-0
# 输出: Frozen

# 3. 提交治理提案解冻
msgd tx gov submit-proposal \
  --type TextProposal \
  --title "Unfreeze IBC Client 07-tendermint-0" \
  --description "..." \
  --deposit 100000000umsg \
  --from validator

# 4. 提案通过后,用新客户端替代
msgd tx ibc client update 07-tendermint-0 \
  --substitute 07-tendermint-1 \
  --from validator

4.5 预防性维护

4.5.1 中继服务高可用部署

# docker-compose.yml
version: '3.8'
services:
  ibc-relayer-1:
    image: cosmos/relayer:v2.5.2
    command:
      - start
      - --config
      - /home/relayer/.relayer/config/config.yaml
    volumes:
      - ./relayer-1:/home/relayer/.relayer
    restart: always
    environment:
      - RELAYER_MNEMONIC=${RELAYER_1_MNEMONIC}
    healthcheck:
      test: ["CMD", "rly", "status"]
      interval: 60s
      timeout: 10s
      retries: 3

  ibc-relayer-2:
    image: cosmos/relayer:v2.5.2
    command:
      - start
      - --config
      - /home/relayer/.relayer/config/config.yaml
    volumes:
      - ./relayer-2:/home/relayer/.relayer
    restart: always
    environment:
      - RELAYER_MNEMONIC=${RELAYER_2_MNEMONIC}
    healthcheck:
      test: ["CMD", "rly", "status"]
      interval: 60s
      timeout: 10s
      retries: 3

  client-updater:
    image: msgchain/client-updater:v1.0.0
    command:
      - run
      - --interval
      - 60m
    environment:
      - MSGD_RPC_ENDPOINT=${MSGD_RPC_ENDPOINT}
      - RELAYER_MNEMONIC=${RELAYER_MNEMONIC}
    depends_on:
      - ibc-relayer-1
      - ibc-relayer-2

4.5.2 客户端存活自动更新脚本

#!/bin/bash
# client-keeper.sh - 自动保持 IBC 客户端存活

set -euo pipefail

CHAIN_ID="msg-chain-1"
NODE="http://localhost:26657"
FROM="relayer"
GAS_PRICES="1000000000umsg"
UPDATE_INTERVAL=240  # 4小时更新一次

update_clients() {
    echo "[$(date)] Checking IBC clients..."

    # 获取所有客户端
    clients=$(msgd query ibc client states \
        --node $NODE \
        --output json | jq -r '.client_states[] | select(.status=="Active") | .client_id')

    for client_id in $clients; do
        echo "  Updating client: $client_id"

        # 获取客户端详细信息
        client_info=$(msgd query ibc client state $client_id \
            --node $NODE \
            --output json)

        # 检查是否需要更新
        trusting_period=$(echo $client_info | jq -r '.client_state.trusting_period')
        # 转为秒
        case $trusting_period in
            *h) seconds=$(( ${trusting_period%h} * 3600 )) ;;
            *d) seconds=$(( ${trusting_period%d} * 86400 )) ;;
            *) seconds=1209600 ;;  # 默认14天
        esac

        # 如果剩余时间小于24小时,更新客户端
        msgd tx ibc client update $client_id \
            --from $FROM \
            --node $NODE \
            --chain-id $CHAIN_ID \
            --gas auto \
            --gas-prices $GAS_PRICES \
            --yes 2>&1 | tail -1

        sleep 2  # 避免速率限制
    done
}

# 主循环
while true; do
    update_clients
    echo "[$(date)] Next update in $UPDATE_INTERVAL minutes"
    sleep $(( UPDATE_INTERVAL * 60 ))
done

#### 4.5.3 批量客户端恢复脚本

```python
#!/usr/bin/env python3
"""
IBC 客户端批量恢复工具
用于升级后批量恢复所有过期客户端
"""
import subprocess
import json
import sys
import time
from typing import List, Dict

class IBCClientRecovery:
    def __init__(self, chain_id: str, node: str, from_key: str):
        self.chain_id = chain_id
        self.node = node
        self.from_key = from_key
        self.gas_prices = "1000000000umsg"

    def get_all_clients(self) -> List[Dict]:
        """获取所有 IBC 客户端"""
        result = subprocess.run([
            "msgd", "query", "ibc", "client", "states",
            "--node", self.node,
            "--output", "json"
        ], capture_output=True, text=True)

        data = json.loads(result.stdout)
        return data.get("client_states", [])

    def get_client_status(self, client_id: str) -> str:
        """检查客户端状态"""
        result = subprocess.run([
            "msgd", "query", "ibc", "client", "status", client_id,
            "--node", self.node,
            "--output", "json"
        ], capture_output=True, text=True)

        data = json.loads(result.stdout)
        return data.get("status", "Unknown")

    def update_client(self, client_id: str) -> bool:
        """更新指定客户端"""
        print(f"  Updating {client_id}...", end=" ", flush=True)

        # 检查是否需要替代客户端
        status = self.get_client_status(client_id)
        if status == "Frozen":
            # 冻结的客户端需要通过治理恢复
            print(f"SKIP (Frozen, needs governance)")
            return False

        result = subprocess.run([
            "msgd", "tx", "ibc", "client", "update", client_id,
            "--from", self.from_key,
            "--node", self.node,
            "--chain-id", self.chain_id,
            "--gas", "auto",
            "--gas-prices", self.gas_prices,
            "--yes", "-o", "json"
        ], capture_output=True, text=True, timeout=60)

        if result.returncode == 0:
            print("OK")
            return True
        else:
            print(f"FAILED: {result.stderr[:100]}")
            return False

    def recover_all_clients(self) -> Dict[str, bool]:
        """恢复所有需要恢复的客户端"""
        clients = self.get_all_clients()
        results = {}

        print(f"Found {len(clients)} clients")

        for client in clients:
            client_id = client["client_id"]
            status = self.get_client_status(client_id)

            if status != "Active":
                print(f"Client {client_id} needs recovery (status: {status})")
                results[client_id] = self.update_client(client_id)
                time.sleep(3)  # 避免速率限制
            else:
                print(f"Client {client_id} is active, skipping")
                results[client_id] = True

        return results

if __name__ == "__main__":
    recovery = IBCClientRecovery(
        chain_id="msg-chain-1",
        node="http://localhost:26657",
        from_key="relayer"
    )

    results = recovery.recover_all_clients()

    failed = [k for k, v in results.items() if not v]
    if failed:
        print(f"\nFailed clients: {failed}")
        sys.exit(1)
    else:
        print(f"\nAll {len(results)} clients recovered successfully")

5. 通道关闭与重启

5.1 通道关闭原因

IBC 通道关闭是比客户端冻结更严重的事件,通常需要治理层面的干预才能恢复。

5.1.1 通道关闭的触发条件

触发条件 关闭类型 恢复难度 常见性
对端链停止 ORDERED 通道自动关闭 高 中等
数据包超时 自动关闭(ORDERED) 高 常见
协议错误 自动关闭 高 罕见
安全事件 手动关闭 中 罕见
治理关闭 治理提案关闭 低 极少
通道升级失败 通道进入 CLOSING 状态 中 罕见

5.1.2 ORDERED vs UNORDERED 通道

// ORDERED 通道
// - 数据包按顺序处理
// - 任一数据包超时即关闭通道
// - 恢复成本高,需治理干预
//
// UNORDERED 通道
// - 数据包可乱序处理
// - 单个数据包超时不影响通道
// - 恢复简单,重新发送即可

// MSG Chain 的通道类型选择
const (
    TransferChannel      = "channel-0"  // UNORDERED(推荐)
    ICAControllerChannel = "channel-1"  // ORDERED(必要)
    ICAMasterChannel     = "channel-2"  // ORDERED(必要)
)

5.1.3 通道关闭的状态转移

OPEN ──► CLOSING ──► CLOSED ──► 治理恢复 ──► OPEN

途中可能的停滞状态:
├── CLOSING: 等待对端确认关闭
├── CLOSED: 已确认关闭,等待恢复
├── INIT: 通道初始化(未完成握手)
└── TRYOPEN: 握手进行中

5.2 通道关闭的影响分析

5.2.1 对应用的影响

通道关闭对不同的 IBC 应用有不同的影响:

ICS-20 Transfer:

ICS-27 Interchain Accounts:

ICS-721 NFT Transfer:

5.2.2 资产安全分析

通道关闭时的资产安全性:

// 通道关闭时的资产状态分析
func AnalyzeChannelCloseAssetImpact(
    ctx sdk.Context,
    transferKeeper ibctransfer.Keeper,
    channelID string,
) (*AssetImpact, error) {
    // 1. 统计该通道的所有待处理数据包
    pendingPackets := transferKeeper.GetAllPendingSendPackets(ctx, channelID)

    // 2. 统计锁定的 IBC 代币
    escrowBalance := transferKeeper.GetEscrowBalance(ctx, channelID)

    // 3. 检查代币轨迹
    denomTraces := transferKeeper.GetAllDenomTraces(ctx)

    return &AssetImpact{
        PendingPackets: len(pendingPackets),
        EscrowBalance:  escrowBalance,
        ActiveDenoms:   len(denomTraces),
    }, nil
}

资产状态总结:

资产状态 安全吗? 说明
已在目标链 ✅ 安全 资产已在目标地址
在 Escrow 中 ✅ 安全 通道恢复后可取出
待发送数据包 ⚠️ 风险 可能超时退回
待接收数据包 ⚠️ 风险 需通道恢复后接收

5.3 通道恢复流程

5.3.1 治理恢复通道

通道关闭后,最可靠的恢复方式是通过治理提案:

方案一:通道重开治理提案

type ReopenChannelProposal struct {
    Title       string
    Description string
    PortId      string
    ChannelId   string
}

func NewReopenChannelHandler(
    channelKeeper ibcchannel.Keeper,
    connectionKeeper ibcconnection.Keeper,
    clientKeeper ibcclient.Keeper,
) govv1.Handler {
    return func(ctx sdk.Context, content govv1.Authority) error {
        proposal, ok := content.(*ReopenChannelProposal)
        if !ok {
            return fmt.Errorf("unexpected proposal type")
        }

        // 1. 验证连接状态
        channel, found := channelKeeper.GetChannel(ctx, proposal.PortId, proposal.ChannelId)
        if !found {
            return fmt.Errorf("channel not found")
        }

        if channel.State != ibcchannel.CLOSED {
            return fmt.Errorf("channel is not in CLOSED state")
        }

        // 2. 验证对端客户端活跃
        connection, found := connectionKeeper.GetConnection(ctx, channel.ConnectionHops[0])
        if !found {
            return fmt.Errorf("connection not found")
        }

        clientState := clientKeeper.GetClientState(ctx, connection.ClientId)
        if clientState.IsFrozen() || clientState.IsExpired() {
            return fmt.Errorf("counterparty client is not active")
        }

        // 3. 重开通道
        channel.State = ibcchannel.OPEN
        channelKeeper.SetChannel(ctx, proposal.PortId, proposal.ChannelId, channel)

        // 4. 发送事件
        ctx.EventManager().EmitEvent(sdk.NewEvent(
            "channel_reopened",
            sdk.NewAttribute("port_id", proposal.PortId),
            sdk.NewAttribute("channel_id", proposal.ChannelId),
        ))

        return nil
    }
}

命令行操作:

# 1. 确认通道状态为 CLOSED
msgd query ibc channel end transfer channel-0

# 2. 提交通道恢复治理提案
msgd tx gov submit-proposal \
  --title "Reopen IBC Channel transfer/channel-0" \
  --description "\
    ## Summary
    Reopen IBC channel transfer/channel-0 on MSG Chain (msg-chain-1)
    after it was closed due to chain upgrade.

    ## Background
    During the v2.0.0 upgrade at height 5000000, some packets timed out
    causing the ORDERED channel to close.

    ## Impact Assessment
    - Escrow balance: 1,000,000 MSG tokens locked
    - Pending packets: 0 (all timed out properly)
    - Recovery plan: Reopen and restart normal operations

    ## Verification
    - Counterparty client: Active
    - Connection: Active
    - Governance authority: msg10d07y265gmmuvt4z0w9aw880jnsr700jm9d9zw
  " \
  --deposit 100000000umsg \
  --from validator

# 3. 提案通过后验证通道恢复
msgd query ibc channel end transfer channel-0 --output json | jq '.channel.state'

# 4. 测试转账
msgd tx ibc transfer transfer channel-0 msg1recipient... 1000umsg \
  --from testuser --gas auto --fees 2500umsg

5.3.2 手动 Reopen 流程

在某些情况下,可以通过治理绕过的手动流程恢复通道:

条件:

手动恢复流程:

#!/bin/bash
# channel-recover.sh - 手动通道恢复脚本

set -euo pipefail

CHAIN_ID="msg-chain-1"
NODE="http://localhost:26657"
FROM="admin"

PORT_ID=${1:-"transfer"}
CHANNEL_ID=${2:-"channel-0"}

echo "=== Channel Recovery Script ==="
echo "Chain: $CHAIN_ID"
echo "Channel: $PORT_ID/$CHANNEL_ID"
echo ""

# 步骤 1: 检查通道状态
echo "[1/6] Checking channel state..."
CHANNEL_INFO=$(msgd query ibc channel end "$PORT_ID" "$CHANNEL_ID" \
    --node "$NODE" --output json 2>/dev/null || echo "{}")
CHANNEL_STATE=$(echo "$CHANNEL_INFO" | jq -r '.channel.state')

echo "  Channel state: $CHANNEL_STATE"

if [ "$CHANNEL_STATE" != "STATE_CLOSED" ]; then
    if [ "$CHANNEL_STATE" == "STATE_OPEN" ]; then
        echo "  Channel is already OPEN, no recovery needed"
        exit 0
    fi
    echo "  ERROR: Unexpected channel state"
    exit 1
fi

# 步骤 2: 验证连接状态
echo "[2/6] Verifying connection..."
CONNECTION_ID=$(echo "$CHANNEL_INFO" | jq -r '.channel.connection_hops[0]')
echo "  Connection: $CONNECTION_ID"

CONNECTION_INFO=$(msgd query ibc connection end "$CONNECTION_ID" \
    --node "$NODE" --output json)
CONNECTION_STATE=$(echo "$CONNECTION_INFO" | jq -r '.connection.state')
echo "  Connection state: $CONNECTION_STATE"

if [ "$CONNECTION_STATE" != "STATE_OPEN" ]; then
    echo "  ERROR: Connection is not OPEN"
    exit 1
fi

# 步骤 3: 验证客户端
echo "[3/6] Verifying client..."
CLIENT_ID=$(echo "$CONNECTION_INFO" | jq -r '.connection.client_id')
echo "  Client: $CLIENT_ID"

CLIENT_STATUS=$(msgd query ibc client status "$CLIENT_ID" \
    --node "$NODE" --output json | jq -r '.status')
echo "  Client status: $CLIENT_STATUS"

if [ "$CLIENT_STATUS" != "Active" ]; then
    echo "  ERROR: Client is not Active"
    exit 1
fi

# 步骤 4: 检查待处理数据包
echo "[4/6] Checking pending packets..."
PENDING_PACKETS=$(msgd query ibc channel packets "$PORT_ID" "$CHANNEL_ID" \
    --node "$NODE" --output json 2>/dev/null | jq '.packets | length')
echo "  Pending packets: $PENDING_PACKETS"

if [ "$PENDING_PACKETS" -gt 0 ]; then
    echo "  WARNING: There are pending packets"
fi

# 步骤 5: 执行通道恢复(通过治理提案)
echo "[5/6] Submitting governance proposal for channel recovery..."
RESULT=$(msgd tx gov submit-proposal \
    --title "Reopen IBC Channel $PORT_ID/$CHANNEL_ID" \
    --description "Manual recovery of $PORT_ID/$CHANNEL_ID after chain upgrade" \
    --type TextProposal \
    --deposit 100000000umsg \
    --from "$FROM" \
    --node "$NODE" \
    --chain-id "$CHAIN_ID" \
    --gas auto \
    --gas-prices "1000000000umsg" \
    --yes -o json 2>&1)

TX_HASH=$(echo "$RESULT" | jq -r '.txhash // "unknown"')
echo "  Governance proposal submitted. TX: $TX_HASH"

# 步骤 6: 等待提案通过并恢复
echo "[6/6] Waiting for proposal to pass..."
echo "  Monitor: msgd query tx $TX_HASH"
echo "  After passage, run: msgd tx gov vote [proposal_id] yes"

echo ""
echo "=== Channel Recovery Initiated ==="

5.4 通道恢复后的验证

5.4.1 功能性验证

func VerifyChannelAfterRecovery(
    ctx sdk.Context,
    channelKeeper ibcchannel.Keeper,
    portID, channelID string,
) error {
    // 1. 验证通道状态
    channel, found := channelKeeper.GetChannel(ctx, portID, channelID)
    if !found {
        return fmt.Errorf("channel not found after recovery")
    }

    if channel.State != ibcchannel.OPEN {
        return fmt.Errorf("channel not OPEN after recovery: %s", channel.State)
    }

    // 2. 验证序列号连续
    nextSeqSend, _ := channelKeeper.GetNextSequenceSend(ctx, portID, channelID)
    nextSeqRecv, _ := channelKeeper.GetNextSequenceRecv(ctx, portID, channelID)

    if nextSeqSend < nextSeqRecv {
        return fmt.Errorf("sequence inconsistency: send=%d recv=%d",
            nextSeqSend, nextSeqRecv)
    }

    // 3. 验证没有残留的承诺数据
    commitments := channelKeeper.GetAllPacketCommitmentsAtChannel(ctx, portID, channelID)
    if len(commitments) > 0 {
        return fmt.Errorf("found %d stale packet commitments", len(commitments))
    }

    return nil
}

5.4.2 端到端测试

#!/bin/bash
# e2e-channel-test.sh

set -euo pipefail

CHAIN_ID="msg-chain-1"
NODE="http://localhost:26657"
FROM="tester"

echo "=== E2E Channel Test ==="

# 测试 ICS-20 转账
echo "1. Testing ICS-20 Transfer..."
TX_RESULT=$(msgd tx ibc transfer transfer channel-0 \
    "msg1test..." 1000umsg \
    --from "$FROM" \
    --node "$NODE" \
    --chain-id "$CHAIN_ID" \
    --gas auto \
    --gas-prices "1000000000umsg" \
    --yes -o json)

TX_HASH=$(echo "$TX_RESULT" | jq -r '.txhash')
echo "   Transfer submitted: $TX_HASH"

# 等待 2 个区块确认
sleep 6

# 验证转账结果
TX_STATUS=$(msgd query tx "$TX_HASH" --node "$NODE" --output json \
    | jq -r '.code')
if [ "$TX_STATUS" = "0" ]; then
    echo "   Transfer successful"
else
    echo "   Transfer failed"
    exit 1
fi

# 测试 ICS-27 ICA
echo "2. Testing ICS-27 ICA..."
# ... ICA 测试逻辑

# 测试 ICS-721 NFT
echo "3. Testing ICS-721 NFT..."
# ... NFT 测试逻辑

echo ""
echo "=== All E2E Tests Passed ==="

5.5 通道关闭预防

5.5.1 使用 UNORDERED 通道

在可能的情况下,优先使用 UNORDERED 通道:

// UNORDERED 通道的优势
// 1. 单个数据包超时不会关闭通道
// 2. 只需要重发超时的数据包
// 3. 恢复过程对用户透明

// 创建 UNORDERED 通道
msgd tx ibc channel open-init transfer ordered \
  --ordering ORDER_UNORDERED \
  --version ics20-1

5.5.2 监控通道健康

#!/usr/bin/env python3
"""
通道健康监控服务
"""
import asyncio
import aiohttp
import json
from datetime import datetime

class ChannelHealthMonitor:
    def __init__(self, endpoints: list, alert_webhook: str):
        self.endpoints = endpoints
        self.alert_webhook = alert_webhook
        self.channel_states = {}

    async def check_channel(self, session, endpoint, port_id, channel_id):
        url = f"{endpoint}/ibc/core/channel/v1/channels/{channel_id}/ports/{port_id}"
        try:
            async with session.get(url) as resp:
                data = await resp.json()
                channel = data.get("channel", {})
                state = channel.get("state", "UNKNOWN")
                return {
                    "channel": f"{port_id}/{channel_id}",
                    "state": state,
                    "next_seq_send": channel.get("next_sequence_send", 0),
                    "next_seq_recv": channel.get("next_sequence_recv", 0),
                    "timestamp": datetime.utcnow().isoformat()
                }
        except Exception as e:
            return {
                "channel": f"{port_id}/{channel_id}",
                "state": "ERROR",
                "error": str(e),
                "timestamp": datetime.utcnow().isoformat()
            }

    async def check_all_channels(self):
        async with aiohttp.ClientSession() as session:
            tasks = []
            for endpoint in self.endpoints:
                # 查询所有通道
                list_url = f"{endpoint}/ibc/core/channel/v1/channels"
                try:
                    async with session.get(list_url) as resp:
                        channels_data = await resp.json()
                        channels = channels_data.get("channels", [])
                        for ch in channels:
                            tasks.append(
                                self.check_channel(
                                    session, endpoint,
                                    ch["port_id"], ch["channel_id"]
                                )
                            )
                except Exception as e:
                    print(f"Error listing channels: {e}")

            results = await asyncio.gather(*tasks)
            return results

    async def run(self):
        while True:
            results = await self.check_all_channels()

            for result in results:
                ch = result["channel"]
                state = result["state"]

                # 检测状态变化
                if ch in self.channel_states:
                    old_state = self.channel_states[ch]
                    if old_state != state and state in ["STATE_CLOSING", "STATE_CLOSED"]:
                        await self.send_alert(
                            f"Channel state change: {ch}: {old_state} -> {state}"
                        )

                self.channel_states[ch] = state

            await asyncio.sleep(60)  # 每分钟检查

    async def send_alert(self, message):
        async with aiohttp.ClientSession() as session:
            await session.post(
                self.alert_webhook,
                json={"text": f"[Channel Monitor] {message}"}
            )

if __name__ == "__main__":
    monitor = ChannelHealthMonitor(
        endpoints=[
            "http://localhost:1317",
            "http://localhost:1318"
        ],
        alert_webhook="https://hooks.alert.msgchain.org/channel"
    )
    asyncio.run(monitor.run())

6. 数据迁移与状态证明

6.1 升级后的状态一致性

6.1.1 状态一致性原则

链升级后,IBC 模块的状态必须与升级前的状态保持一致性:

升级前状态 ──► 升级处理函数 ──► 迁移 ──► 升级后状态

  一致性要求:
  ├── 所有轻客户端状态可验证
  ├── 所有连接状态可验证
  ├── 所有通道状态可验证
  ├── 所有数据包承诺可验证
  └── 所有代币面额轨迹可追溯

6.1.2 状态迁移验证器

// 升级后状态一致性验证器
type StateConsistencyVerifier struct {
    ibcKeeper *ibckeeper.Keeper
    cdc       codec.Codec
}

func (v *StateConsistencyVerifier) VerifyAll(ctx sdk.Context) error {
    checks := []struct {
        name string
        fn   func(sdk.Context) error
    }{
        {"client-consistency", v.verifyClientConsistency},
        {"connection-consistency", v.verifyConnectionConsistency},
        {"channel-consistency", v.verifyChannelConsistency},
        {"packet-consistency", v.verifyPacketConsistency},
        {"denom-consistency", v.verifyDenomConsistency},
        {"param-consistency", v.verifyParamConsistency},
    }

    for _, check := range checks {
        if err := check.fn(ctx); err != nil {
            return fmt.Errorf("%s: %w", check.name, err)
        }
    }

    return nil
}

func (v *StateConsistencyVerifier) verifyClientConsistency(ctx sdk.Context) error {
    clients := v.ibcKeeper.ClientKeeper.GetAllClients(ctx)
    for _, client := range clients {
        // 验证每个客户端都可解码
        clientState, ok := v.ibcKeeper.ClientKeeper.GetClientState(ctx, client.ClientId)
        if !ok {
            return fmt.Errorf("client state missing: %s", client.ClientId)
        }

        // 验证客户端状态的一致性
        if clientState.GetLatestHeight().IsZero() {
            return fmt.Errorf("client %s has zero height", client.ClientId)
        }

        // 验证客户端参数未在升级中损坏
        if clientState.GetTrustingPeriod() <= 0 {
            return fmt.Errorf("client %s has invalid trusting period", client.ClientId)
        }
    }
    return nil
}

func (v *StateConsistencyVerifier) verifyDenomConsistency(ctx sdk.Context) error {
    traces := v.ibcKeeper.TransferKeeper.GetAllDenomTraces(ctx)

    // 验证所有代币轨迹的完整性
    for _, trace := range traces {
        if trace.BaseDenom == "" {
            return fmt.Errorf("denom trace %s has empty base denom", trace.IBCDenom())
        }

        // 验证路径格式
        if len(trace.Path) > 0 {
            pathParts := strings.Split(trace.Path, "/")
            for _, part := range pathParts {
                if !strings.HasPrefix(part, "transfer/") && !strings.HasPrefix(part, "channel-") {
                    // 验证路径格式
                }
            }
        }
    }

    return nil
}

6.2 状态证明验证

6.2.1 Merkle 证明

IBC 协议依赖 Merkle 证明来验证跨链状态:

// 状态证明验证
func VerifyStateProof(
    proof *ibcexported.MerkleProof,
    stateRoot commitmenttypes.MerkleRoot,
    path commitmenttypes.MerklePath,
    value []byte,
) error {
    // 验证 Merkle 路径
    if err := proof.VerifyMembership(stateRoot, path, value); err != nil {
        return fmt.Errorf("proof verification failed: %w", err)
    }

    return nil
}

6.2.2 升级后的证明兼容性

升级可能导致证明格式变化,需要特别注意:

升级类型 证明影响 处理方式
IAVL 版本升级 Merkle 树结构不变 兼容
SDK 版本升级 Merkle 路径可能变化 需验证
CometBFT 升级 ProofSpec 不变 兼容
存储迁移 键值布局变化 需迁移

6.3 数据迁移策略

6.3.1 存储键值迁移

// IBC 存储键值迁移
func MigrateIBCStore(ctx sdk.Context, store sdk.KVStore, cdc codec.BinaryCodec) error {
    // 旧的键值前缀
    oldClientPrefix := []byte{0x02, 0x01}
    newClientPrefix := []byte{0x02, 0x02}

    // 迭代旧存储
    iterator := store.Iterator(oldClientPrefix, sdk.PrefixEndBytes(oldClientPrefix))
    defer iterator.Close()

    for ; iterator.Valid(); iterator.Next() {
        // 读取旧键值
        oldKey := iterator.Key()
        value := iterator.Value()

        // 计算新键
        newKey := append(newClientPrefix, oldKey[len(oldClientPrefix):]...)

        // 写入新存储
        store.Set(newKey, value)

        // 删除旧存储
        store.Delete(oldKey)
    }

    return nil
}

6.3.2 自动迁移脚本

#!/usr/bin/env python3
"""
IBC 状态迁移工具 - 用于链升级后的状态验证和修复
"""
import json
import sys
import requests
from typing import Dict, List, Optional

class IBCStateMigrator:
    def __init__(self, rpc_endpoint: str, rest_endpoint: str):
        self.rpc = rpc_endpoint
        self.rest = rest_endpoint

    def export_genesis(self, height: Optional[int] = None) -> Dict:
        """导出指定高度的创世状态"""
        params = {"height": height} if height else {}
        resp = requests.get(
            f"{self.rest}/cosmos/genesis/v1beta1/genesis",
            params=params
        )
        return resp.json()

    def validate_ibc_genesis(self, genesis: Dict) -> List[str]:
        """验证创世文件中的 IBC 状态"""
        errors = []

        app_state = genesis.get("app_state", {})
        ibc_state = app_state.get("ibc", {})

        client_genesis = ibc_state.get("client_genesis", {})
        connection_genesis = ibc_state.get("connection_genesis", {})
        channel_genesis = ibc_state.get("channel_genesis", {})

        # 验证客户端状态
        clients = client_genesis.get("clients", [])
        for client in clients:
            client_id = client.get("client_id", "")
            if not client_id:
                errors.append("Client missing client_id")

            client_state = client.get("client_state", {})
            if not client_state.get("trusting_period"):
                errors.append(f"{client_id}: missing trusting_period")

        # 验证连接状态
        connections = connection_genesis.get("connections", [])
        for conn in connections:
            conn_id = conn.get("id", "")
            if not conn_id:
                errors.append("Connection missing id")

        # 验证通道状态
        channels = channel_genesis.get("channels", [])
        for ch in channels:
            ch_id = ch.get("channel_id", "")
            if not ch_id:
                errors.append("Channel missing channel_id")

            state = ch.get("state", "")
            if state not in ["INIT", "TRYOPEN", "OPEN", "CLOSED"]:
                errors.append(f"{ch_id}: invalid state {state}")

        return errors

    def migrate_ibc_genesis(self, old_genesis: Dict,
                           target_version: str) -> Dict:
        """迁移 IBC 创世状态到目标版本"""
        genesis = json.loads(json.dumps(old_genesis))  # 深拷贝

        ibc_state = genesis.get("app_state", {}).get("ibc", {})

        # 添加新版本特有的字段
        client_genesis = ibc_state.get("client_genesis", {})
        params = client_genesis.get("params", {})

        if target_version >= "v8":
            # v8 新增的参数
            params["allowed_clients"] = [
                "07-tendermint",
                "06-solomachine",
                "09-localhost"
            ]

        if target_version >= "v9":
            # v9 新增的参数
            params["channel_upgrade_enabled"] = True
            params["max_concurrent_channel_upgrades"] = 10

        client_genesis["params"] = params
        ibc_state["client_genesis"] = client_genesis

        # 添加通道升级状态
        if target_version >= "v8":
            channel_genesis = ibc_state.get("channel_genesis", {})
            channel_genesis["next_channel_upgrade_sequences"] = []

        return genesis

    def compare_genesis(self, gen_a: Dict, gen_b: Dict) -> Dict:
        """比较两个创世文件的差异"""
        diffs = {}

        def compare_dict(d1: Dict, d2: Dict, path: str = ""):
            all_keys = set(list(d1.keys()) + list(d2.keys()))
            for key in all_keys:
                full_path = f"{path}.{key}" if path else key

                if key not in d1:
                    diffs[full_path] = {"action": "added", "value": d2[key]}
                elif key not in d2:
                    diffs[full_path] = {"action": "removed", "value": d1[key]}
                elif isinstance(d1[key], dict) and isinstance(d2[key], dict):
                    compare_dict(d1[key], d2[key], full_path)
                elif d1[key] != d2[key]:
                    diffs[full_path] = {
                        "action": "modified",
                        "old": d1[key],
                        "new": d2[key]
                    }

        compare_dict(gen_a, gen_b)
        return diffs

if __name__ == "__main__":
    migrator = IBCStateMigrator(
        rpc_endpoint="http://localhost:26657",
        rest_endpoint="http://localhost:1317"
    )

    # 导出升级前的创世状态
    old_genesis = migrator.export_genesis(height=4999999)

    # 验证旧状态
    errors = migrator.validate_ibc_genesis(old_genesis)
    if errors:
        print("Pre-upgrade validation errors:")
        for err in errors:
            print(f"  - {err}")

    # 迁移到 v8
    migrated = migrator.migrate_ibc_genesis(old_genesis, "v8")

    # 验证迁移后的状态
    errors = migrator.validate_ibc_genesis(migrated)
    if errors:
        print("Post-migration validation errors:")
        for err in errors:
            print(f"  - {err}")
    else:
        print("Migration validation passed")

    # 比较差异
    diffs = migrator.compare_genesis(old_genesis, migrated)
    print(f"\nGenesis changes: {len(diffs)}")
    for path, change in list(diffs.items())[:20]:
        print(f"  [{change['action']}] {path}")

6.4 数据完整性校验

6.4.1 哈希一致性检查

func VerifyDataIntegrity(
    store sdk.KVStore,
    expectedRoot []byte,
) error {
    // 计算 IBC 子存储的 Merkle 根
    actualRoot := computeIBCStoreMerkleRoot(store)

    if !bytes.Equal(actualRoot, expectedRoot) {
        return fmt.Errorf(
            "data integrity violation: expected %x, got %x",
            expectedRoot, actualRoot,
        )
    }

    return nil
}

func computeIBCStoreMerkleRoot(store sdk.KVStore) []byte {
    // 实现 IAVL 树的 Merkle 根计算
    // 实际实现会遍历存储并构建树
    return nil
}

6.4.2 一致性检查清单

升级后的 IBC 状态验证应包含以下检查:

□ 客户端数量与升级前一致
□ 所有客户端状态有效(未冻结、未过期)
□ 连接数量与升级前一致
□ 所有连接处于 OPEN 状态
□ 通道数量与升级前一致
□ 所有通道状态正确
□ 待处理数据包数量合理
□ 代币面额轨迹完整
□ Escrow 余额与代币轨迹一致
□ 模块参数符合预期

7. MSG Chain 的 IBC 升级历史与兼容性策略

7.1 MSG Chain 技术栈概览

组件 版本 说明
Chain ID msg-chain-1 主网
Cosmos SDK v0.47.x 即将升级到 v0.50.x
ibc-go v7.3.x 即将升级到 v8.x
CometBFT v0.37.x 与 SDK v0.47 配套
CosmWasm v1.5.x 智能合约支持
Bech32 前缀 msg 地址格式
Governance v1(SDK 47) 治理模块

7.2 升级历史

7.2.1 v1.0.0 - 主网上线

时间: 2025 Q1
升级内容:

升级要点:

7.2.2 v1.1.0 - 安全补丁

时间: 2025 Q2
升级内容:

升级类型: 补丁升级(无共识变更)
IBC 影响: 无
恢复操作: 无需操作

7.2.3 v1.2.0 - 功能增强

时间: 2025 Q3
升级内容:

升级类型: 功能升级
IBC 影响: 低(仅新增模块)
重要变更:

7.2.4 v1.3.0 - IBC 增强

时间: 2025 Q4
升级内容:

升级类型: 补丁升级
IBC 影响: 低
重要变更:

7.2.5 v2.0.0 - 重大升级(规划中)

时间: 2026 Q3
升级内容:

升级类型: 重大升级
IBC 影响: 高
重要变更:

7.3 兼容性策略

7.3.1 多版本兼容策略

MSG Chain 采用以下策略确保 IBC 兼容性:

策略一:渐进式升级

// 渐进式升级:先升级链,再升级 IBC 连接
type UpgradeSchedule struct {
    Phases []UpgradePhase
}

type UpgradePhase struct {
    Name           string
    Height         int64
    IBCCoordinated bool
    Duration       time.Duration
}

// 推荐的渐进式升级计划
var MSGChainUpgradeV2 = UpgradeSchedule{
    Phases: []UpgradePhase{
        {
            Name:           "genesis-migration",
            Height:         5000000,
            IBCCoordinated: false,
            Duration:       2 * time.Hour,
        },
        {
            Name:           "client-recovery",
            Height:         5000010,
            IBCCoordinated: true,
            Duration:       4 * time.Hour,
        },
        {
            Name:           "channel-recovery",
            Height:         5000020,
            IBCCoordinated: true,
            Duration:       24 * time.Hour,
        },
        {
            Name:           "full-operations",
            Height:         5000100,
            IBCCoordinated: true,
            Duration:       0,
        },
    },
}

策略二:向后兼容 API

// 为 IBC 查询提供向后兼容的 API 版本
type IBCQueryRouter struct {
    v7Handler ibckeeper.QueryServer
    v8Handler ibckeeper.QueryServer
}

func (r *IBCQueryRouter) Channel(ctx context.Context,
    req *ibcchannel.QueryChannelRequest) (*ibcchannel.QueryChannelResponse, error) {
    // 先尝试 v8 版本
    resp, err := r.v8Handler.Channel(ctx, req)
    if err != nil {
        // 回退到 v7 版本
        return r.v7Handler.Channel(ctx, req)
    }
    return resp, nil
}

策略三:连接健康检查

type ConnectionHealth struct {
    ConnectionID string
    ClientID     string
    Status       string
    LastUpdated  time.Time
    PacketLag    uint64
}

func CheckAllConnections(ctx sdk.Context,
    clientKeeper ibcclient.Keeper,
    connectionKeeper ibcconnection.Keeper,
    channelKeeper ibcchannel.Keeper) []ConnectionHealth {

    var health []ConnectionHealth

    connections := connectionKeeper.GetAllConnections(ctx)
    for _, conn := range connections {
        clientState := clientKeeper.GetClientState(ctx, conn.ClientId)

        h := ConnectionHealth{
            ConnectionID: conn.Id,
            ClientID:     conn.ClientId,
            Status:       "healthy",
        }

        if clientState.IsFrozen() {
            h.Status = "frozen"
        } else if clientState.IsExpired() {
            h.Status = "expired"
        }

        health = append(health, h)
    }

    return health
}

7.3.2 通道兼容性矩阵

对端链 通道 ID 端口 类型 版本 升级状态
Cosmos Hub channel-0 transfer UNORDERED ics20-1 ✅ 已升级
Osmosis channel-1 transfer UNORDERED ics20-1 ✅ 已升级
Neutron channel-2 transfer UNORDERED ics20-1 ✅ 已升级
Osmosis channel-3 icacontroller ORDERED ics27-1 ⚠️ 需升级
Neutron channel-4 icahost ORDERED ics27-1 ⚠️ 需升级

7.4 升级沟通机制

7.4.1 公告模板

# MSG Chain 升级公告:v2.0.0

## 概述
- **升级名称**: v2.0.0(IBC 重大升级)
- **升级高度**: 5,000,000
- **预计时间**: 2026-07-15 14:00 UTC
- **影响**: IBC 连接将暂停 2-4 小时

## 变更内容
### ibc-go v8 迁移
- [x] 通道升级支持
- [x] ICS-27 增强
- [x] API 重构
- [x] 存储格式变更

### 升级步骤
1. 验证人请在升级高度前更新二进制
2. 升级后执行状态迁移
3. 恢复 IBC 连接(详情见恢复指南)

## 对端链协调
| 对端链 | 协调状态 | 联系人 | 备注 |
|--------|---------|-------|------|
| Cosmos Hub | ✅ 已确认 | validator@cosmos | - |
| Osmosis | ✅ 已确认 | security@osmosis | - |
| Neutron | ⏳ 待确认 | - | - |

## 回退计划
如升级失败,验证人需在升级高度前恢复旧版本二进制。

7.4.2 协调通信

# 升级协调配置文件
upgrade_coordination:
  chains:
    - name: "cosmoshub-4"
      contacts:
        - email: "validators@cosmos.network"
        - discord: "Cosmos Validators"
      ibl_channels:
        - "channel-0"
        - "channel-1"
      upgrade_window: "48h"

    - name: "osmosis-1"
      contacts:
        - email: "security@osmosis.zone"
        - matrix: "#osmosis-validators"
      ibl_channels:
        - "channel-1"
        - "channel-3"
      upgrade_window: "24h"

  notification:
    - type: "discord"
      webhook: "https://discord.com/api/webhooks/..."
    - type: "telegram"
      bot_token: "${TELEGRAM_BOT_TOKEN}"
      chat_id: "-100..."

7.5 升级回退策略

7.5.1 回退条件

以下情况需要触发升级回退:

  1. 共识失败:升级后区块无法达成共识(AppHash 不匹配)
  2. 状态损坏:状态迁移导致数据丢失或不一致
  3. 严重 Bug:升级后的二进制存在关键安全漏洞
  4. IBC 断裂:升级导致所有 IBC 连接无法恢复
  5. 性能严重退化:TPS 下降超过 50%

7.5.2 回退步骤

#!/bin/bash
# rollback.sh - 升级回退脚本

set -euo pipefail

CHAIN_HOME="${HOME}/.msgd"
BACKUP_DIR="${CHAIN_HOME}/cosmovisor/backups"
ROLLBACK_HEIGHT=$1  # 回滚到的目标高度

echo "=== Rollback Procedure ==="
echo "Target height: $ROLLBACK_HEIGHT"
echo ""

# 步骤 1: 停止节点
echo "[1/6] Stopping node..."
systemctl stop msgd

# 步骤 2: 恢复旧二进制
echo "[2/6] Restoring old binary..."
OLD_BINARY="${BACKUP_DIR}/$(date +%Y%m%d)/bin/msgd"
if [ -f "$OLD_BINARY" ]; then
    cp "$OLD_BINARY" /usr/local/bin/msgd
    echo "  Restored old binary from: $OLD_BINARY"
else
    echo "  ERROR: Old binary not found at $OLD_BINARY"
    exit 1
fi

# 步骤 3: 恢复数据快照
echo "[3/6] Restoring data snapshot..."
SNAPSHOT="${BACKUP_DIR}/$(date +%Y%m%d)/data.tar.gz"
if [ -f "$SNAPSHOT" ]; then
    rm -rf "${CHAIN_HOME}/data"
    tar -xzf "$SNAPSHOT" -C "$CHAIN_HOME"
    echo "  Restored data from: $SNAPSHOT"
else
    echo "  ERROR: Snapshot not found at $SNAPSHOT"
    exit 1
fi

# 步骤 4: 重置到目标高度
echo "[4/6] Resetting to height $ROLLBACK_HEIGHT..."
msgd tendermint unsafe-reset-all \
    --height "$ROLLBACK_HEIGHT" \
    --home "$CHAIN_HOME"

# 步骤 5: 启动节点
echo "[5/6] Starting node..."
systemctl start msgd

# 等待节点同步
sleep 10
CURRENT_HEIGHT=$(msgd status | jq -r '.SyncInfo.latest_block_height')
echo "  Current height: $CURRENT_HEIGHT"

# 步骤 6: 验证状态
echo "[6/6] Verifying state..."
msgd query ibc client states --output json | jq '.client_states | length'
msgd query ibc channel channels --output json | jq '.channels | length'

echo ""
echo "=== Rollback Completed ==="

8. 跨链协调

8.1 升级沟通机制

8.1.1 协调层次

跨链升级协调涉及多个层次的沟通:

层次一:验证人层面
├── 升级通知(48h 前)
├── 升级细节说明
├── 二进制分发
└── 升级演习

层次二:对端链层面
├── 升级时间窗口协商
├── IBC 暂停协调
├── 通道恢复计划
└── 联系人信息确认

层次三:社区层面
├── 升级公告发布
├── 常见问题解答
├── 实时状态更新
└── 升级后回顾

8.1.2 跨链升级通知 API

// 跨链升级通知消息
type CrossChainUpgradeNotice struct {
    SourceChain    string `json:"source_chain"`
    UpgradeName    string `json:"upgrade_name"`
    UpgradeHeight  int64  `json:"upgrade_height"`
    UpgradeTime    string `json:"upgrade_time"`
    IBCChannels    []struct {
        ChannelID string `json:"channel_id"`
        PortID    string `json:"port_id"`
        Action    string `json:"action"` // pause | migrate | resume
    } `json:"ibc_channels"`
    Contacts       []struct {
        Name    string `json:"name"`
        Role    string `json:"role"`
        Contact string `json:"contact"`
    } `json:"contacts"`
}

// 通过 IBC 发送升级通知(使用标准通道)
func SendUpgradeNotice(
    ctx sdk.Context,
    ibcKeeper *ibckeeper.Keeper,
    channelID string,
    notice *CrossChainUpgradeNotice,
) error {
    data, err := json.Marshal(notice)
    if err != nil {
        return err
    }

    packet := ibcchannel.Packet{
        Sequence:       ibcKeeper.ChannelKeeper.GetNextSequenceSend(ctx, "upgrade", channelID),
        SourcePort:     "upgrade",
        SourceChannel:  channelID,
        DestinationPort: "upgrade",
        DestinationChannel: channelID,
        Data:           data,
        TimeoutHeight:  ibcclient.Height{RevisionNumber: 1, RevisionHeight: uint64(notice.UpgradeHeight + 1000)},
        TimeoutTimestamp: 0,
    }

    return ibcKeeper.ChannelKeeper.SendPacket(ctx, packet)
}

8.2 调度窗口

8.2.1 窗口选择工具

#!/usr/bin/env python3
"""
跨链升级调度窗口优化器
"""
from datetime import datetime, timedelta
from typing import List, Dict
import json

class UpgradeSlotOptimizer:
    def __init__(self, chain_configs: Dict):
        self.chains = chain_configs

    def calculate_optimal_window(
        self,
        source_chain: str,
        preferred_day: str = "Wednesday",
        preferred_hour_utc: int = 8,
        buffer_hours: int = 4,
        min_notice_hours: int = 48
    ) -> Dict:
        """计算最优升级窗口"""
        now = datetime.utcnow()

        # 计算最低通知时间
        min_notice_time = now + timedelta(hours=min_notice_hours)

        # 找到下一个首选日期
        days_of_week = ["Monday", "Tuesday", "Wednesday",
                       "Thursday", "Friday", "Saturday", "Sunday"]
        target_day = days_of_week.index(preferred_day)
        current_day = now.weekday()
        days_until = (target_day - current_day + 7) % 7

        proposed_date = now + timedelta(days=days_until)
        proposed_time = proposed_date.replace(
            hour=preferred_hour_utc,
            minute=0,
            second=0,
            microsecond=0
        )

        # 检查是否满足最小通知时间
        if proposed_time < min_notice_time:
            proposed_time += timedelta(weeks=1)

        # 检查冲突
        conflicts = self.check_conflicts(source_chain, proposed_time, buffer_hours)

        window_end = proposed_time + timedelta(hours=buffer_hours)

        return {
            "proposed_window": {
                "start": proposed_time.isoformat(),
                "end": window_end.isoformat(),
                "duration_hours": buffer_hours
            },
            "conflicts": conflicts,
            "participating_chains": self.get_participating_chains(source_chain)
        }

    def check_conflicts(
        self,
        source_chain: str,
        proposed_time: datetime,
        buffer_hours: int
    ) -> List[Dict]:
        """检查升级窗口冲突"""
        conflicts = []

        for chain_name, chain_info in self.chains.items():
            if chain_name == source_chain:
                continue

            if chain_info.get("planned_upgrades"):
                for upgrade in chain_info["planned_upgrades"]:
                    upgrade_start = datetime.fromisoformat(upgrade["window_start"])
                    upgrade_end = datetime.fromisoformat(upgrade["window_end"])

                    # 检查时间窗口重叠
                    window_start = proposed_time
                    window_end = proposed_time + timedelta(hours=buffer_hours)

                    if window_start < upgrade_end and window_end > upgrade_start:
                        conflicts.append({
                            "chain": chain_name,
                            "upgrade": upgrade["name"],
                            "conflict_window": {
                                "start": upgrade["window_start"],
                                "end": upgrade["window_end"]
                            }
                        })

        return conflicts

    def get_participating_chains(self, chain_id: str) -> List[Dict]:
        """获取参与升级协调的对端链"""
        chains = []
        for name, info in self.chains.items():
            if name != chain_id:
                chains.append({
                    "name": name,
                    "ibc_channels": info.get("ibc_channels", []),
                    "contact": info.get("contact", "unknown")
                })
        return chains

# 示例配置
chain_configs = {
    "msg-chain-1": {
        "name": "MSG Chain",
        "timezone": "UTC",
        "ibc_channels": ["channel-0", "channel-1", "channel-2"],
        "contact": "validators@msgchain.org",
        "planned_upgrades": []
    },
    "cosmoshub-4": {
        "name": "Cosmos Hub",
        "timezone": "UTC",
        "ibc_channels": ["channel-392"],
        "contact": "validators@cosmos.network",
        "planned_upgrades": [
            {
                "name": "v14",
                "window_start": "2026-08-01T10:00:00Z",
                "window_end": "2026-08-01T14:00:00Z"
            }
        ]
    },
    "osmosis-1": {
        "name": "Osmosis",
        "timezone": "UTC",
        "ibc_channels": ["channel-4"],
        "contact": "security@osmosis.zone",
        "planned_upgrades": []
    }
}

if __name__ == "__main__":
    optimizer = UpgradeSlotOptimizer(chain_configs)
    window = optimizer.calculate_optimal_window(
        source_chain="msg-chain-1",
        preferred_day="Wednesday",
        preferred_hour_utc=8,
        buffer_hours=4
    )
    print(json.dumps(window, indent=2))

8.2.2 推荐的调度策略

策略 描述 适用场景 风险等级
独立升级 不协调对端链,升级后客户端恢复 补丁升级 低
协调暂停 协调对端链暂停 IBC 活动 主版本升级 中
协调升级 同时对端链升级 重大升级 高
分阶段升级 逐步升级各连接 复杂网络 中

8.3 回退计划

8.3.1 回退决策矩阵

升级阶段 问题类型 决策 操作
升级前(< 100 区块) 二进制发布错误 延迟升级 重新发布二进制
升级中(0-10 区块) AppHash 不匹配 紧急回滚 恢复快照
升级后(10-1000 区块) 状态不一致 状态修复 执行迁移补丁
升级后(> 1000 区块) IBC 连接失败 通道恢复 治理恢复
升级后(> 1000 区块) 安全漏洞 紧急补丁 发布补丁版本

8.3.2 回退计划模板

# 升级回退计划 — MSG Chain v2.0.0

## 触发条件
- [ ] 升级后 30 分钟内未达到共识
- [ ] AppHash 与预期不匹配
- [ ] 超过 3 个验证人报告错误
- [ ] IBC 连接在 6 小时内无法恢复

## 回退步骤
1. **停止升级**:所有验证人恢复旧版本二进制
2. **数据恢复**:从 Cosmovisor 备份恢复
3. **状态回滚**:回滚到升级前的高度
4. **IBC 恢复**:验证连接和通道状态
5. **重新规划**:确定新的升级时间

## 联系方式
- 升级协调人:@coordinator (Discord)
- 技术负责人:@tech-lead (Telegram)
- 备份验证人:@backup-validator

## 升级后回顾
计划在升级完成后 48 小时内召开回顾会议。

8.4 跨链测试协调

8.4.1 联合测试计划

# 跨链升级联合测试计划
cross_chain_upgrade_test:
  participants:
    - chain: msg-chain-1
      validators: 3
      relayer: 2
    - chain: cosmoshub-4
      validators: 2
      relayer: 1
    - chain: osmosis-1
      validators: 2
      relayer: 1

  test_scenarios:
    - name: "basic-ibc-transfer"
      description: "升级前后测试代币转账"
      expected_duration: "30min"

    - name: "channel-recovery"
      description: "模拟通道关闭后恢复"
      expected_duration: "1h"

    - name: "stress-test"
      description: "升级后大量 IBC 交易测试"
      expected_duration: "2h"

  success_criteria:
    - all_chains_recovered: true
    - transfer_success_rate: ">99%"
    - max_downtime: "<4h"

9. 回滚策略

9.1 链分叉与 IBC 一致性

9.1.1 分叉对 IBC 的影响

链分叉是 IBC 协议面临的最严重挑战之一。当发生链分叉时:

正常情况:
区块 N ──► 区块 N+1 ──► 区块 N+2 ──► ...

分叉情况:
           ┌──► 区块 N+1 (A) ──► 区块 N+2 (A) ──► ...
区块 N ────┤
           └──► 区块 N+1 (B) ──► 区块 N+2 (B) ──► ...

IBC 影响:
- 轻客户端可能检测到冲突
- 跨链交易可能被回滚
- 需要治理协调恢复

9.1.2 分叉检测

// 分叉检测逻辑
func DetectFork(
    ctx sdk.Context,
    clientKeeper ibcclient.Keeper,
    clientID string,
    conflictingHeader ibctm.Header,
) (bool, error) {
    // 获取当前客户端状态
    clientState, ok := clientKeeper.GetClientState(ctx, clientID)
    if !ok {
        return false, fmt.Errorf("client state not found")
    }

    // 检查高度是否冲突
    height := conflictingHeader.GetHeight()
    existingConsensus, ok := clientKeeper.GetClientConsensusState(ctx, clientID, height)
    if !ok {
        return false, nil // 该高度没有已有共识状态,不是分叉
    }

    // 比较区块哈希
    existingHash := existingConsensus.GetHash()
    newHash := conflictingHeader.Header.GetLastBlockId().GetHash()

    if !bytes.Equal(existingHash, newHash) {
        // 检测到分叉!
        return true, nil
    }

    return false, nil
}

9.2 回滚后的交易重放

9.2.1 交易重放机制

当链回滚时,回滚高度以上的交易会丢失。重放机制确保这些交易不会永久丢失:

// 交易重放管理器
type TxReplayManager struct {
    mempool       *mempool.TxMempool
    eventStore    *EventStore
    replayQueue   *ReplayQueue
}

type ReplayQueue struct {
    // 存储回滚高度以上的交易
    transactions []sdk.Tx
    // 交易的 IBC 上下文
    ibcContexts  map[string]*IBCContext
}

type IBCContext struct {
    SourceChannel string
    SourcePort    string
    DestinationChannel string
    DestinationPort   string
    Sequence      uint64
    TimeoutHeight ibcclient.Height
}

func (m *TxReplayManager) CaptureTxsForReplay(ctx sdk.Context,
    fromHeight, toHeight int64) error {
    // 收集回滚范围内的所有交易
    for height := fromHeight; height <= toHeight; height++ {
        block, err := m.getBlock(height)
        if err != nil {
            return err
        }

        for _, tx := range block.Data.Txs {
            // 解码交易
            decodedTx, err := m.decodeTx(tx)
            if err != nil {
                continue
            }

            // 如果是 IBC 交易,记录上下文
            if containsIBCMsg(decodedTx) {
                context := extractIBCContext(decodedTx)
                m.replayQueue.ibcContexts[tx.Hash()] = context
            }

            m.replayQueue.transactions = append(
                m.replayQueue.transactions, decodedTx)
        }
    }

    return nil
}

func (m *TxReplayManager) ReplayTxs(ctx sdk.Context) (uint64, error) {
    var replayed uint64

    for _, tx := range m.replayQueue.transactions {
        // 验证交易是否仍然有效
        if err := m.validateTx(tx); err != nil {
            continue // 跳过无效交易
        }

        // 对于 IBC 交易,检查时序
        if context, ok := m.replayQueue.ibcContexts[tx.Hash()]; ok {
            if !m.checkIBCContext(ctx, context) {
                continue // IBC 上下文不再有效
            }
        }

        // 重新执行交易
        _, err := m.executeTx(ctx, tx)
        if err == nil {
            replayed++
        }
    }

    return replayed, nil
}

func (m *TxReplayManager) checkIBCContext(
    ctx sdk.Context, context *IBCContext) bool {
    // 检查通道是否仍然存在
    channel, found := m.channelKeeper.GetChannel(
        ctx, context.SourcePort, context.SourceChannel)
    if !found || channel.State != ibcchannel.OPEN {
        return false
    }

    // 检查序列号是否仍然有效
    // 如果序列号已被消耗,交易不能重放
    nextSeqSend, _ := m.channelKeeper.GetNextSequenceSend(
        ctx, context.SourcePort, context.SourceChannel)
    if context.Sequence < nextSeqSend {
        return false // 该序列号已被使用
    }

    return true
}

9.2.2 重放策略

交易类型 重放策略 安全考量 成功概率
普通转账 直接重放 余额检查 高
IBC 转账 重放(需检查序列号) 避免双重发送 中
IBC 接收 不重放(依赖源链) 无风险 N/A
ICA 交易 需检查 ICA 状态 避免重复执行 低
合约调用 需检查合约状态 避免状态冲突 中

9.3 回滚后的 IBC 恢复

9.3.1 完整的恢复流程

#!/bin/bash
# post-rollback-ibc-recovery.sh

set -euo pipefail

CHAIN_HOME="${HOME}/.msgd"
ROLLBACK_HEIGHT=$1
TIMEOUT=${2:-300}  # 超时时间(秒)

echo "=== Post-Rollback IBC Recovery ==="
echo "Rollback height: $ROLLBACK_HEIGHT"
echo ""

# 阶段一:初始验证
echo "Phase 1: Initial Verification"
echo "-----------------------------"

# 1.1 确认链运行正常
echo "  [1/6] Checking chain status..."
for i in $(seq 1 $TIMEOUT); do
    STATUS=$(msgd status --home "$CHAIN_HOME" 2>/dev/null | jq -r '.SyncInfo.latest_block_height' 2>/dev/null || echo "syncing")
    if [ "$STATUS" != "syncing" ] && [ "$STATUS" -gt "$ROLLBACK_HEIGHT" ]; then
        echo "    Chain resumed at height $STATUS"
        break
    fi
    if [ $i -eq $TIMEOUT ]; then
        echo "    ERROR: Chain did not resume within timeout"
        exit 1
    fi
    sleep 2
done

# 1.2 验证 AppHash
CURRENT_HASH=$(msgd status --home "$CHAIN_HOME" | jq -r '.SyncInfo.latest_app_hash')
echo "  [2/6] Current app hash: $CURRENT_HASH"

# 1.3 检查验证人集
VALIDATOR_COUNT=$(msgd query staking validators --home "$CHAIN_HOME" --output json | jq '.validators | length')
echo "  [3/6] Active validators: $VALIDATOR_COUNT"

# 阶段二:IBC 状态评估
echo ""
echo "Phase 2: IBC State Assessment"
echo "-----------------------------"

# 2.1 检查客户端
echo "  [4/6] Checking IBC clients..."
CLIENT_COUNT=$(msgd query ibc client states --home "$CHAIN_HOME" --output json 2>/dev/null | jq '.client_states | length' || echo "0")
echo "    Total clients: $CLIENT_COUNT"

# 获取过期客户端
EXPIRED_CLIENTS=$(msgd query ibc client states --home "$CHAIN_HOME" --output json 2>/dev/null | \
    jq -r '.client_states[] | select(.status != "Active") | .client_id' || echo "")
if [ -n "$EXPIRED_CLIENTS" ]; then
    echo "    WARNING: Expired clients detected:"
    for client in $EXPIRED_CLIENTS; do
        echo "      - $client"
    done
fi

# 2.2 检查连接
echo "  [5/6] Checking IBC connections..."
CONNECTION_COUNT=$(msgd query ibc connection connections --home "$CHAIN_HOME" --output json 2>/dev/null | jq '.connections | length' || echo "0")
echo "    Total connections: $CONNECTION_COUNT"

# 2.3 检查通道
echo "  [6/6] Checking IBC channels..."
CHANNEL_COUNT=$(msgd query ibc channel channels --home "$CHAIN_HOME" --output json 2>/dev/null | jq '.channels | length' || echo "0")
CLOSED_CHANNELS=$(msgd query ibc channel channels --home "$CHAIN_HOME" --output json 2>/dev/null | \
    jq -r '.channels[] | select(.state == "STATE_CLOSED") | .channel_id' || echo "")
echo "    Total channels: $CHANNEL_COUNT"
if [ -n "$CLOSED_CHANNELS" ]; then
    echo "    WARNING: Closed channels detected:"
    for ch in $CLOSED_CHANNELS; do
        echo "      - $ch"
    done
fi

# 阶段三:IBC 恢复
echo ""
echo "Phase 3: IBC Recovery"
echo "----------------------"

# 3.1 恢复过期客户端
if [ -n "$EXPIRED_CLIENTS" ]; then
    echo "  Restoring expired clients..."
    for client in $EXPIRED_CLIENTS; do
        echo "    Updating $client..."
        msgd tx ibc client update "$client" \
            --from validator \
            --home "$CHAIN_HOME" \
            --chain-id msg-chain-1 \
            --gas auto \
            --gas-prices "1000000000umsg" \
            --yes -o json 2>/dev/null || true
        sleep 3
    done
fi

# 3.2 恢复关闭的通道
if [ -n "$CLOSED_CHANNELS" ]; then
    echo "  Restoring closed channels (requires governance)..."
    for ch in $CLOSED_CHANNELS; do
        echo "    Channel $ch needs governance proposal for recovery"
        # 这里需要治理提案,不能自动恢复
    done
fi

# 阶段四:验证
echo ""
echo "Phase 4: Final Verification"
echo "---------------------------"

# 4.1 验证 IBC 转账功能
echo "  Testing IBC transfer..."
TEST_TX=$(msgd tx ibc transfer transfer channel-0 \
    "msg1test..." 1000umsg \
    --from tester \
    --home "$CHAIN_HOME" \
    --chain-id msg-chain-1 \
    --gas auto \
    --gas-prices "1000000000umsg" \
    --yes -o json 2>/dev/null || echo "failed")
echo "    Result: $TEST_TX"

# 4.2 汇总状态
echo ""
echo "=== Recovery Summary ==="
echo "Chain: msg-chain-1"
echo "Current height: $(msgd status --home "$CHAIN_HOME" 2>/dev/null | jq -r '.SyncInfo.latest_block_height')"
echo "Clients: $CLIENT_COUNT"
echo "Connections: $CONNECTION_COUNT"
echo "Channels: $CHANNEL_COUNT"
echo "Status: $([ -n "$EXPIRED_CLIENTS" ] || [ -n "$CLOSED_CHANNELS" ] && echo "RECOVERING" || echo "OK")"

9.3.2 IBC 回滚后的状态一致性验证

// 回滚后状态一致性验证
func PostRollbackConsistencyCheck(
    ctx sdk.Context,
    ibcKeeper *ibckeeper.Keeper,
    rollbackHeight int64,
) error {
    // 1. 验证所有客户端在回滚高度处有有效状态
    clients := ibcKeeper.ClientKeeper.GetAllClients(ctx)
    for _, client := range clients {
        clientState, ok := ibcKeeper.ClientKeeper.GetClientState(ctx, client.ClientId)
        if !ok {
            return fmt.Errorf("client %s state missing after rollback", client.ClientId)
        }

        // 客户端的最新高度不应高于回滚高度
        latestHeight := clientState.GetLatestHeight()
        if latestHeight.GetRevisionHeight() > uint64(rollbackHeight) {
            return fmt.Errorf(
                "client %s has height %d > rollback height %d",
                client.ClientId, latestHeight.GetRevisionHeight(), rollbackHeight,
            )
        }
    }

    // 2. 验证所有连接完整性
    connections := ibcKeeper.ConnectionKeeper.GetAllConnections(ctx)
    for _, conn := range connections {
        if conn.State != ibcconnection.OPEN {
            // 连接可能因回滚而断开
            log.Printf("Connection %s is in state %s after rollback", conn.Id, conn.State)
        }
    }

    // 3. 验证通道和序列号
    channels := ibcKeeper.ChannelKeeper.GetAllChannels(ctx)
    for _, channel := range channels {
        // 回滚后,序列号可能回退
        // 需要确保没有数据包序列号间隙
        nextSeqSend, _ := ibcKeeper.ChannelKeeper.GetNextSequenceSend(
            ctx, channel.PortId, channel.ChannelId)

        // 验证所有待处理数据包
        commitments := ibcKeeper.ChannelKeeper.GetAllPacketCommitmentsAtChannel(
            ctx, channel.PortId, channel.ChannelId)

        for _, commitment := range commitments {
            if commitment.Sequence >= nextSeqSend {
                return fmt.Errorf(
                    "commitment sequence %d >= next send sequence %d for channel %s/%s",
                    commitment.Sequence, nextSeqSend,
                    channel.PortId, channel.ChannelId,
                )
            }
        }
    }

    return nil
}

9.4 IBC 时序与回滚恢复

9.4.1 时序验证

回滚后,IBC 的时序验证是确保数据包不会重复的关键:

// 回滚后的时序验证
func TimingVerificationAfterRollback(
    ctx sdk.Context,
    channelKeeper ibcchannel.Keeper,
    portID, channelID string,
) error {
    channel, found := channelKeeper.GetChannel(ctx, portID, channelID)
    if !found {
        return fmt.Errorf("channel not found")
    }

    // 获取升级后的序列号范围
    nextSeqSend, _ := channelKeeper.GetNextSequenceSend(ctx, portID, channelID)
    nextSeqRecv, _ := channelKeeper.GetNextSequenceRecv(ctx, portID, channelID)

    // 获取所有未完成的承诺
    commitments := channelKeeper.GetAllPacketCommitmentsAtChannel(ctx, portID, channelID)

    for _, commitment := range commitments {
        // 检查超时高度
        timeoutHeight := commitment.TimeoutHeight

        // 如果回滚后超时高度已过,数据包应被视为超时
        if timeoutHeight.GetRevisionHeight() > 0 &&
           timeoutHeight.GetRevisionHeight() <= uint64(ctx.BlockHeight()) {
            log.Printf("Packet %d timed out after rollback", commitment.Sequence)
        }

        // 检查超时时间戳
        timeoutTimestamp := commitment.TimeoutTimestamp
        if timeoutTimestamp > 0 && timeoutTimestamp < uint64(ctx.BlockTime().UnixNano()) {
            log.Printf("Packet %d timestamp timeout after rollback", commitment.Sequence)
        }
    }

    return nil
}

9.4.2 数据包去重

回滚后可能出现的包重复问题及处理策略:

场景 问题 解决方案 自动处理
已确认包被回滚 包重复 检查序列号是否已使用 ✅
已超时包被回滚 包状态错乱 重新评估超时条件 ✅
待确认包被回滚 包丢失 重新发送 ✅
ICA 执行回滚 状态不一致 需治理协调 ❌

10. 升级自动化

10.1 CI/CD 升级测试

10.1.1 GitHub Actions 工作流

# .github/workflows/ibc-upgrade-test.yml
name: IBC Upgrade Test

on:
  pull_request:
    branches: [main, release/*]
  workflow_dispatch:
    inputs:
      upgrade_type:
        description: 'Type of upgrade to test'
        required: true
        default: 'minor'
        type: choice
        options:
          - patch
          - minor
          - major

env:
  GO_VERSION: '1.21'
  CHAIN_ID: 'msg-chain-test-1'
  IBC_VERSION: 'v8.3.0'

jobs:
  upgrade-compatibility:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Setup Go
        uses: actions/setup-go@v5
        with:
          go-version: ${{ env.GO_VERSION }}

      - name: Build old binary
        run: |
          git checkout $(git describe --tags --abbrev=0)
          make build
          cp build/msgd /tmp/msgd-old

      - name: Build new binary
        run: |
          git checkout ${{ github.sha }}
          make build
          cp build/msgd /tmp/msgd-new

      - name: Setup test environment
        run: |
          # 创建测试目录
          mkdir -p /tmp/ibc-test

          # 初始化测试链
          /tmp/msgd-old init ibc-test-node \
            --chain-id ${{ env.CHAIN_ID }} \
            --home /tmp/ibc-test/node1

          # 添加测试账户
          /tmp/msgd-old keys add test-validator \
            --keyring-backend test \
            --home /tmp/ibc-test/node1

          # 创建创世交易
          /tmp/msgd-old genesis add-genesis-account \
            $(/tmp/msgd-old keys show test-validator -a --keyring-backend test --home /tmp/ibc-test/node1) \
            1000000000umsg \
            --home /tmp/ibc-test/node1

          /tmp/msgd-old genesis gentx test-validator \
            100000000umsg \
            --keyring-backend test \
            --home /tmp/ibc-test/node1 \
            --chain-id ${{ env.CHAIN_ID }}

          /tmp/msgd-old genesis collect-gentxs \
            --home /tmp/ibc-test/node1

      - name: Start old chain
        run: |
          /tmp/msgd-old start \
            --home /tmp/ibc-test/node1 \
            --rpc.laddr tcp://0.0.0.0:26657 \
            --grpc.address 0.0.0.0:9090 \
            > /tmp/ibc-test/node1.log 2>&1 &

          for i in $(seq 1 30); do
            if /tmp/msgd-old status --home /tmp/ibc-test/node1 2>/dev/null; then
              echo "Chain started"
              break
            fi
            sleep 2
          done

      - name: Deploy upgrade proposal
        run: |
          CURRENT_HEIGHT=$(/tmp/msgd-old status \
            --home /tmp/ibc-test/node1 | \
            jq -r '.SyncInfo.latest_block_height')

          UPGRADE_HEIGHT=$((CURRENT_HEIGHT + 20))

          /tmp/msgd-old tx gov submit-proposal \
            --type SoftwareUpgrade \
            --title "Test Upgrade" \
            --description "Test IBC upgrade compatibility" \
            --upgrade-name "v2.0.0" \
            --upgrade-height $UPGRADE_HEIGHT \
            --upgrade-info "{}" \
            --deposit 100000000umsg \
            --from test-validator \
            --keyring-backend test \
            --home /tmp/ibc-test/node1 \
            --chain-id ${{ env.CHAIN_ID }} \
            --gas auto \
            --gas-prices "1000000000umsg" \
            --yes

          sleep 3

          /tmp/msgd-old tx gov vote 1 yes \
            --from test-validator \
            --keyring-backend test \
            --home /tmp/ibc-test/node1 \
            --chain-id ${{ env.CHAIN_ID }} \
            --gas auto \
            --gas-prices "1000000000umsg" \
            --yes

      - name: Perform upgrade
        run: |
          UPGRADE_HEIGHT=$((CURRENT_HEIGHT + 20))

          echo "Waiting for upgrade height $UPGRADE_HEIGHT..."
          while true; do
            CURRENT_HEIGHT=$(/tmp/msgd-old status \
              --home /tmp/ibc-test/node1 2>/dev/null | \
              jq -r '.SyncInfo.latest_block_height' 2>/dev/null || echo "0")

            if [ "$CURRENT_HEIGHT" -ge "$UPGRADE_HEIGHT" ]; then
              echo "Upgrade height reached"
              break
            fi
            sleep 1
          done

          kill %1 2>/dev/null || true
          sleep 5

          cp /tmp/msgd-new /tmp/ibc-test/node1/cosmovisor/current/bin/msgd

          /tmp/msgd-new start \
            --home /tmp/ibc-test/node1 \
            --rpc.laddr tcp://0.0.0.0:26657 \
            --grpc.address 0.0.0.0:9090 \
            > /tmp/ibc-test/node1-new.log 2>&1 &

          for i in $(seq 1 30); do
            if /tmp/msgd-new status --home /tmp/ibc-test/node1 2>/dev/null; then
              echo "New chain started"
              break
            fi
            sleep 2
          done

      - name: Verify IBC state
        run: |
          echo "Checking IBC module..."

          CLIENTS=$(/tmp/msgd-new query ibc client states \
            --home /tmp/ibc-test/node1 --output json 2>/dev/null | \
            jq '.client_states | length' || echo "0")
          echo "IBC clients: $CLIENTS"

          CONNECTIONS=$(/tmp/msgd-new query ibc connection connections \
            --home /tmp/ibc-test/node1 --output json 2>/dev/null | \
            jq '.connections | length' || echo "0")
          echo "IBC connections: $CONNECTIONS"

          CHANNELS=$(/tmp/msgd-new query ibc channel channels \
            --home /tmp/ibc-test/node1 --output json 2>/dev/null | \
            jq '.channels | length' || echo "0")
          echo "IBC channels: $CHANNELS"

          PARAMS=$(/tmp/msgd-new query ibc params \
            --home /tmp/ibc-test/node1 --output json 2>/dev/null || echo "{}")
          echo "IBC params: $PARAMS"

      - name: Run IBC integration tests
        run: |
          echo "Running IBC integration tests..."

          /tmp/msgd-new query ibc transfer params \
            --home /tmp/ibc-test/node1 || \
            echo "Transfer module check: PASS"

          /tmp/msgd-new query ibc channel channels \
            --home /tmp/ibc-test/node1 > /dev/null && \
            echo "Channel query: PASS"

          echo "IBC integration tests completed"

      - name: Cleanup
        if: always()
        run: |
          kill %1 %2 2>/dev/null || true
          rm -rf /tmp/ibc-test /tmp/msgd-old /tmp/msgd-new

10.1.2 本地测试脚本

#!/bin/bash
# local-ibc-upgrade-test.sh

set -euo pipefail

TEST_DIR="/tmp/ibc-upgrade-test-$(date +%s)"
CHAIN_ID="msg-chain-test-1"
MONIKER="test-validator"

echo "=== Local IBC Upgrade Test ==="
echo "Test directory: $TEST_DIR"

# 清理
cleanup() {
    echo "Cleaning up..."
    kill %1 %2 2>/dev/null || true
    rm -rf "$TEST_DIR"
}
trap cleanup EXIT

# 准备
echo ""
echo "Step 1: Setup test environment"
mkdir -p "$TEST_DIR"

# 检查依赖
for cmd in msgd jq curl; do
    if ! command -v $cmd &> /dev/null; then
        echo "ERROR: $cmd not found"
        exit 1
    fi
done

# 获取当前版本
CURRENT_VERSION=$(msgd version 2>/dev/null || echo "unknown")
echo "Current msgd version: $CURRENT_VERSION"

# 初始化测试链
msgd init "$MONIKER" \
    --chain-id "$CHAIN_ID" \
    --home "$TEST_DIR/node1" 2>/dev/null

# 配置最小 gas 价格
sed -i 's/minimum-gas-prices = ""/minimum-gas-prices = "1000000000umsg"/' \
    "$TEST_DIR/node1/config/app.toml"

# 添加测试密钥
echo "lock guitar ... virtue" | \
    msgd keys add validator \
    --recover \
    --keyring-backend test \
    --home "$TEST_DIR/node1" 2>/dev/null

VALIDATOR_ADDR=$(msgd keys show validator \
    -a --keyring-backend test --home "$TEST_DIR/node1")

# 添加创世账户
msgd genesis add-genesis-account \
    "$VALIDATOR_ADDR" 1000000000000umsg \
    --home "$TEST_DIR/node1"

# 创建创世交易
msgd genesis gentx validator \
    500000000000umsg \
    --keyring-backend test \
    --home "$TEST_DIR/node1" \
    --chain-id "$CHAIN_ID" 2>/dev/null

msgd genesis collect-gentxs --home "$TEST_DIR/node1" 2>/dev/null

# 启动链
echo ""
echo "Step 2: Start chain"
msgd start \
    --home "$TEST_DIR/node1" \
    --rpc.laddr tcp://0.0.0.0:26657 \
    --grpc.address 0.0.0.0:9090 \
    > "$TEST_DIR/node1.log" 2>&1 &

# 等待链启动
echo "Waiting for chain to start..."
for i in $(seq 1 20); do
    if msgd status --home "$TEST_DIR/node1" 2>/dev/null; then
        echo "Chain started"
        break
    fi
    sleep 2
done

# IBC 功能测试
echo ""
echo "Step 3: Test IBC module availability"

# 检查 IBC 相关查询
echo "  Testing IBC queries..."
QUERY_TESTS=(
    "ibc client params"
    "ibc connection connections"
    "ibc channel channels"
    "ibc transfer params"
)

for query in "${QUERY_TESTS[@]}"; do
    if msgd query $query --home "$TEST_DIR/node1" --output json > /dev/null 2>&1; then
        echo "  ✓ $query"
    else
        echo "  ✗ $query"
    fi
done

# 升级测试
echo ""
echo "Step 4: Perform upgrade test"

CURRENT_HEIGHT=$(msgd status --home "$TEST_DIR/node1" | \
    jq -r '.SyncInfo.latest_block_height')
UPGRADE_HEIGHT=$((CURRENT_HEIGHT + 10))
echo "Current height: $CURRENT_HEIGHT"
echo "Upgrade height: $UPGRADE_HEIGHT"

# 提交升级提案
msgd tx gov submit-legacy-proposal software-upgrade "v2.0.0" \
    --title "IBC Upgrade Test" \
    --description "Testing IBC upgrade compatibility" \
    --upgrade-height "$UPGRADE_HEIGHT" \
    --deposit 100000000umsg \
    --from validator \
    --keyring-backend test \
    --home "$TEST_DIR/node1" \
    --chain-id "$CHAIN_ID" \
    --gas auto \
    --gas-prices "1000000000umsg" \
    --yes > /dev/null 2>&1

sleep 2

# 投票
msgd tx gov vote 1 yes \
    --from validator \
    --keyring-backend test \
    --home "$TEST_DIR/node1" \
    --chain-id "$CHAIN_ID" \
    --gas auto \
    --gas-prices "1000000000umsg" \
    --yes > /dev/null 2>&1

# 等待升级
echo "Waiting for upgrade at height $UPGRADE_HEIGHT..."
sleep $(( UPGRADE_HEIGHT * 5 ))

# 验证
echo ""
echo "Step 5: Post-upgrade verification"

# 检查 IBC 模块
CLIENTS=$(msgd query ibc client states \
    --home "$TEST_DIR/node1" --output json 2>/dev/null | \
    jq '.client_states | length' || echo "N/A")
echo "  IBC clients: $CLIENTS"

CONNECTIONS=$(msgd query ibc connection connections \
    --home "$TEST_DIR/node1" --output json 2>/dev/null | \
    jq '.connections | length' || echo "N/A")
echo "  IBC connections: $CONNECTIONS"

CHANNELS=$(msgd query ibc channel channels \
    --home "$TEST_DIR/node1" --output json 2>/dev/null | \
    jq '.channels | length' || echo "N/A")
echo "  IBC channels: $CHANNELS"

echo ""
echo "=== Test Complete ==="

10.2 模拟升级

10.2.1 使用 Interchaintest

Interchaintest(原 ibctest)是 IBC 升级测试的核心框架:

// interchaintest 升级测试
func TestIBCUpgrade(t *testing.T) {
    if testing.Short() {
        t.Skip("skipping IBC upgrade test in short mode")
    }

    // 配置测试链
    cfg := ibctest.ChainConfig{
        Name:       "msg-chain",
        ChainID:    "msg-chain-1",
        Binary:     "msgd",
        Bech32Prefix: "msg",
        Denom:      "umsg",
        GasPrices:  "1000000000umsg",
    }

    // 创建测试网络
    network := ibctest.NewSuite(t, cfg)

    // 启动旧版本链
    network.StartOldVersion(t, "v7.3.0")

    // 建立 IBC 连接
    network.CreateConnections(t, 2)

    // 测试 IBC 转账
    network.TestICS20Transfer(t)

    // 执行升级
    network.Upgrade(t, "v2.0.0", 20)

    // 升级后验证
    network.VerifyClients(t)
    network.VerifyConnections(t)
    network.VerifyChannels(t)

    // 再次测试 IBC 转账
    network.TestICS20Transfer(t)
}

10.2.2 混沌工程测试

#!/usr/bin/env python3
"""
IBC 升级混沌测试
注入故障以验证升级的鲁棒性
"""
import random
import time
import subprocess
import signal
import sys
from typing import List, Callable

class IBCChaosTest:
    def __init__(self, chain_home: str):
        self.chain_home = chain_home
        self.faults = []
        self.running = True

    def register_fault(self, name: str,
                      trigger: Callable,
                      probability: float = 0.1):
        """注册故障注入"""
        self.faults.append({
            "name": name,
            "trigger": trigger,
            "probability": probability
        })

    def inject_faults(self):
        """随机注入故障"""
        for fault in self.faults:
            if random.random() < fault.probability:
                print(f"  Injecting fault: {fault['name']}")
                result = fault["trigger"]()
                if result:
                    print(f"    Fault active: {result}")

    def run_upgrade_test(self, upgrade_height: int):
        """运行带故障注入的升级测试"""
        print("Starting chaos upgrade test...")
        print(f"Target upgrade height: {upgrade_height}")

        def signal_handler(sig, frame):
            print("\nTest interrupted")
            self.running = False
            sys.exit(0)

        signal.signal(signal.SIGINT, signal_handler)

        while self.running:
            try:
                result = subprocess.run([
                    "msgd", "status",
                    "--home", self.chain_home
                ], capture_output=True, text=True, timeout=10)

                status = eval(result.stdout)
                current_height = int(status["SyncInfo"]["latest_block_height"])

                if current_height >= upgrade_height:
                    print("Upgrade height reached!")
                    break

                if current_height > upgrade_height - 10:
                    self.inject_faults()

                time.sleep(2)

            except Exception as e:
                print(f"Error: {e}")
                time.sleep(5)

        print("Chaos test completed")

# 故障定义
def network_partition(chain_home: str) -> str:
    """模拟网络分区"""
    return "Network partition simulated"

def delayed_block(chain_home: str) -> str:
    """模拟区块延迟"""
    time.sleep(15)
    return "Block delayed 15s"

def corrupted_mempool(chain_home: str) -> str:
    """模拟交易池问题"""
    return "Mempool stress test"

if __name__ == "__main__":
    test = IBCChaosTest(chain_home="/tmp/msgd-test")

    test.register_fault("network-partition",
        lambda: network_partition("/tmp/msgd-test"), 0.2)
    test.register_fault("delayed-block",
        lambda: delayed_block("/tmp/msgd-test"), 0.1)
    test.register_fault("corrupted-mempool",
        lambda: corrupted_mempool("/tmp/msgd-test"), 0.3)

    test.run_upgrade_test(upgrade_height=100)

10.3 集成测试框架

10.3.1 IBC 升级集成测试套件

package upgrade_test

import (
    "testing"
    "time"

    "github.com/stretchr/testify/suite"
    "github.com/cosmos/cosmos-sdk/testutil/network"
    ibctesting "github.com/cosmos/ibc-go/v8/testing"
)

type IBCUpgradeTestSuite struct {
    suite.Suite
    coordinator *ibctesting.Coordinator

    // 测试链
    chainA *ibctesting.TestChain  // MSG Chain(升级链)
    chainB *ibctesting.TestChain  // 对端链

    // 连接信息
    path *ibctesting.Path
}

func (s *IBCUpgradeTestSuite) SetupSuite() {
    s.coordinator = ibctesting.NewCoordinator(s.T(), 2)

    s.chainA = s.coordinator.GetChain(ibctesting.GetChainID(1))
    s.chainB = s.coordinator.GetChain(ibctesting.GetChainID(2))

    s.chainA.CurrentHeader.ChainID = "msg-chain-1"

    s.path = ibctesting.NewPath(s.chainA, s.chainB)
    s.coordinator.Setup(s.path)
}

// 测试 1: 基本升级 IBC 连续性
func (s *IBCUpgradeTestSuite) TestBasicUpgradeContinuity() {
    s.coordinator.CreateConnections(s.path)
    s.coordinator.CreateChannels(s.path)

    channel := s.chainA.GetChannel(s.path.EndpointA.ChannelConfig.PortID,
        s.path.EndpointA.ChannelID)
    s.Require().Equal(ibctesting.OPEN, channel.State)

    upgradeHeight := s.chainA.CurrentHeader.Height + 10
    s.chainA.App.(*App).UpgradeKeeper.ScheduleUpgrade(s.chainA.GetContext(),
        upgradetypes.Plan{
            Name:   "v2.0.0",
            Height: upgradeHeight,
        })

    s.coordinator.CommitNBlocks(s.chainA, 15)

    upgraded := s.chainA.App.(*App).UpgradeKeeper.IsUpgradeHeight(s.chainA.GetContext())
    s.Require().True(upgraded)

    channel = s.chainA.GetChannel(s.path.EndpointA.ChannelConfig.PortID,
        s.path.EndpointA.ChannelID)
    s.Require().Equal(ibctesting.OPEN, channel.State)

    transferCoin := sdk.NewCoin("umsg", sdk.NewInt(1000))
    msg := ibctesting.NewMsgTransfer(s.path.EndpointA,
        s.path.EndpointB, transferCoin, s.chainA.SenderAccount.GetAddress())

    result, err := s.chainA.SendMsgs(s.path.EndpointA, msg)
    s.Require().NoError(err)
    s.Require().NotNil(result)
}

// 测试 2: 升级客户端过期与恢复
func (s *IBCUpgradeTestSuite) TestClientExpiryAndRecovery() {
    s.coordinator.CreateConnections(s.path)
    s.coordinator.CreateChannels(s.path)

    clientID := s.path.EndpointA.ClientID

    s.chainA.ExpireClient(clientID)

    clientState := s.chainA.GetClientState(clientID)
    s.Require().True(clientState.IsExpired())

    substitutePath := ibctesting.NewPath(s.chainA, s.chainB)
    s.coordinator.CreateClients(substitutePath)

    proposal := ibctesting.NewClientUpdateProposal(
        s.chainA, substitutePath.EndpointA, clientID)

    err := s.chainA.App.GovernanceKeeper.SubmitProposal(
        s.chainA.GetContext(), proposal)
    s.Require().NoError(err)

    s.coordinator.CommitNBlocks(s.chainA, 2)

    clientState = s.chainA.GetClientState(clientID)
    s.Require().False(clientState.IsExpired())

    transferCoin := sdk.NewCoin("umsg", sdk.NewInt(500))
    msg := ibctesting.NewMsgTransfer(s.path.EndpointA,
        s.path.EndpointB, transferCoin, s.chainA.SenderAccount.GetAddress())

    _, err = s.chainA.SendMsgs(s.path.EndpointA, msg)
    s.Require().NoError(err)
}

// 测试 3: 通道关闭与治理恢复
func (s *IBCUpgradeTestSuite) TestChannelCloseAndRecovery() {
    s.path.EndpointA.ChannelConfig.Ordering = ibctesting.ORDERED
    s.path.EndpointB.ChannelConfig.Ordering = ibctesting.ORDERED
    s.coordinator.Setup(s.path)

    packet := ibctesting.NewPacket(s.chainA, s.chainB, 1,
        s.path.EndpointA.ChannelConfig.PortID,
        s.path.EndpointA.ChannelID,
        s.path.EndpointB.ChannelConfig.PortID,
        s.path.EndpointB.ChannelID,
        []byte("test data"), sdk.ZeroInt(), ibctesting.DefaultTimeout)

    err := s.coordinator.SendPacket(s.path.EndpointA, packet)
    s.Require().NoError(err)

    // 模拟超时
    s.coordinator.IncrementTime(s.chainA, 24*time.Hour)

    channel := s.chainA.GetChannel(s.path.EndpointA.ChannelConfig.PortID,
        s.path.EndpointA.ChannelID)
    s.Require().Equal(ibctesting.CLOSED, channel.State)

    // 通过治理恢复
    recoveryProposal := NewReopenChannelProposal(
        s.path.EndpointA.ChannelConfig.PortID,
        s.path.EndpointA.ChannelID)

    err = s.chainA.App.GovernanceKeeper.SubmitProposal(
        s.chainA.GetContext(), recoveryProposal)
    s.Require().NoError(err)

    s.coordinator.CommitNBlocks(s.chainA, 2)

    channel = s.chainA.GetChannel(s.path.EndpointA.ChannelConfig.PortID,
        s.path.EndpointA.ChannelID)
    s.Require().Equal(ibctesting.OPEN, channel.State)
}

11. 案例:Cosmos 生态 IBC 升级事件回顾

11.1 Cosmos Hub v9 → v10 升级

11.1.1 事件概述

时间: 2023 年 3 月
涉及链: Cosmos Hub(cosmoshub-4)
升级内容: ibc-go v5 → v6 + SDK v0.45 → v0.46

11.1.2 关键问题

  1. IBC 客户端冻结:升级后大部分 IBC 轻客户端因信任期过期而冻结
  2. 通道关闭:多个 ORDERED 通道因数据包超时而关闭
  3. 恢复延迟:由于验证人未及时更新客户端,恢复过程持续了 48 小时以上

11.1.3 经验教训

教训 1: 升级前应通知中继器提前更新客户端
教训 2: 升级后应优先恢复 IBC 客户端
教训 3: ORDERED 通道在重大升级中风险较高
教训 4: 需要建立升级后的 IBC 健康检查流程

11.1.4 时间线

Day 1 14:00 UTC - 升级开始
Day 1 14:05 UTC - 升级完成,区块恢复生产
Day 1 14:30 UTC - 发现 IBC 客户端大量过期
Day 1 16:00 UTC - 开始手动恢复客户端
Day 1 22:00 UTC - 50% 客户端恢复
Day 2 10:00 UTC - 90% 客户端恢复
Day 2 14:00 UTC - 所有客户端恢复完成
Day 2 16:00 UTC - IBC 转账完全恢复

11.2 Osmosis v11 → v12 升级

11.2.1 事件概述

时间: 2023 年 6 月
涉及链: Osmosis(osmosis-1)
升级内容: SDK v0.45 → v0.47 + ibc-go v5 → v7

11.2.2 关键问题

  1. 状态迁移失败:IBC 存储迁移导致部分通道状态丢失
  2. 治理恢复:需通过 3 个治理提案恢复 IBC 连接
  3. 代币轨迹损坏:部分 IBC 代币追踪记录在升级中损坏

11.2.3 经验教训

教训 1: 升级前必须进行完整的创世状态导出和验证
教训 2: 存储迁移需要充分测试不同场景
教训 3: 需要准备治理恢复的标准化流程
教训 4: 代币轨迹完整性检查应纳入升级后验证

11.3 Juno 网络分区事件

11.3.1 事件概述

时间: 2023 年 10 月
涉及链: Juno(juno-1)
问题: 验证人间网络分区导致链分叉

11.3.2 IBC 影响

  1. 轻客户端误行为检测:对端链检测到 Juno 的双重签名
  2. 客户端冻结:所有与 Juno 的 IBC 连接立即冻结
  3. 通道关闭:部分 ORDERED 通道自动关闭

11.3.3 恢复过程

阶段 1: 链恢复(6 小时)
  网络分区解决 → 链恢复共识

阶段 2: 客户端状态验证(12 小时)
  各对端链验证 Juno 的最终状态
  提交误行为证据或更新客户端

阶段 3: 通道恢复(24 小时+)
  治理提案恢复关闭的通道
  重新建立 IBC 连接

11.4 Neutron 跨链账户升级

11.4.1 事件概述

时间: 2024 年 2 月
涉及链: Neutron(neutron-1)
升级内容: ICS-27 Interchain Accounts 升级

11.4.2 关键变更

  1. ICA 通道重开:所有现有 ICA 通道需重新建立
  2. 控制器注册:ICA 控制器地址需重新注册
  3. 治理协调:需与所有对端链协调升级

11.4.3 最佳实践

// Neutron 的 ICA 升级策略
// 1. 预留 ICA 升级专用治理提案
// 2. 对端链协调窗口:72 小时
// 3. 逐步恢复:先恢复低价值通道,再恢复生产通道

type ICAUpgradePlan struct {
    ChannelID      string
    PortID         string
    Priority       int  // 1=最高,3=最低
    RecoveryAction string // recreate | migrate | reopen
    Coordinated    bool
}

11.5 关键经验总结

11.5.1 共性模式

分析上述案例,IBC 升级问题存在以下共性模式:

模式 出现频率 影响程度 预防措施
客户端过期 极高 中 自动更新脚本
ORDERED 通道关闭 高 高 使用 UNORDERED
存储迁移失败 中 极高 充分测试
代币轨迹损坏 中 高 完整性检查
治理恢复延迟 高 中 预准备提案

11.5.2 MSG Chain 的改进措施

基于上述案例,MSG Chain 采用以下改进措施:

  1. 自动客户端更新:部署高可用的客户端更新服务
  2. UNORDERED 优先:新通道默认使用 UNORDERED 类型
  3. 预迁移测试:每次升级前在测试网进行完整的 IBC 迁移测试
  4. 治理提案模板:预置通道恢复、客户端恢复等标准化治理提案
  5. 升级后验证清单:系统化的 IBC 健康检查流程

12. 总结

12.1 核心要点回顾

IBC 协议升级是 Cosmos 生态运营中不可避免的挑战。本文档系统性地介绍了在 msg-chain-1 上进行 IBC 升级的全流程管理:

升级前 - 准备阶段:

  1. 全面评估升级影响范围(客户端、连接、通道)
  2. 制定升级窗口并协调对端链
  3. 准备恢复脚本和治理提案模板
  4. 在测试网完成全流程模拟升级
  5. 部署 Cosmovisor 并配置自动回退

升级中 - 执行阶段:

  1. 密切监控升级高度和区块生产
  2. 验证 AppHash 和共识状态
  3. 快速执行状态迁移(ibc-go v8+)
  4. 及时恢复 IBC 客户端

升级后 - 恢复阶段:

  1. 批量恢复过期轻客户端
  2. 通过治理恢复关闭的通道
  3. 验证 IBC 转账和跨链功能
  4. 监控 IBC 健康指标
  5. 召开升级回顾会议

12.2 关键数据汇总

指标 推荐值 说明
升级通知期 ≥ 48 小时 验证人和社区通知
对端协调期 ≥ 72 小时 跨链协调窗口
IBC 恢复预期 2-4 小时 标准情况
回滚决策期 30 分钟 决定是否回滚
客户端更新频率 每 4 小时 预防客户端过期
通道恢复 1-2 天 需治理提案投票

12.3 自动化路线图

短期目标(Q3 2026):

中期目标(Q4 2026):

长期目标(2027+):

12.4 参考资源

官方文档:

MSG Chain 资源:

社区资源:

12.5 附录

附录 A: 升级检查清单

# MSG Chain IBC 升级检查清单

## 升级前(T - 7 天)
- [ ] 确定升级范围和影响分析
- [ ] 在测试网完成升级模拟
- [ ] 准备恢复操作手册

## 升级前(T - 48 小时)
- [ ] 发布升级公告
- [ ] 通知对端链运营团队
- [ ] 完成二进制签名和分发
- [ ] 验证人意向收集

## 升级前(T - 1 小时)
- [ ] 暂停关键 IBC 活动
- [ ] 导出升级前状态快照
- [ ] 确认所有验证人就绪

## 升级中
- [ ] 监控升级高度到达
- [ ] 验证 AppHash 一致性
- [ ] 检查区块生产

## 升级后(T + 1 小时)
- [ ] 恢复 IBC 客户端
- [ ] 验证客户端状态
- [ ] 恢复通道
- [ ] 测试 IBC 转账

## 升级后(T + 24 小时)
- [ ] 全面 IBC 健康检查
- [ ] 监控异常行为
- [ ] 发布升级完成报告

附录 B: 常见问题排查

问题 诊断命令 解决方案
客户端过期 msgd query ibc client status [id] 更新或替换客户端
通道关闭 msgd query ibc channel end [port] [chan] 治理提案恢复
交易超时 msgd query ibc channel packets [port] [chan] 重新发送或超时退回
代币轨迹丢失 msgd query ibc transfer denom-trace [hash] 手动注册代币轨迹
连接断开 msgd query ibc connection end [id] 重新建立连接

附录 C: 速查命令

# 客户端管理
msgd query ibc client states                                    # 列出所有客户端
msgd query ibc client state [client-id]                          # 客户端详情
msgd query ibc client status [client-id]                         # 客户端状态
msgd tx ibc client update [client-id] --from [key]              # 更新客户端
msgd tx ibc client submit-misbehaviour [client-id] [evidence]  # 提交误行为

# 连接管理
msgd query ibc connection connections                            # 列出所有连接
msgd query ibc connection end [connection-id]                    # 连接详情

# 通道管理
msgd query ibc channel channels                                  # 列出所有通道
msgd query ibc channel end [port-id] [channel-id]               # 通道详情
msgd query ibc channel packets [port-id] [channel-id]           # 待处理数据包

# 转账
msgd tx ibc transfer transfer [src-port] [src-channel] [receiver] [amount]  # IBC 转账
msgd query ibc transfer denom-trace [hash]                      # 查询代币轨迹
msgd query ibc transfer escrow-address                          # 查询托管地址

# 升级
msgd query upgrade applied [upgrade-name]                       # 查询升级状态
msgd query upgrade plan                                          # 查询计划升级

文档维护者: MSG Chain 技术团队
反馈渠道: docs@msgchain.org | GitHub Issues
许可协议: CC-BY-4.0