# 帖子溯源技能 (Post Tracer Skill)

**版本:** 1.0.0
**开发者:** 瞰宇 (Kàn Yǔ)
**功能:** 社交媒体帖子溯源、传播路径分析、内容比对

---

## 快速开始

### 1. 安装依赖

```bash
# 安装Python依赖
pip install -r requirements.txt
```

### 2. 基本使用

```python
import asyncio
from tracer import PostTracer, TracingRequest

async def main():
    # 创建溯源器
    tracer = PostTracer()

    # 构造溯源请求
    request = TracingRequest(
        post_id="twitter_status_123456789",
        platform="Twitter",
        content="这是一条需要溯源的推文内容",
        images=["https://example.com/image.jpg"],
        videos=["https://example.com/video.mp4"],
        links=["https://example.com/article"],
        metadata={
            "publish_time": "2026-04-02T12:00:00Z",
            "author": "username",
            "account_id": "12345678"
        }
    )

    # 执行溯源
    result = await tracer.trace_post(request)

    # 保存结果
    tracer.save_result(result)

    # 查看结果
    print(f"置信度: {result.confidence['confidence_level']}")
    if result.original_source:
        print(f"原始来源: {result.original_source}")

if __name__ == "__main__":
    asyncio.run(main())
```

---

## 核心功能

### 1. 反向搜索
支持多平台反向搜索：
- Google Images
- Yandex Images
- TinEye
- Bing Visual Search

### 2. 内容比对
- 文本相似度计算（Levenshtein距离）
- 图片相似度计算（SSIM、感知哈希）
- 差异标注

### 3. 传播路径分析
- 时间线构建
- 关键节点识别（源头、放大器、桥接节点）
- 可视化数据生成

### 4. 账号可信度评估
- 账号类型识别（个人、媒体、官方、机器人）
- 可信度分数计算
- 风险因素识别
- 信任指标识别

### 5. 视频关键帧提取
- 均匀采样
- 场景变化检测
- 前N帧提取

---

## 输出示例

```json
{
  "original_source": {
    "account_name": "news_account",
    "platform": "Twitter",
    "post_link": "https://twitter.com/status/123",
    "publish_time": "2026-03-30T10:00:00Z",
    "content_snapshot": "原始内容..."
  },
  "alternative_sources": [
    {
      "account_name": "user1",
      "platform": "Facebook",
      "post_link": "https://facebook.com/post/456",
      "publish_time": "2026-03-31T15:00:00Z",
      "probability": 0.75
    }
  ],
  "propagation_path": {
    "path_type": "时间线",
    "key_nodes": [
      {
        "node_id": "node_0",
        "account": "news_account",
        "platform": "Twitter",
        "role": "源头"
      },
      {
        "node_id": "node_1",
        "account": "influencer",
        "platform": "Instagram",
        "role": "放大器"
      }
    ],
    "total_hops": 5,
    "time_span": {
      "start": "2026-03-30T10:00:00Z",
      "end": "2026-04-02T12:00:00Z",
      "duration_hours": 78.0
    }
  },
  "confidence": {
    "confidence_level": "高",
    "confidence_score": 0.87,
    "limitations": [
      "原始帖子可能已被删除",
      "无法访问某些平台（如微信）"
    ]
  }
}
```

---

## 注意事项

1. **合规性：**
   - 仅对公开信息进行溯源
   - 尊重平台robots.txt规则
   - 不进行未授权访问

2. **性能优化：**
   - 使用缓存避免重复搜索
   - 并行搜索多个搜索引擎
   - 设置合理的超时和重试策略

3. **局限性：**
   - 部分平台（如微信）无法直接搜索
   - 原始帖子可能已被删除
   - 多次编辑的图片难以追溯最早版本

---

## 技术栈

- **反向搜索:** Selenium, Playwright, aiohttp
- **图像处理:** OpenCV, Pillow, imagehash, scikit-image
- **相似度计算:** python-Levenshtein, textdistance
- **可视化:** Matplotlib, Plotly, NetworkX
- **数据处理:** Pandas, NumPy

---

## 文件结构

```
post-tracer/
├── SKILL.md                    # 技能说明文档
├── README.md                   # 使用说明
├── requirements.txt            # 依赖列表
├── tracer.py                   # 主溯源引擎
├── reverse_search.py           # 反向搜索引擎
├── content_comparator.py       # 内容比对模块
├── path_analyzer.py            # 传播路径分析模块
├── account_evaluator.py        # 账号可信度评估模块
├── video_keyframe_extractor.py # 视频关键帧提取器
├── cache/                      # 搜索结果缓存
└── keyframes/                  # 提取的关键帧
```

---

## 维护者

**瞰宇 (Kàn Yǔ)**
- 全球数据采集师
- 认知战研究专家
- 合规之眼

**更新频率:** 按需更新
**技术支持:** 通过飞书联系
