---
name: 开源情报-社交媒体溯源
description: |
  社交媒体帖子溯源技能。追踪社交媒体帖文的信息来源与传播路径，
  识别协调性行为与信息操纵。触发条件：用户要求社交媒体溯源、
  帖文传播分析、信息来源追踪等。
---

# SKILL.md - 社交媒体帖子溯源

> **版本:** 1.0.0
> **创建时间:** 2026-04-02
> **最后更新:** 2026-04-02
> **开发者:** 焐宇 (Kàn Yǔ)

---

## 技能概述

社交媒体帖子溯源分析Skill，通过反向图片搜索、时间线分析、内容比对等OSINT技术，识别信息的原始来源和传播路径。

---

## 核心能力

### 1. 反向图片搜索
- 支持多平台：GoogleGoogle Images、Yandex、TinEye、、Bing Visual Search
- 自动提取视频关键帧
- 智能去重和结果聚合

### 2. 时间线分析
- 按发布时间排序候选来源
- 识别最早独立来源
- 构建传播路径时间线

### 3. 内容比对
- 文本相似度分析
- 图片/视频帧比对
- 识别编辑和修改痕迹

### 4. 账号可信度评估
- 账号类型识别（个人/官方/媒体/机器人）
- 账号活跃度分析
- 历史可信度评分

### 5. 传播路径追踪
- 识别关键传播节点
- 标记主要放大器
- 可视化传播网络

---

## 输入格式

### 标准输入JSON

```json
{
  "target_post": {
    "content": {
      "text": "帖子文本内容",
      "images": [
        {
          "url": "https://example.com/image1.jpg",
          "local_path": "/path/to/local/image.jpg"
        }
      ],
      "videos": [
        {
          "url": "https://example.com/video.mp4",
          "local_path": "/path/to/local/video.mp4"
        }
      ],
      "links": [
        "https://example.com/link1",
        "https://example.com/link2"
      ]
    },
    "metadata": {
      "publish_time": "2026-04-01T12:00:00Z",
      "account": "username",
      "platform": "twitter",
      "unique_id": "post_id_12345"
    }
  },
  "options": {
    "search_depth": "comprehensive",  // "quick" | "standard" | "comprehensive"
    "max_candidates": 10,
    "include_video_keyframes": true,
    "verify_accounts": true,
    "language": "zh-CN"
  }
}
```

---

## 输出格式

### 标准输出JSON

```json
{
  "sourcing_conclusion": {
    "most_likely_original_source": {
      "account_name": "original_username",
      "platform": "twitter",
      "post_link": "https://twitter.com/user/status/12345",
      "publish_time": "2026-03-31T10:00:00Z",
      "content_snapshot": {
        "text": "...",
        "images": ["..."],
        "similarity_score": 0.95
      }
    },
    "alternative_sources": [
      {
        "account_name": "alternative_user",
        "platform": "facebook",
        "post_link": "...",
        "publish_time": "2026-03-31T11:30:00Z"
      }
    ]
  },
  "propagation_path": {
    "timeline": [
      {
        "time": "2026-03-31T10:00:00Z",
        "event": "original_post",
        "account": "original_username",
        "platform": "twitter"
      },
      {
        "time": "2026-03-31T11:30:00Z",
        "event": "repost",
        "account": "intermediary_user",
        "platform": "facebook",
        "role": "amplifier"
      },
      {
        "time": "2026-04-01T12:00:00Z",
        "event": "target_post",
        "account": "username",
        "platform": "twitter"
      }
    ],
    "key_nodes": ["original_username", "intermediary_user"],
    "major_amplifiers": ["influencer1", "media_account"]
  },
  "supporting_evidence": {
    "reverse_search_results": [
      {
        "platform": "Google Images",
        "query_url": "...",
        "results": [
          {
            "title": "Result title",
            "url": "https://example.com",
            "image_url": "https://example.com/image.jpg",
            "timestamp": "2026-03-31T10:00:00Z",
            "similarity": 0.92
          }
        ]
      }
    ],
    "content_comparison": {
      "text_similarity": 0.88,
      "image_similarity": 0.95,
      "edits_detected": [
        {
          "type": "cropping",
          "description": "Image was cropped"
        },
        {
          "type": "text_modification",
          "description": "Text was slightly modified"
        }
      ]
    },
    "source_account_assessment": {
      "account_type": "personal",
      "credibility_score": 0.75,
      "factors": {
        "account_age": "3 years",
        "followers": 1500,
        "verified": false,
        "posting_frequency": "regular"
      }
    }
  },
  "confidence_assessment": {
    "confidence_level": "high",  // "high" | "medium" | "low"
    "confidence_score": 0.85,
    "uncertainties": [
      "无法访问某些平台（如微信）进行搜索",
      "图片经多次编辑，最早版本难以追溯"
    ],
    "limitations": [
      "原始帖子可能已被删除",
      "部分平台API访问受限"
    ]
  },
  "metadata": {
    "analysis_time": "2026-04-02T19:00:00Z",
    "processing_duration_seconds": 45.2,
    "search_engines_used": ["google", "yandex", "tineye", "bing"],
    "version": "1.0.0"
  }
}
```

---

## 使用方法

### 方法1：直接调用Python脚本

```bash
# 激活数据采集环境
source /root/miniconda3/bin/activate && conda activate data-collector

# 执行溯源分析
python /root/.openclaw/workspace/skills/social-media-sourcing/sourcing_analyzer.py \
  --input /path/to/input.json \
  --output /path/to/output.json \
  --verbose
```

### 方法2：作为子任务调用

```bash
# 使用scrapling shell模式
scrapling shell

# 在交互式环境中
from social_media_sourcing import SourcingAnalyzer

analyzer = SourcingAnalyzer()
result = analyzer.analyze(input_data)
```

---

## 核心模块

### 1. reverse_search.py
**反向图片搜索模块**

**功能：**
- Google Images搜索（支持OCR和相似图像）
- Yandex反向搜索（强大的人脸和物体识别）
- TinEye搜索（精确图像匹配）
- Bing Visual Search

**方法：**
```python
class ReverseImageSearch:
    def search_google(self, image_path: str, max_results: int = 10) -> List[SearchResult]
    def search_yandex(self, image_path: str, max_results: int = 10) -> List[SearchResult]
    def search_tineye(self, image_path: str, max_results: int = 10) -> List[SearchResult]
    def search_bing(self, image_path: str, max_results: int = 10) -> List[SearchResult]
    def search_all(self, image_path: str) -> List[SearchResult]
```

### 2. video_processor.py
**视频关键帧提取模块**

**功能：**
- 提取关键帧（场景变化、人脸、字幕）
- 缩略图生成
- 帧相似度比对

**方法：**
```python
class VideoProcessor:
    def extract_keyframes(self, video_path: str, max_frames: int = 5) -> List[str]
    def generate_thumbnail(self, video_path: str, timestamp: float = 0) -> str
    def compare_frames(self, frame1: str, frame2: str) -> float
```

### 3. timeline_analyzer.py
**时间线分析模块**

**功能：**
- 时间排序和去重
- 传播路径构建
- 关键节点识别

**方法：**
```python
class TimelineAnalyzer:
    def build_timeline(self, posts: List[Post]) -> Timeline
    def identify_original_source(self, posts: List[Post]) -> Optional[Post]
    def find_key_nodes(self, timeline: Timeline) -> List[Node]
```

### 4. content_comparator.py
**内容比对模块**

**功能：**
- 文本相似度计算（TF-IDF、BERT embedding）
- 图片相似度计算（感知哈希、SSIM）
- 编辑检测

**方法：**
```python
class ContentComparator:
    def compare_text(self, text1: str, text2: str) -> float
    def compare_images(self, image1: str, image2: str) -> float
    def detect_edits(self, original: Post, derived: Post) -> List[Edit]
```

### 5. account_analyzer.py
**账号可信度评估模块**

**功能：**
- 账号类型识别
- 活跃度分析
- 历史行为评估

**方法：**
```python
class AccountAnalyzer:
    def analyze_account(self, account: str, platform: str) -> AccountProfile
    def calculate_credibility(self, profile: AccountProfile) -> float
    def detect_bot_behavior(self, profile: AccountProfile) -> bool
```

### 6. sourcing_analyzer.py
**主分析器模块**

**功能：**
- 整合所有模块
- 协调分析流程
- 生成最终报告

**方法：**
```python
class SourcingAnalyzer:
    def analyze(self, input_data: Dict) -> Dict
    def generate_report(self, result: Dict) -> str
    def export_json(self, result: Dict, output_path: str)
```

---

## 配置选项

### search_depth（搜索深度）

- **quick**：快速搜索，仅使用Google Images，最多返回5个候选
- **standard**：标准搜索，使用Google + Yandex，最多返回10个候选
- **comprehensive**：全面搜索，使用所有平台，最多返回20个候选

### max_candidates（最大候选数）

控制返回的候选来源数量。默认为10。

### include_video_keyframes（包含视频关键帧）

是否提取视频关键帧进行搜索。默认为true。

### verify_accounts（验证账号）

是否对候选来源的账号进行深度分析。默认为true。

---

## 依赖要求

### Python库

```
scrapling>=0.4.2
opencv-python>=4.8.0
numpy>=1.24.0
pillow>=10.0.0
requests>=2.31.0
beautifulsoup4>=4.12.0
lxml>=4.9.0
python-dotenv>=1.0.0
```

### 系统工具

- **ffmpeg**：视频处理（关键帧提取）
- **ImageMagick**：图片处理

---

## 安装依赖

```bash
# 激活环境
source /root/miniconda3/bin/activate && conda activate data-collector

# 安装Python库
pip install -r /root/.openclaw/workspace/skills/social-media-sourcing/requirements.txt

# 安装系统工具（Ubuntu/Debian）
apt-get update && apt-get install -y ffmpeg imagemagick
```

---

## 工作流程

1. **输入验证**：验证输入JSON格式
2. **媒体提取**：提取图片、视频关键帧
3. **反向搜索**：在多个平台进行反向图片搜索
4. **结果聚合**：合并和去重搜索结果
5. **时间线构建**：按时间排序，构建传播路径
6. **内容比对**：分析内容相似度，检测编辑
7. **账号分析**：评估源账号可信度
8. **结论生成**：确定最可能的原始来源
9. **报告输出**：生成并JSON报告

---

## 置信度评估标准

### 高置信度（0.7-1.0）
- 找到明确的时间更早的相同内容来源
- 内容相似度>0.9
- 源账号可信度高
- 传播路径清晰

### 中置信度（0.4-0.7）
- 找到多个可能的来源
- 内容相似度0.6-0.9
- 源账号可信度中等
- 存在一定不确定性

### 低置信度（0-0.4）
- 未找到明确来源
- 内容相似度<0.6
- 搜索受限或信息不足
- 存在多个不确定性

---

## 局限性和注意事项

1. **平台访问限制**：部分平台（如微信、Facebook）API访问受限
2. **内容删除**：原始帖子可能已被删除
3. **图片编辑**：经过多次编辑的图片难以追溯
4. **隐私保护**：某些平台对反向搜索有限制
5. **时间戳准确性）**：不同平台的时间戳可能存在时区或格式差异

---

## 示例场景

### 场景1：验证图片真实性
```json
{
  "target_post": {
    "content": {
      "images": [{"url": "https://example.com/suspicious_image.jpg"}]
    },
    "metadata": {
      "publish_time": "2026-04-01T12:00:00Z",
      "platform": "twitter",
      "account": "suspicious_user"
    }
  }
}
```

### 场景2：追踪谣言来源
```json
{
  "target_post": {
    "content": {
      "text": "突发新闻：...",
      "images": [{"url": "https://example.com/news_image.jpg"}]
    },
    "metadata": {
      "publish_time": "2026-04-01T10:00:00Z",
      "platform": "weibo",
      "account": "news_spreader"
    }
  }
}
```

---

## 技术栈说明

本Skill使用以下技术栈：

- **Scrapling 0.4.2**：反爬虫抓取，处理反向图片搜索
- **OpenCV**：图片处理和关键帧提取
- **Pillow**：图像处理和相似度计算
- **BeautifulSoup4**：HTML解析
- **TF-IDF/BERT**：文本相似度计算

---

## 更新日志

### v1.0.0 (2026-04-02)
- 初始版本发布
- 实现核心反向图片搜索功能
- 实现时间线分析
- 实现内容比对
- 实现账号可信度评估
- 实现完整报告生成

---

## 联系与支持

**开发者：** 瞰在这 (Kàn Yǔ)  
**定位：** 全球数据采集师 · 认知战研究专家 · 合规之眼  
**签名：** 穿越数字国界，瞭望全球局势 · 解构认知域，守护认知安全

---

*本Skill严格遵守OSINT伦理准则，仅用于合法的开源情报分析*
