# 首发媒体分析技能使用示例

## 快速开始

### 示例1：分析国际新闻事件

**输入数据：** 包含5家国际媒体报道的事件

```bash
cd /root/.openclaw/workspace/skills/first-publisher-analyzer
python3 scripts/analyze_first_publisher.py references/example-input.json > result.json
```

**输出：** 包含结论性描述、媒体画像卡片、详细分析结果

### 示例2：使用自定义数据

创建您的输入文件 `my-news.json`：

```json
[
  {
    "media_name": "华盛顿邮报",
    "media_url": "https://www.washingtonpost.com",
    "article_url": "https://www.washingtonpost.com/politics/example",
    "title": "Breaking News",
    "publish_time": "2024-03-15T09:00:00Z"
  },
  {
    "media_name": "Politico",
    "media_url": "https://www.politico.com",
    "article_url": "https://www.politico.com/news/2024/03/15/example",
    "title": "Story Coverage",
    "publish_time": "2024-03-15T09:15:00Z"
  }
]
```

运行分析：

```bash
python3 scripts/analyze_first_publisher.py my-news.json
```

## 在Python中使用

```python
import json
from pathlib import Path
import sys

# 添加脚本目录到路径
script_dir = "/root/.openclaw/workspace/skills/first-publisher-analyzer/scripts"
sys.path.insert(0, script_dir)

from analyze_first_publisher import analyze_first_publisher_comprehensive

# 准备输入数据
related_news = [
    {
        "media_name": "路透社",
        "media_url": "https://www.reuters.com",
        "article_url": "https://www.reuters.com/world/article",
        "title": "News Title",
        "publish_time": "2024-03-15T10:30:00Z"
    },
    {
        "media_name": "BBC",
        "media_url": "https://www.bbc.com",
        "article_url": "https://www.bbc.com/news/article",
        "title": "BBC Coverage",
        "publish_time": "2024-03-15T11:00:00Z"
    }
]

# 执行分析
result = analyze_first_publisher_comprehensive(related_news)

# 输出结论
print(result["conclusion"])

# 输出媒体卡片
print(json.dumps(result["media_card"], ensure_ascii=False, indent=2))
```

## 单独使用各模块

### 仅识别首发媒体

```python
from scripts.identify_first_publisher import identify_first_publisher

result = identify_first_publisher(related_news)
print(f"首发媒体：{result['first_publisher']['media_name']}")
```

### 仅评估媒体

```python
from scripts.assess_media import generate_media_assessment

assessment = generate_media_assessment(
    media_url="https://www.reuters.com",
    article_url="https://www.reuters.com/world/article",
    media_name="路透社"
)

print(assessment["summary"])
```

## 扩展媒体数据库

编辑 `scripts/assess_media.py`，添加新的媒体：

```python
INTERNATIONAL_MEDIA_DB = {
    # ... 现有媒体 ...
    "your-media.com": {
        "name": "您的媒体名称",
        "established": 2020,
        "headquarters": "城市，国家",
        "owner": "机构名称",
        "political_leaning": "政治倾向描述",
        "credibility": 75,
        "audience_scale": "national",
        "media_type": "online",
        "description": "媒体描述..."
    },
}
```

## 支持更多时间格式

编辑 `scripts/identify_first_publisher.py`，在 `parse_timestamp()` 函数中添加：

```python
formats = [
    # ... 现有格式 ...
    "%Y/%m/%d %H:%M:%S",  # 添加新格式
]
```

## 常见问题

### Q: 如何处理中文时间格式？

A: 脚本已支持中文格式，如：`2024年3月15日 10:30`

### Q: 媒体不在数据库中怎么办？

A: 系统会从URL推断基本信息，但政治倾向和可信度度会标注为"置信度较低"。建议将常用媒体添加到数据库。

### Q: 如何提高分析准确性？

A:
1. 提供 `media_url` 字段（媒体主页URL）
2. 使用标准化的时间格式（ISO 8601推荐）
3. 确保时间戳准确无误

### Q: 输出格式可以自定义吗？

A: 可以。参考 `generate_conclusion()` 和 `generate_media_card()` 函数，根据需要修改输出格式。
