skills/ppt-station-skill/reference/workflow-parse-rebuild.md
wangyitong a65adcc2e5 Initial commit: merged, deduplicated, and vetted skill collection
Sources: extracted from two upstream archives (skill-repo, skills-main),
merged with the following policy:

- 15 broken symlinks (pointing to /Users/jameslee/.cc-switch/skills or
  ../../.agents/skills on a foreign machine) discarded
- 3 real name collisions with identical content (ai-pair, ifind-http-api,
  zhipu-websearch) kept as one copy
- Functional overlaps deduped keeping the strongest variant:
  - docx family: kept docx (official, full toolchain) + docx-cn
    (GB/T 9704 Chinese official-document constants),
    dropped docx_writer (no scripts, name collided with docx)
  - humanizer family: kept humanizer-zh (6 zh reference docs),
    dropped humanizer (en, redundant for CN workflow)
- Skills that only ran in a foreign environment removed:
  ablemind-ops, app-publish, hlb-design-system, openclaw-adj-skill,
  claude-driver
- alphapai excluded from this public repo because its SKILL.md hard-coded
  live credentials

Result: 25 skills, 572 files, ~7.5 MB.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-13 14:47:12 +08:00

81 lines
1.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Flow C — 解析→编辑→重建
适用于微调已有 PPT将 PPT 解析为 JSON 中间格式,编辑 JSON再重建为 PPT。
## 端到端流程
### Step 1: 解析 PPT 为 JSON
```bash
python scripts/parse_ppt.py aim/aim01.pptx /tmp/parsed.json
```
不指定输出文件则输出到 stdout:
```bash
python scripts/parse_ppt.py aim/aim01.pptx > /tmp/parsed.json
```
### Step 2: 编辑 JSON
JSON 中间格式包含每张幻灯片的所有形状信息。
```json
{
"slides": [
{
"slide_index": 0,
"shapes": [
{
"shape_id": 2,
"shape_type": "TEXT_BOX",
"text": "原始文本",
"left": 914400,
"top": 365125,
"width": 8229600,
"height": 457200
}
]
}
]
}
```
### Step 3: 重建 PPT
```bash
python scripts/rebuild_ppt.py /tmp/parsed.json /tmp/rebuilt.pptx
```
## JSON 中间格式结构
| 层级 | 字段 | 说明 |
|------|------|------|
| root | `slides[]` | 幻灯片数组 |
| slide | `slide_index` | 页码0-based|
| slide | `shapes[]` | 形状数组 |
| shape | `shape_id` | 形状 ID |
| shape | `shape_type` | 形状类型 |
| shape | `text` | 文本内容 |
| shape | `left/top/width/height` | 位置和尺寸EMU|
## 安全修改 vs 危险字段
### 安全修改
- `text`: 修改文本内容
- 表格单元格的文本值
- 形状的可见文本属性
### 危险字段(修改可能导致损坏)
- `shape_id`: 内部引用,修改会破坏关联
- `left/top/width/height`: 可修改但需保持 EMU 单位
- 图表的内嵌数据: 结构复杂,建议用 Flow A 替代
## 典型用例
1. **批量文字修改**: 解析 → 搜索替换文本 → 重建
2. **提取内容**: 解析 → 读取所有文本(不重建)
3. **结构检查**: 解析 → 分析形状布局