科研技能库/文献编目
知识管理
未发现用户侧风险

文献编目

按主题、方法和结论对文献进行分类与组织;适用于批量读取包含 PDF、MD、DOCX、TXT 文件的文件夹,并输出结构化 CSV,辅助文献综述与注释管理。

文件预览

4 个文件
assets
references
SKILL.md
6.3 KB · 可预览
---
name: bibliography
description: Classifies and organizes literature by theme, method, and conclusion; use when you need to batch-read a folder of PDF/MD/DOCX/TXT files and output a structured CSV for literature reviews and annotation management.
license: MIT
author: AIPOCH
---
> **Source**: [https://github.com/aipoch/medical-research-skills](https://github.com/aipoch/medical-research-skills)

# Bibliography

## When to Use

- You are conducting a literature review and need consistent summaries plus structured metadata (theme/method/conclusion) across many papers.
- You have a mixed-format reading folder (`.pdf`, `.md`, `.docx`, `.txt`) and want a single CSV for downstream analysis (e.g., Excel, R, Python).
- You need to organize annotations by **keywords (theme)**, **experimental methods (method)**, and **key conclusions (conclusion)**.
- You want a two-step pipeline: first generate a human-readable summary Markdown, then generate a machine-friendly CSV from that Markdown.
- You need robust handling of PDFs by converting them to Markdown first (via `pdf-extract`) and then using **only Markdown content** for extraction.

## Key Features

- Batch scans an input directory for `.pdf`, `.md`, `.docx`, and `.txt` literature files.
- Converts PDFs to Markdown via `pdf-extract`, then **ignores non-Markdown artifacts** (e.g., image folders).
- Extracts and normalizes, per document:
  - Title
  - Summary (prefer original abstract)
  - Keywords (theme)
  - Experimental Methods (method names only)
  - Key Conclusions (single sentence)
  - Commentary (one-sentence, tactful evaluation)
- Produces exactly **two outputs**:
  1. A consolidated **Summary Markdown** saved under `outputs/`
  2. A single **CSV** generated from that Summary Markdown
- Enforces UTF-8 output to prevent garbled characters; fills missing fields with `"Not recognized"` instead of leaving blanks.
- Uses the CSV field order and headers defined in `assets/bibliography_template.csv`.

## Dependencies

- `pdf-extract` (version: not specified; required when PDFs are present)
- Input formats supported (no external version constraints specified):
  - PDF
  - Markdown (`.md`)
  - DOCX (`.docx`)
  - Plain text (`.txt`)

## Example Usage

### Goal
Read all literature files in a folder, generate a consolidated summary Markdown, then generate a CSV following `assets/bibliography_template.csv`.

### Inputs
- Input directory (example): `./inputs/literature/`
- Output directory: `./outputs/`
- Output CSV path (example): `./outputs/bibliography.csv`

### Expected Outputs (exactly two files)
- `./outputs/bibliography_summary.md`
- `./outputs/bibliography.csv`

### Example Summary Markdown Structure (generated first)

```md
# Bibliography Summary

## Document 1
- Title: <Title>
- Summary: <Prefer the original Abstract; if missing, use the closest equivalent section>
- Keywords: keyword1 | keyword2 | keyword3
- Experimental Methods: <method1; method2; ... (names only)>
- Key Conclusions: <one sentence covering all main points>
- Commentary: <one tactful sentence>

## Document 2
...
```

### Example CSV (generated from the Summary Markdown)

The CSV must follow the header order defined in:

- `assets/bibliography_template.csv`

Rules:
- One row per document.
- No empty cells; use `Not recognized` when extraction fails.
- Save as UTF-8.

## Implementation Details

### 1) Input Reading and Normalization

- Traverse the input directory and process files with extensions:
  - `.pdf`, `.md`, `.docx`, `.txt`

**PDF handling**
- If PDFs exist, convert them to Markdown using `pdf-extract`.
- Use only the generated `.md` content; ignore image directories or other byproducts.
- Locate `pdf-extract` as follows:
  1. First, look for a sibling skill directory containing `SKILL.md` at the same level as this skill’s parent directory.
  2. If not found, ask the user to confirm the actual `pdf-extract` path.

**DOCX handling**
- Extract body text while preserving title/paragraph order as much as possible.

**MD/TXT handling**
- Read text directly.
- If garbled characters appear or key fields cannot be recognized, attempt to detect and read using the original encoding (commonly `GB18030` / `GBK`) before extraction.

### 2) Generate the Summary Markdown First (Single Source of Truth)

Before producing the CSV, generate a consolidated Summary Markdown containing, for each document:

- **Title**
- **Summary**
  - Prefer the original **Abstract**.
  - If no “Abstract” exists, use the closest equivalent section (e.g., “Summary”, “Highlights”, or an “Objective–Method–Result–Conclusion” style segment).
- **Keywords**
- **Experimental Methods**
- **Key Conclusions**
- **Commentary**
  - Exactly one sentence.
  - Avoid harsh criticism; if the work has low value, use tactful phrasing.

This Summary Markdown must be saved with **UTF-8** encoding and stored under `outputs/`. The CSV must be generated **only** from this Markdown (not directly from raw files).

### 3) Field Extraction Rules (Theme / Method / Conclusion)

- **Keywords (theme)**
  - Prefer the original keywords from the document.
  - Separate multiple keywords with `|`.
  - If no keywords are found, generate **3–5** keyword phrases based on the abstract and append:
    - `(generated based on abstract)`
- **Experimental Methods (method)**
  - Output method names only (no long descriptions).
- **Key Conclusions (conclusion)**
  - One sentence that covers all main points.

### 4) CSV Output Constraints

- Output exactly **one** CSV file at the end.
- CSV field order and headers must match `assets/bibliography_template.csv`.
- Encoding must be **UTF-8** to avoid garbled characters.
- If any field cannot be extracted, write `Not recognized` (never leave empty).
- Only two files may be generated in total:
  1. Summary Markdown
  2. CSV
- No temporary/intermediate/auxiliary files may be left behind (including extracted text dumps, caches, logs, images, backups). If conversion/extraction requires intermediate artifacts, keep them in memory or ensure all non-target files are deleted before final output.
- Do not use PowerShell to directly write/manipulate CSV/Markdown to avoid encoding/newline issues; always generate and save using UTF-8.

### Reference

- Detailed rules and field descriptions: `references/guide.md`

SKILL.md

元数据
name文献编目
description按主题、方法和结论对文献进行分类与组织;适用于批量读取包含 PDF、MD、DOCX、TXT 文件的文件夹,并输出结构化 CSV,辅助文献综述与注释管理。
licenseMIT
authorAIPOCH

来源: https://github.com/aipoch/medical-research-skills

文献编目

使用场景

  • 您正在进行文献综述,需要为大量文献生成一致的摘要以及结构化元数据(主题/方法/结论)。
  • 您有一个包含多种格式(.pdf、.md、.docx、.txt)的阅读文件夹,希望为后续分析(例如 Excel、R、Python)生成一个统一的 CSV 文件。
  • 您需要按关键词(主题)、**实验方法(方法)和关键结论(结论)**来组织注释。
  • 您希望采用两步流水线:首先生成人类可读的摘要 Markdown,再基于该 Markdown 生成机器友好的 CSV。
  • 您需要稳健的 PDF 处理方式:先将 PDF 转换为 Markdown(通过 pdf-extract),然后仅使用 Markdown 内容进行信息提取。

主要特性

  • 批量扫描输入目录中的 .pdf、.md、.docx 和 .txt 文献文件。
  • 通过 pdf-extract 将 PDF 转换为 Markdown,然后忽略非 Markdown 产物(如图像文件夹)。
  • 对每份文档提取并规范化以下字段:
    • 标题
    • 摘要(优先采用原始摘要)
    • 关键词(主题)
    • 实验方法(仅方法名称)
    • 关键结论(一句话概括)
    • 评论(一句委婉的评价)
  • 仅生成两个输出:
    1. 一份合并的摘要 Markdown,保存在 outputs/ 下
    2. 一个从该摘要 Markdown 生成的单一 CSV 文件
  • 强制使用 UTF-8 编码输出,防止乱码;缺失字段用 "未识别" 填充,避免留空。
  • CSV 字段顺序和表头遵循 assets/bibliography_template.csv 中的定义。

依赖项

  • pdf-extract(版本未指定;当存在 PDF 文件时必需)
  • 支持的输入格式(无外部版本约束):
    • PDF
    • Markdown(.md)
    • DOCX(.docx)
    • 纯文本(.txt)

使用示例

目标

读取某个文件夹中的所有文献文件,生成一份合并的摘要 Markdown,然后根据 assets/bibliography_template.csv 生成 CSV。

输入

  • 输入目录(示例):./inputs/literature/
  • 输出目录:./outputs/
  • 输出 CSV 路径(示例):./outputs/bibliography.csv

预期输出(仅两个文件)

  • ./outputs/bibliography_summary.md
  • ./outputs/bibliography.csv

示例摘要 Markdown 结构(首先生成)

md
# 文献摘要汇总

## 文档 1
- 标题: <标题>
- 摘要: <优先使用原始摘要;若无,则使用最相近的章节>
- 关键词: 关键词1 | 关键词2 | 关键词3
- 实验方法: <方法1; 方法2; ... (仅名称)>
- 关键结论: <一句涵盖所有要点的话>
- 评论: <一句委婉评价>

## 文档 2
...

示例 CSV(由摘要 Markdown 生成)

CSV 必须遵循如下定义的表头顺序:

  • assets/bibliography_template.csv

规则:

  • 每个文档占用一行。
  • 无空白单元格;提取失败时使用 未识别。
  • 以 UTF-8 编码保存。

实现细节

1) 输入读取与规范化

  • 遍历输入目录,处理以下扩展名的文件:
    • .pdf、.md、.docx、.txt

PDF 处理

  • 如果存在 PDF,则使用 pdf-extract 将其转换为 Markdown。
  • 仅使用生成的 .md 内容;忽略图像目录或其他副产品。
  • 按以下顺序查找 pdf-extract:
    1. 首先,在本技能父目录的同级目录下查找包含 SKILL.md 的兄弟技能目录。
    2. 若未找到,请用户确认 pdf-extract 的实际路径。

DOCX 处理

  • 提取正文文本,尽可能保留标题和段落顺序。

MD/TXT 处理

  • 直接读取文本。
  • 若出现乱码或关键字段无法识别,先尝试检测并以原始编码(常见为 GB18030 / GBK)读取,再进行提取。

2) 首先生成摘要 Markdown(单一事实来源)

在生成 CSV 之前,先产生一份合并的摘要 Markdown,对每个文档包含:

  • 标题
  • 摘要
    • 优先使用原始摘要。
    • 若无“摘要”,则使用最相近的章节(如“总结”、“亮点”或“目的-方法-结果-结论”式段落)。
  • 关键词
  • 实验方法
  • 关键结论
  • 评论
    • 仅限一句话。
    • 避免严厉批评;若工作价值较低,使用委婉措辞。

该摘要 Markdown 必须以 UTF-8 编码保存,存放于 outputs/ 下。CSV 必须仅根据此 Markdown 生成(而非直接源自原始文件)。

3) 字段提取规则(主题 / 方法 / 结论)

  • 关键词(主题)
    • 优先使用文档中的原始关键词。
    • 多个关键词用 | 分隔。
    • 若未找到关键词,基于摘要生成 3-5 个关键词短语,并追加:
      • (基于摘要生成)
  • 实验方法(方法)
    • 仅输出方法名称(无详细描述)。
  • 关键结论(结论)
    • 一句涵盖所有要点的话。

4) CSV 输出约束

  • 最终仅输出一个 CSV 文件。
  • CSV 字段顺序和表头必须与 assets/bibliography_template.csv 一致。
  • 编码必须为 UTF-8,防止乱码。
  • 如果某个字段无法提取,则填入 未识别(不得留空)。
  • 总共只能生成两个文件:
    1. 摘要 Markdown
    2. CSV
  • 不得残留任何临时/中间/辅助文件(包括提取的文本转储、缓存、日志、图像、备份)。如果转换/提取过程需要中间产物,应将其保留在内存中,或确保在最终输出前删除所有非目标文件。
  • 不要使用 PowerShell 直接写入或操作 CSV/Markdown,以避免编码和换行问题;始终使用 UTF-8 生成并保存。

参考

  • 详细规则与字段说明:references/guide.md