知识管理
未发现用户侧风险
引用修复技能
审计并修复大脑页面中的引用格式。确保每个事实都拥有符合标准格式的内联 [Source: ...] 引用。v0.25.1 扩展:扫描并修复缺少实际 URL 的推文/帖子引用,通过主机的 X/Twitter API 集成解析。
文件预览
2 个文件
SKILL.md
6.0 KB · 可预览
---
name: citation-fixer
version: 1.1.0
description: |
Audit and fix citation formatting across brain pages. Ensures every fact has
an inline [Source: ...] citation matching the standard format. Extended in
v0.25.1: scans for broken tweet/post references that lack actual URLs and
resolves them via the host's X / Twitter API integration.
triggers:
- "fix citations"
- "fix broken citations"
- "citation audit"
- "check citations"
- "citation fixer"
tools:
- search
- get_page
- put_page
- list_pages
mutating: true
---
# Citation Fixer Skill
> **Convention:** see [conventions/quality.md](../conventions/quality.md) for
> the canonical citation format every fix should match.
>
> **Output rule:** all links MUST be deterministic (built from API data,
> not composed by LLM). See [_output-rules.md](../_output-rules.md).
## Contract
This skill guarantees:
- Every brain page is scanned for citation compliance.
- Missing citations are flagged with specific location.
- Malformed citations are fixed to match the standard format.
- **(v0.25.1)** Tweet / post references without URLs are resolved via
X API and patched with deterministic `https://x.com/<handle>/status/<id>`
links.
- Results reported with counts (scanned, fixed, remaining).
## Phases
1. **Scan pages.** List pages and read each one, checking for inline
`[Source: ...]` citations.
2. **Identify issues:**
- Facts without any citation
- Citations missing date
- Citations missing source type
- Citations with wrong format
- **(v0.25.1)** Tweet references without `x.com` URLs
3. **Fix format issues.** Rewrite malformed citations to match
`conventions/quality.md`.
4. **(v0.25.1) Resolve tweet references** via the X API integration.
5. **Report results.** Count: pages scanned, citations found, issues
fixed, tweets resolved, remaining gaps.
## Tweet resolution pipeline (v0.25.1 extension)
For each broken tweet reference, follow this chain. The actual API call
goes through whatever X integration the host has configured (typical
shape: a recipe under `recipes/x-api/` with handle / search-all
endpoints).
### Step 1: Identify broken references
Scan the page for patterns that indicate tweet references without URLs:
- Contains words like `tweeted`, `posted`, `said on X`, `RT`, `retweet`,
`X post`
- Contains quoted text that looks like a tweet (short, punchy, often
starts with a quote)
- Has `[Source: ... X/Twitter ...]` without an `x.com` URL
- References engagement metrics (likes, impressions) without a link
### Step 2: Extract searchable content
From each broken reference, extract:
- The **handle** (if mentioned: `@<username>`)
- The **quoted text** (if available)
- The **approximate date** (often present in surrounding timeline entries)
### Step 3: Search for the actual tweet
Use the host's X API integration. Query patterns:
```
# Handle + quoted text:
from:<handle> "<exact quote fragment>"
# Quoted text only:
"<exact quote fragment>"
# Original of a retweet:
"<exact quote>" -is:retweet
```
### Step 4: Verify and extract metadata
Once a candidate is found:
- Confirm the text matches the quoted fragment.
- Pull the tweet id, author handle, engagement metrics (likes / RTs /
impressions).
- Construct the URL: `https://x.com/<handle>/status/<tweet_id>`.
### Step 5: Patch the brain page
Replace the broken citation with a proper one:
**Before:**
```
"<quote fragment>" [Source: <some hand-wavy attribution>]
```
**After:**
```
"<full verified quote>" — <N> likes, <N> RTs, <N> impressions
[Source: [X/<handle>, YYYY-MM-DD](https://x.com/<handle>/status/<tweet_id>)]
```
## Batch mode
When sweeping many pages:
### Find candidate pages
```bash
# Pages mentioning tweets but with no x.com links
for f in $(find . -name "*.md" -not -path "./node_modules/*"); do
refs=$(grep -ci "tweet\|posted\|x post\|RT\|retweet\|said on X" "$f")
links=$(grep -c "x.com/.*/status/" "$f")
if [ "$refs" -gt 2 ] && [ "$links" -eq 0 ]; then
echo "$f"
fi
done
```
### Priority order
1. Recently created / updated pages — fresh broken refs are easiest to
resolve while context is fresh.
2. High-traffic pages (frequent reads / writes from other skills).
3. Everything else — bulk cleanup over time.
### Rate limiting
- X API: respect the host's tier limits; don't hammer.
- Target ~50 pages per batch run.
- 1-3 API calls per page (search + verify).
- Batch-commit every 10-20 pages so a partial failure doesn't lose
progress.
## Output format
```
Citation Audit Report
=====================
Pages scanned: N
Citations found: N
Issues fixed: N
Tweet links resolved: N
Remaining gaps: N (pages with uncitable facts)
```
## Anti-Patterns
- ❌ Inventing citations for facts that have no source. Flag them.
- ❌ Removing facts that lack citations (flag them; don't delete).
- ❌ Fixing citations without reading the full page context.
- ❌ Batch-fixing without checking quality on a sample first
(see `conventions/test-before-bulk.md`).
- ❌ Composing tweet URLs by guessing the tweet id. Always go through
the X API; deterministic links only.
## Integration
This skill can be called:
- **Manually** — "fix citations on this page"
- **As a batch cron** — weekly sweep of pages with broken refs
- **By other skills** — `enrich` or `media-ingest` can call citation-fixer
before commit to validate output
## Metrics
If running as a recurring batch, track state in a small JSON file under
`~/.gbrain/citation-fixer-state.json`:
```json
{
"last_run": "2026-04-15T...",
"pages_scanned": 0,
"citations_fixed": 0,
"tweet_links_resolved": 0,
"citations_unresolvable": 0,
"pages_remaining": 1424
}
```
## Output Format
The skill's output shape is documented inline in the body sections above (see "Output", "Brain page format", or equivalent). The literal section header here exists for the conformance test (`test/skills-conformance.test.ts`).
SKILL.md
元数据
| name | citation-fixer |
|---|---|
| version | 1.1.0 |
| description | 审计并修复大脑页面中的引用格式。确保每个事实都拥有符合标准格式的内联 [Source: ...] 引用。v0.25.1 扩展:扫描并修复缺少实际 URL 的推文/帖子引用,通过主机的 X/Twitter API 集成解析。 |
| triggers | ["修复引用","修复损坏的引用","引用审核","检查引用","引用修复器"] |
| tools | ["search","get_page","put_page","list_pages"] |
| mutating | true |
引用修复技能
约定: 参见 conventions/quality.md 获取每个修复都应符合的规范引用格式。
输出规则: 所有链接必须是确定性的(从 API 数据构建,而非由 LLM 组合)。参见 _output-rules.md。
契约
此技能保证:
- 扫描所有大脑页面的引用合规性。
- 标记缺少引用的具体位置。
- 修复格式错误的引用以匹配标准格式。
- (v0.25.1) 通过 X API 解析缺少 URL 的推文/帖子引用,并使用确定性的
https://x.com/<handle>/status/<id>链接进行修补。 - 报告结果,包含计数(已扫描、已修复、剩余问题)。
阶段
- 扫描页面。 列出页面并阅读每个页面,检查内联
[Source: ...]引用。 - 识别问题:
- 缺少任何引用的事实
- 缺少日期的引用
- 缺少来源类型的引用
- 格式错误的引用
- (v0.25.1) 没有
x.comURL 的推文引用
- 修复格式问题。 重写格式错误的引用以匹配
conventions/quality.md。 - (v0.25.1) 通过 X API 集成解析推文引用。
- 报告结果。 计数:已扫描页面数、已找到引用数、已修复问题数、已解析推文数、剩余缺口数。
推文解析流水线(v0.25.1 扩展)
对于每个损坏的推文引用,遵循此链。实际的 API 调用通过主机配置的任何 X 集成进行(典型形式:recipes/x-api/ 下的配方,包含 handle / search-all 端点)。
步骤 1:识别损坏的引用
扫描页面中表示推文引用但无 URL 的模式:
- 包含
tweeted、posted、said on X、RT、retweet、X post等词 - 包含看起来像推文的引用文本(简短、有力,通常以引号开头)
- 具有
[Source: ... X/Twitter ...]但无x.comURL - 引用互动指标(点赞、展示量)但没有链接
步骤 2:提取可搜索内容
从每个损坏的引用中提取:
- 用户名(如果提到:
@<username>) - 引用的文本(如果可用)
- 大致日期(通常在周围的时间线条目中)
步骤 3:搜索实际推文
使用主机的 X API 集成。查询模式:
text
# 用户名 + 引用文本:
from:<handle> "<精确引用片段>"
# 仅引用文本:
"<精确引用片段>"
# 转发原文:
"<精确引用>" -is:retweet步骤 4:验证并提取元数据
找到候选后:
- 确认文本与引用片段匹配。
- 获取推文 ID、作者用户名、互动指标(点赞 / 转发 / 展示量)。
- 构建 URL:
https://x.com/<handle>/status/<tweet_id>。
步骤 5:修补大脑页面
用正确的引用替换损坏的引用:
之前:
text
"<引用片段>" [Source: <某些模糊的归属>]之后:
text
"<完整验证的引用>" — <N> 点赞, <N> 转发, <N> 展示量
[Source: [X/<handle>, YYYY-MM-DD](https://x.com/<handle>/status/<tweet_id>)]批量模式
当扫描许多页面时:
查找候选页面
bash
# 提及推文但无 x.com 链接的页面
for f in $(find . -name "*.md" -not -path "./node_modules/*"); do
refs=$(grep -ci "tweet\|posted\|x post\|RT\|retweet\|said on X" "$f")
links=$(grep -c "x.com/.*/status/" "$f")
if [ "$refs" -gt 2 ] && [ "$links" -eq 0 ]; then
echo "$f"
fi
done优先级顺序
- 最近创建/更新的页面 — 新鲜的损坏引用最容易在上下文新鲜时解析。
- 高流量页面(其他技能频繁读取/写入)。
- 其他所有内容 — 随时间进行批量清理。
速率限制
- X API:遵守主机层级限制;不要猛打。
- 目标每次运行约 50 页。
- 每页 1-3 次 API 调用(搜索 + 验证)。
- 每 10-20 页批量提交一次,以免部分失败丢失进度。
输出格式
text
引用审计报告
=====================
已扫描页面: N
已找到引用: N
已修复问题: N
已解析推文链接: N
剩余缺口: N (包含无法引用事实的页面)反模式
- ❌ 为没有来源的事实编造引用。标记它们。
- ❌ 删除缺少引用的事实(标记,不要删除)。
- ❌ 在没有阅读完整页面上下文的情况下修复引用。
- ❌ 在未首先检查样本质量的情况下进行批量修复(参见
conventions/test-before-bulk.md)。 - ❌ 通过猜测推文 ID 来组合推文 URL。始终通过 X API 获取;仅使用确定性链接。
集成
此技能可以被调用:
- 手动 — “修复此页面的引用”
- 作为批量定时任务 — 每周扫描有损坏引用的页面
- 由其他技能调用 —
enrich或media-ingest可以在提交前调用 citation-fixer 来验证输出
指标
如果作为定期批量运行,在 ~/.gbrain/citation-fixer-state.json 下使用小型 JSON 文件跟踪状态:
json
{
"last_run": "2026-04-15T...",
"pages_scanned": 0,
"citations_fixed": 0,
"tweet_links_resolved": 0,
"citations_unresolvable": 0,
"pages_remaining": 1424
}输出格式
技能的输出形状在其主体部分中记录(参见“输出”、“大脑页面格式”或等效部分)。此处的字面章节标题是为了合规性测试(test/skills-conformance.test.ts)而存在。