科研技能库/学术论文评审
论文阅读
未发现用户侧风险

学术论文评审

当用户请求评审、分析、批评或总结学术论文、研究文章、预印本或科学出版物时使用此技能。支持全面的结构化评审,涵盖方法论评估、贡献评价、文献定位和建设性反馈生成。触发条件包括涉及论文网址、上传的PDF、arXiv链接或诸如“评审这篇论文”、“分析这项研究”、“总结这项研究”或“撰写同行评审”之类的请求。

文件预览

1 个文件
SKILL.md
11.9 KB · 可预览
---
name: academic-paper-review
description: Use this skill when the user requests to review, analyze, critique, or summarize academic papers, research articles, preprints, or scientific publications. Supports comprehensive structured reviews covering methodology assessment, contribution evaluation, literature positioning, and constructive feedback generation. Trigger on queries involving paper URLs, uploaded PDFs, arXiv links, or requests like "review this paper", "analyze this research", "summarize this study", or "write a peer review".
---

# Academic Paper Review Skill

## Overview

This skill produces structured, peer-review-quality analyses of academic papers and research publications. It follows established academic review standards used by top-tier venues (NeurIPS, ICML, ACL, Nature, IEEE) to provide rigorous, constructive, and balanced assessments.

The review covers **summary, strengths, weaknesses, methodology assessment, contribution evaluation, literature positioning, and actionable recommendations** — all grounded in evidence from the paper itself.

## Core Capabilities

- Parse and comprehend academic papers from uploaded PDFs or fetched URLs
- Generate structured reviews following top-venue review templates
- Assess methodology rigor (experimental design, statistical validity, reproducibility)
- Evaluate novelty and significance of contributions
- Position the work within the broader research landscape via targeted literature search
- Identify limitations, gaps, and potential improvements
- Produce both detailed review and concise executive summary formats
- Support papers in any scientific domain (CS, biology, physics, social sciences, etc.)

## When to Use This Skill

**Always load this skill when:**

- User provides a paper URL (arXiv, DOI, conference proceedings, journal link)
- User uploads a PDF of a research paper or preprint
- User asks to "review", "analyze", "critique", "assess", or "summarize" a research paper
- User wants to understand the strengths and weaknesses of a study
- User requests a peer-review-style evaluation of academic work
- User asks for help preparing a review for a conference or journal submission

## Review Methodology

### Phase 1: Paper Comprehension

Thoroughly read and understand the paper before forming any judgments.

#### Step 1.1: Identify Paper Metadata

Extract and record:

| Field | Description |
|-------|-------------|
| **Title** | Full paper title |
| **Authors** | Author list and affiliations |
| **Venue / Status** | Publication venue, preprint server, or submission status |
| **Year** | Publication or submission year |
| **Domain** | Research field and subfield |
| **Paper Type** | Empirical, theoretical, survey, position paper, systems paper, etc. |

#### Step 1.2: Deep Reading Pass

Read the paper systematically:

1. **Abstract & Introduction** — Identify the claimed contributions and motivation
2. **Related Work** — Note how authors position their work relative to prior art
3. **Methodology** — Understand the proposed approach, model, or framework in detail
4. **Experiments / Results** — Examine datasets, baselines, metrics, and reported outcomes
5. **Discussion & Limitations** — Note any self-identified limitations
6. **Conclusion** — Compare concluded claims against actual evidence presented

#### Step 1.3: Key Claims Extraction

List the paper's main claims explicitly:

```
Claim 1: [Specific claim about contribution or finding]
Evidence: [What evidence supports this claim in the paper]
Strength: [Strong / Moderate / Weak]

Claim 2: [...]
...
```

### Phase 2: Critical Analysis

#### Step 2.1: Literature Context Search

Use web search to understand the research landscape:

```
Search queries:
- "[paper topic] state of the art [current year]"
- "[key method name] comparison benchmark"
- "[authors] previous work [topic]"
- "[specific technique] limitations criticism"
- "survey [research area] recent advances"
```

Use `web_fetch` on key related papers or surveys to understand where this work fits.

#### Step 2.2: Methodology Assessment

Evaluate the methodology using the following framework:

| Criterion | Questions to Ask | Rating |
|-----------|-----------------|--------|
| **Soundness** | Is the approach technically correct? Are there logical flaws? | 1-5 |
| **Novelty** | What is genuinely new vs. incremental improvement? | 1-5 |
| **Reproducibility** | Are details sufficient to reproduce? Code/data available? | 1-5 |
| **Experimental Design** | Are baselines fair? Are ablations adequate? Are datasets appropriate? | 1-5 |
| **Statistical Rigor** | Are results statistically significant? Error bars reported? Multiple runs? | 1-5 |
| **Scalability** | Does the approach scale? Are computational costs discussed? | 1-5 |

#### Step 2.3: Contribution Significance Assessment

Evaluate the significance level:

| Level | Description | Criteria |
|-------|-------------|----------|
| **Landmark** | Fundamentally changes the field | New paradigm, widely applicable breakthrough |
| **Significant** | Strong contribution advancing the state of the art | Clear improvement with solid evidence |
| **Moderate** | Useful contribution with some limitations | Incremental but valid improvement |
| **Marginal** | Minimal advance over existing work | Small gains, narrow applicability |
| **Below threshold** | Does not meet publication standards | Fundamental flaws, insufficient evidence |

#### Step 2.4: Strengths and Weaknesses Analysis

For each strength or weakness, provide:
- **What**: Specific observation
- **Where**: Section/figure/table reference
- **Why it matters**: Impact on the paper's claims or utility

### Phase 3: Review Synthesis

#### Step 3.1: Assemble the Structured Review

Produce the final review using the template below.

## Review Output Template

```markdown
# Paper Review: [Paper Title]

## Paper Metadata
- **Authors**: [Author list]
- **Venue**: [Publication venue or preprint server]
- **Year**: [Year]
- **Domain**: [Research field]
- **Paper Type**: [Empirical / Theoretical / Survey / Systems / Position]

## Executive Summary

[2-3 paragraph summary of the paper's core contribution, approach, and main findings.
State your overall assessment upfront: what the paper does well, where it falls short,
and whether the contribution is sufficient for the claimed venue/impact level.]

## Summary of Contributions

1. [First claimed contribution — one sentence]
2. [Second claimed contribution — one sentence]
3. [Additional contributions if any]

## Strengths

### S1: [Concise strength title]
[Detailed explanation with specific references to sections, figures, or tables in the paper.
Explain WHY this is a strength and its significance.]

### S2: [Concise strength title]
[...]

### S3: [Concise strength title]
[...]

## Weaknesses

### W1: [Concise weakness title]
[Detailed explanation with specific references. Explain the impact of this weakness on
the paper's claims. Suggest how it could be addressed.]

### W2: [Concise weakness title]
[...]

### W3: [Concise weakness title]
[...]

## Methodology Assessment

| Criterion | Rating (1-5) | Assessment |
|-----------|:---:|------------|
| Soundness | X | [Brief justification] |
| Novelty | X | [Brief justification] |
| Reproducibility | X | [Brief justification] |
| Experimental Design | X | [Brief justification] |
| Statistical Rigor | X | [Brief justification] |
| Scalability | X | [Brief justification] |

## Questions for the Authors

1. [Specific question that would clarify a concern or ambiguity]
2. [Question about methodology choices or alternative approaches]
3. [Question about generalizability or practical applicability]

## Minor Issues

- [Typos, formatting issues, unclear figures, notation inconsistencies]
- [Missing references that should be cited]
- [Suggestions for improved clarity]

## Literature Positioning

[How does this work relate to the current state of the art?
Are key related works cited? Are comparisons fair and comprehensive?
What important related work is missing?]

## Recommendations

**Overall Assessment**: [Accept / Weak Accept / Borderline / Weak Reject / Reject]

**Confidence**: [High / Medium / Low] — [Justification for confidence level]

**Contribution Level**: [Landmark / Significant / Moderate / Marginal / Below threshold]

### Actionable Suggestions for Improvement
1. [Specific, constructive suggestion]
2. [Specific, constructive suggestion]
3. [Specific, constructive suggestion]
```

## Review Principles

### Constructive Criticism
- **Always suggest how to fix it** — Don't just point out problems; propose solutions
- **Give credit where due** — Acknowledge genuine contributions even in flawed papers
- **Be specific** — Reference exact sections, equations, figures, and tables
- **Separate minor from major** — Distinguish fatal flaws from fixable issues

### Objectivity Standards
- ❌ "This paper is poorly written" (vague, unhelpful)
- ✅ "Section 3.2 introduces notation X without formal definition, making the proof in Theorem 1 difficult to follow. Consider adding a notation table after the problem formulation." (specific, actionable)

### Ethical Review Practices
- Do NOT dismiss work based on author reputation or affiliation
- Evaluate the work on its own merits
- Flag potential ethical concerns (bias in datasets, dual-use implications) constructively
- Maintain confidentiality of unpublished work

## Adaptation by Paper Type

| Paper Type | Focus Areas |
|------------|-------------|
| **Empirical** | Experimental design, baselines, statistical significance, ablations, reproducibility |
| **Theoretical** | Proof correctness, assumption reasonableness, tightness of bounds, connection to practice |
| **Survey** | Comprehensiveness, taxonomy quality, coverage of recent work, synthesis insights |
| **Systems** | Architecture decisions, scalability evidence, real-world deployment, engineering contributions |
| **Position** | Argument coherence, evidence for claims, impact potential, fairness of characterizations |

## Common Pitfalls to Avoid

- ❌ Reviewing the paper you wish was written instead of the paper that was submitted
- ❌ Demanding additional experiments that are unreasonable in scope
- ❌ Penalizing the paper for not solving a different problem
- ❌ Being overly influenced by writing quality versus technical contribution
- ❌ Treating absence of comparison to your own work as a weakness
- ❌ Providing only a summary without critical analysis

## Quality Checklist

Before finalizing the review, verify:

- [ ] Paper was read completely (not just abstract and introduction)
- [ ] All major claims are identified and evaluated against evidence
- [ ] At least 3 strengths and 3 weaknesses are provided with specific references
- [ ] The methodology assessment table is complete with ratings and justifications
- [ ] Questions for authors target genuine ambiguities, not rhetorical critiques
- [ ] Literature search was conducted to contextualize the contribution
- [ ] Recommendations are actionable and constructive
- [ ] The overall assessment is consistent with the identified strengths and weaknesses
- [ ] The review tone is professional and respectful
- [ ] Minor issues are separated from major concerns

## Output Format

- Output the complete review in **Markdown** format
- Save the review to `/mnt/user-data/outputs/review-{paper-topic}.md` when working in sandbox
- Present the review to the user using the `present_files` tool

## Notes

- This skill complements the `deep-research` skill — load both when the user wants the paper reviewed in the context of the broader field
- For papers behind paywalls, work with whatever content is accessible (abstract, publicly available versions, preprint mirrors)
- Adapt the review depth to the user's needs: a brief assessment for quick triage versus a full review for submission preparation
- When reviewing multiple papers comparatively, maintain consistent criteria across all reviews
- Always disclose limitations of your review (e.g., "I could not verify the proofs in Appendix B in detail")

SKILL.md

元数据
nameacademic-paper-review
description当用户请求评审、分析、批评或总结学术论文、研究文章、预印本或科学出版物时使用此技能。支持全面的结构化评审,涵盖方法论评估、贡献评价、文献定位和建设性反馈生成。触发条件包括涉及论文网址、上传的PDF、arXiv链接或请求如“评审这篇论文”、“分析这项研究”、“总结这项研究”或“撰写同行评审”。

学术论文评审技能

概述

本技能针对学术论文和研究出版物生成结构化的、同行评审质量的分析。遵循顶级会议(NeurIPS、ICML、ACL、Nature、IEEE)使用的既定学术评审标准,提供严谨、建设性和平衡的评估。

评审覆盖摘要、优点、缺点、方法评估、贡献评价、文献定位和可操作的建议——全部基于论文本身提供的证据。

核心能力

  • 从上传的PDF或获取的URL中解析并理解学术论文
  • 按照顶级评审模板生成结构化评审
  • 评估方法严谨性(实验设计、统计有效性、可重复性)
  • 评价贡献的新颖性和重要性
  • 通过针对性文献搜索,将工作置于更广泛的研究背景中
  • 识别局限性、不足和可能的改进
  • 生成详细评审和简洁的执行摘要两种格式
  • 支持任何科学领域的论文(计算机科学、生物学、物理学、社会科学等)

何时使用此技能

在以下情况始终加载此技能:

  • 用户提供论文URL(arXiv、DOI、会议论文集、期刊链接)
  • 用户上传研究论文或预印本PDF
  • 用户要求“评审”、“分析”、“批评”、“评估”或“总结”一篇研究论文
  • 用户想了解某项研究的优点和缺点
  • 用户请求对学术工作进行同行评审式的评价
  • 用户请求帮助准备会议或期刊投稿的审稿意见

评审方法论

第一阶段:论文理解

在形成任何判断之前,彻底阅读并理解论文。

步骤1.1:识别论文元数据

提取并记录:

字段描述
标题完整论文标题
作者作者列表及其所属机构
发表地点/状态发表刊物、预印本服务器或投稿状态
年份发表或投稿年份
领域研究领域和子领域
论文类型实证、理论、综述、立场论文、系统论文等

步骤1.2:深度阅读

系统阅读论文:

  1. 摘要与引言——识别声称的贡献和动机
  2. 相关工作——注意作者如何将其工作与现有研究进行对比
  3. 方法论——详细理解提出的方法、模型或框架
  4. 实验/结果——检查数据集、基线、指标和报告的结果
  5. 讨论与局限性——注意作者自己指出的局限性
  6. 结论——将总结出的主张与实际呈现的证据进行对比

步骤1.3:提取关键主张

明确列出论文的主要主张:

text
声称1:[关于贡献或发现的具体主张]
证据:[论文中支持该主张的证据]
强度:[强/中/弱]

声称2:[...]
...

第二阶段:批判性分析

步骤2.1:文献背景搜索

使用网络搜索了解研究背景:

text
搜索查询:
- “【论文主题】当前最新进展 【当前年份】”
- “【关键方法名称】对比基准测试”
- “【作者】先前工作 【主题】”
- “【特定技术】局限性批评”
- “综述 【研究领域】 近期进展”

使用 web_fetch 获取关键相关论文或综述,了解本工作所处位置。

步骤2.2:方法论评估

使用以下框架评估方法论:

标准待问问题评分
正确性方法在技术上是否正确?是否存在逻辑缺陷?1-5
新颖性什么是真正新颖的,还是渐进式改进?1-5
可重复性提供足够细节以复现吗?代码/数据可用吗?1-5
实验设计基线对比公平吗?消融实验充分吗?数据集合适吗?1-5
统计严谨性结果有统计显著性吗?报告了误差线吗?多次运行?1-5
可扩展性方法可扩展吗?讨论了计算成本吗?1-5

步骤2.3:贡献重要性评估

评估重要性级别:

级别描述标准
里程碑根本改变领域新范式,广泛适用的突破
显著强有力地推进了最先进水平有明显提升且证据扎实
中等有用的贡献但有一些局限性增量但有效的改进
边缘对现有工作提升最小小幅提升,应用范围窄
低于标准不符合发表标准根本性缺陷,证据不足

步骤2.4:优缺点分析

对每个优点或缺点,提供:

  • 什么:具体观察
  • 哪里:章节/图表/表格引用
  • 为什么重要:对论文主张或效用的影响

第三阶段:评审综合

步骤3.1:汇集结构化评审

使用下面的模板生成最终评审。

评审输出模板

markdown
# 论文评审:[论文标题]

## 论文元数据
- **作者**:[作者列表]
- **发表地点**:[发表刊物或预印本服务器]
- **年份**:[年份]
- **领域**:[研究领域]
- **论文类型**:[实证/理论/综述/系统/立场]

## 执行摘要

[2-3段总结论文的核心贡献、方法和主要发现。
首先给出总体评估:论文做得好的地方、不足之处,以及贡献是否足以达到所声称的发表场合/影响力水平。]

## 贡献总结

1. [第一个声称的贡献——一句话]
2. [第二个声称的贡献——一句话]
3. [若有更多贡献]

## 优点

### S1:[简明优点标题]
[详细解释,引用论文中具体的章节、图表或表格。
解释这为什么是一个优点及其重要性。]

### S2:[简明优点标题]
[...]

### S3:[简明优点标题]
[...]

## 缺点

### W1:[简明缺点标题]
[详细解释,并引用具体位置。说明该缺点对论文主张的影响。建议如何解决。]

### W2:[简明缺点标题]
[...]

### W3:[简明缺点标题]
[...]

## 方法论评估

| 标准 | 评分(1-5) | 评估 |
|-----------|:---:|------------|
| 正确性 | X | [简要理由] |
| 新颖性 | X | [简要理由] |
| 可重复性 | X | [简要理由] |
| 实验设计 | X | [简要理由] |
| 统计严谨性 | X | [简要理由] |
| 可扩展性 | X | [简要理由] |

## 向作者提问

1. [能澄清某个疑虑或模糊之处的具体问题]
2. [关于方法论选择或替代方案的问题]
3. [关于泛化能力或实际应用性的问题]

## 次要问题

- [笔误、格式问题、图表不清晰、符号不一致]
- [应引用但遗漏的参考文献]
- [改进清晰度的建议]

## 文献定位

[该工作与当前最先进水平的关系?
关键相关工作是否被引用?对比是否公平且全面?
遗漏了哪些重要的相关工作?]

## 建议

**总体评估**:[接受/弱接受/边界/弱拒绝/拒绝]

**信心**:[高/中/低]——[信心水平的理由]

**贡献水平**:[里程碑/显著/中等/边缘/低于标准]

### 可操作的改进建议
1. [具体、建设性的建议]
2. [具体、建设性的建议]
3. [具体、建设性的建议]

评审原则

建设性批评

  • 总是建议如何修复问题——不要只指出问题,要提出解决方案
  • 给予应有的肯定——即使论文有缺陷,也要认可真正的贡献
  • 具体明确——引用确切的章节、等式、图表和表格
  • 区分次要和主要问题——将致命缺陷和可修复的问题区分开

客观性标准

  • ❌ “这篇论文写得不好”(模糊、无帮助)
  • ✅ “第3.2节引入了符号X但没有正式定义,导致定理1的证明难以理解。建议在问题表述后添加符号表。”(具体、可操作)

评审伦理规范

  • 不要因为作者声誉或所属机构而轻视工作
  • 根据工作本身的价值进行评估
  • 以建设性方式指出潜在的伦理问题(数据集中的偏见、双重用途影响)
  • 保持未发表工作的保密性

按论文类型的适配

论文类型关注重点
实证实验设计、基线、统计显著性、消融实验、可重复性
理论证明正确性、假设合理性、界的紧密性、与实践的联系
综述全面性、分类法质量、近期工作覆盖度、综合洞察
系统架构决策、可扩展性证据、实际部署、工程贡献
立场论点连贯性、主张的证据、影响潜力、描述的公平性

常见陷阱避免

  • ❌ 评审你希望写成的论文,而不是实际提交的论文
  • ❌ 要求进行范围不合理的大量额外实验
  • ❌ 因为论文没有解决另一个不同问题而扣分
  • ❌ 过分受写作质量影响而非技术贡献
  • ❌ 将未与你自己工作比较视为缺点
  • ❌ 只提供总结而没有批判性分析

质量检查清单

最终确定评审前,验证:

  • 论文已完全阅读(不仅是摘要和引言)
  • 所有主要主张均已识别并对照证据评估
  • 至少提供3个优点和3个缺点,并附有具体引用
  • 方法论评估表完整,包含评分和理由
  • 对作者的问题针对真正的模糊之处,而非修辞性批评
  • 已进行文献搜索以定位贡献
  • 建议可操作且具有建设性
  • 总体评估与识别出的优点和缺点一致
  • 评审语气专业且尊重
  • 次要问题与主要关切分开

输出格式

  • 以Markdown格式输出完整评审
  • 在沙箱中工作时,将评审保存到 /mnt/user-data/outputs/review-{论文主题}.md
  • 使用 present_files 工具向用户展示评审

备注

  • 本技能与 deep-research 技能互补——当用户希望将论文置于更广泛领域背景下进行评审时,可同时加载
  • 对于付费墙后的论文,尽量利用可访问的内容(摘要、公开版本、预印本镜像)
  • 根据用户需求调整评审深度:快速筛选用简要评估,投稿准备用完整评审
  • 当比较多篇论文时,在所有评审中保持标准一致
  • 始终披露你评审的局限性(例如,“我未能详细验证附录B中的证明”)