文献综述
使用多个学术数据库(PubMed、arXiv、bioRxiv、Semantic Scholar 等)进行全面、系统的文献综述。用于进行系统文献综述、荟萃分析、研究综合或跨生物医学、科学和技术领域的全面文献检索。生成格式专业的 Markdown 文档和 PDF,附带验证过的引用,支持多种引用风格(APA、Nature、Vancouver 等)。
文件预览
---
name: literature-review
description: Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.). This skill should be used when conducting systematic literature reviews, meta-analyses, research synthesis, or comprehensive literature searches across biomedical, scientific, and technical domains. Creates professionally formatted markdown documents and PDFs with verified citations in multiple citation styles (APA, Nature, Vancouver, etc.).
allowed-tools: Read Write Edit Bash
license: MIT license
metadata:
skill-author: K-Dense Inc.
---
# Literature Review
## Overview
Conduct systematic, comprehensive literature reviews following rigorous academic methodology. Search multiple literature sources, synthesize findings thematically, verify citations for accuracy, and generate professional output documents in markdown and PDF formats.
## When to Use This Skill
Use this skill when:
- Conducting a systematic literature review for research or publication
- Synthesizing current knowledge on a specific topic across multiple sources
- Performing meta-analysis or scoping reviews
- Writing the literature review section of a research paper or thesis
- Investigating the state of the art in a research domain
- Identifying research gaps and future directions
- Requiring verified citations and professional formatting
## Core Workflow
Literature reviews follow a structured, multi-phase workflow:
### Phase 1: Planning and Scoping
1. **Define Research Question**: Use PICO framework (Population, Intervention, Comparison, Outcome) for clinical/biomedical reviews
- Example: "What is the efficacy of CRISPR-Cas9 (I) for treating sickle cell disease (P) compared to standard care (C)?"
2. **Establish Scope and Objectives**:
- Define clear, specific research questions
- Determine review type (narrative, systematic, scoping, meta-analysis)
- Set boundaries (time period, geographic scope, study types)
3. **Develop Search Strategy**:
- Identify 2-4 main concepts from research question
- List synonyms, abbreviations, and related terms for each concept
- Plan Boolean operators (AND, OR, NOT) to combine terms
- Select minimum 3 complementary databases
4. **Set Inclusion/Exclusion Criteria**:
- Date range (e.g., last 10 years: 2015-2024)
- Language (typically English, or specify multilingual)
- Publication types (peer-reviewed, preprints, reviews)
- Study designs (RCTs, observational, in vitro, etc.)
- Document all criteria clearly
### Phase 2: Systematic Literature Search
1. **Multi-Database Search**:
Select databases appropriate for the domain:
**Biomedical & Life Sciences:**
- Use PubMed/PMC for biomedical literature retrieval
- Include bioRxiv/medRxiv preprints when the review scope requires recent non-peer-reviewed evidence
- Use specialized biomedical databases for pathway, structure, variant, and target evidence when relevant
**General Scientific Literature:**
- Search arXiv via direct API (preprints in physics, math, CS, q-bio)
- Search Semantic Scholar via API (200M+ papers, cross-disciplinary)
- Use Google Scholar for comprehensive coverage (manual or careful scraping)
**Specialized Databases:**
- Include protein structure evidence when needed
- Include cancer genomics evidence when needed
- Include demographic, economic, or health statistics when needed
- Use specialized databases as appropriate for the domain
2. **Document Search Parameters**:
```markdown
## Search Strategy
### Database: PubMed
- **Date searched**: 2024-10-25
- **Date range**: 2015-01-01 to 2024-10-25
- **Search string**:
```
("CRISPR"[Title] OR "Cas9"[Title])
AND ("sickle cell"[MeSH] OR "SCD"[Title/Abstract])
AND 2015:2024[Publication Date]
```
- **Results**: 247 articles
```
Repeat for each database searched.
3. **Export and Aggregate Results**:
- Export results in JSON format from each database
- Combine all results into a single file
- Use `scripts/search_databases.py` for post-processing:
```bash
python search_databases.py combined_results.json \
--deduplicate \
--format markdown \
--output aggregated_results.md
```
### Phase 3: Screening and Selection
1. **Deduplication**:
```bash
python search_databases.py results.json --deduplicate --output unique_results.json
```
- Removes duplicates by DOI (primary) or title (fallback)
- Document number of duplicates removed
2. **Title Screening**:
- Review all titles against inclusion/exclusion criteria
- Exclude obviously irrelevant studies
- Document number excluded at this stage
3. **Abstract Screening**:
- Read abstracts of remaining studies
- Apply inclusion/exclusion criteria rigorously
- Document reasons for exclusion
4. **Full-Text Screening**:
- Obtain full texts of remaining studies
- Conduct detailed review against all criteria
- Document specific reasons for exclusion
- Record final number of included studies
5. **Create PRISMA Flow Diagram**:
```
Initial search: n = X
├─ After deduplication: n = Y
├─ After title screening: n = Z
├─ After abstract screening: n = A
└─ Included in review: n = B
```
### Phase 4: Data Extraction and Quality Assessment
1. **Extract Key Data** from each included study:
- Study metadata (authors, year, journal, DOI)
- Study design and methods
- Sample size and population characteristics
- Key findings and results
- Limitations noted by authors
- Funding sources and conflicts of interest
2. **Assess Study Quality**:
- **For RCTs**: Use Cochrane Risk of Bias tool
- **For observational studies**: Use Newcastle-Ottawa Scale
- **For systematic reviews**: Use AMSTAR 2
- Rate each study: High, Moderate, Low, or Very Low quality
- Consider excluding very low-quality studies
3. **Organize by Themes**:
- Identify 3-5 major themes across studies
- Group studies by theme (studies may appear in multiple themes)
- Note patterns, consensus, and controversies
### Phase 5: Synthesis and Analysis
1. **Create Review Document** from template:
```bash
cp assets/review_template.md my_literature_review.md
```
2. **Write Thematic Synthesis** (NOT study-by-study summaries):
- Organize Results section by themes or research questions
- Synthesize findings across multiple studies within each theme
- Compare and contrast different approaches and results
- Identify consensus areas and points of controversy
- Highlight the strongest evidence
Example structure:
```markdown
#### 3.3.1 Theme: CRISPR Delivery Methods
Multiple delivery approaches have been investigated for therapeutic
gene editing. Viral vectors (AAV) were used in 15 studies^1-15^ and
showed high transduction efficiency (65-85%) but raised immunogenicity
concerns^3,7,12^. In contrast, lipid nanoparticles demonstrated lower
efficiency (40-60%) but improved safety profiles^16-23^.
```
3. **Critical Analysis**:
- Evaluate methodological strengths and limitations across studies
- Assess quality and consistency of evidence
- Identify knowledge gaps and methodological gaps
- Note areas requiring future research
4. **Write Discussion**:
- Interpret findings in broader context
- Discuss clinical, practical, or research implications
- Acknowledge limitations of the review itself
- Compare with previous reviews if applicable
- Propose specific future research directions
### Phase 6: Citation Verification
**CRITICAL**: All citations must be verified for accuracy before final submission.
1. **Verify All DOIs**:
```bash
python scripts/verify_citations.py my_literature_review.md
```
This script:
- Extracts all DOIs from the document
- Verifies each DOI resolves correctly
- Retrieves metadata from CrossRef
- Generates verification report
- Outputs properly formatted citations
2. **Review Verification Report**:
- Check for any failed DOIs
- Verify author names, titles, and publication details match
- Correct any errors in the original document
- Re-run verification until all citations pass
3. **Format Citations Consistently**:
- Choose one citation style and use throughout (see `references/citation_styles.md`)
- Common styles: APA, Nature, Vancouver, Chicago, IEEE
- Use verification script output to format citations correctly
- Ensure in-text citations match reference list format
### Phase 7: Document Generation
1. **Generate PDF**:
```bash
python scripts/generate_pdf.py my_literature_review.md \
--citation-style apa \
--output my_review.pdf
```
Options:
- `--citation-style`: apa, nature, chicago, vancouver, ieee
- `--no-toc`: Disable table of contents
- `--no-numbers`: Disable section numbering
- `--check-deps`: Check if pandoc/xelatex are installed
2. **Review Final Output**:
- Check PDF formatting and layout
- Verify all sections are present
- Ensure citations render correctly
- Check that figures/tables appear properly
- Verify table of contents is accurate
3. **Quality Checklist**:
- [ ] All DOIs verified with verify_citations.py
- [ ] Citations formatted consistently
- [ ] PRISMA flow diagram included (for systematic reviews)
- [ ] Search methodology fully documented
- [ ] Inclusion/exclusion criteria clearly stated
- [ ] Results organized thematically (not study-by-study)
- [ ] Quality assessment completed
- [ ] Limitations acknowledged
- [ ] References complete and accurate
- [ ] PDF generates without errors
## Database-Specific Search Guidance
### PubMed / PubMed Central
Access via PubMed, PMC, or Entrez-compatible APIs:
```bash
# Search PubMed
pubmed search "CRISPR gene editing" --limit 100
# Search with filters
# Use PubMed Advanced Search Builder to construct complex queries
# Then execute via PubMed/Entrez-compatible access
```
**Search tips**:
- Use MeSH terms: `"sickle cell disease"[MeSH]`
- Field tags: `[Title]`, `[Title/Abstract]`, `[Author]`
- Date filters: `2020:2024[Publication Date]`
- Boolean operators: AND, OR, NOT
- See MeSH browser: https://meshb.nlm.nih.gov/search
### bioRxiv / medRxiv
Use preprint servers as literature sources inside the review workflow:
```bash
search biorxiv "CRISPR sickle cell" --limit 50
```
**Important considerations**:
- Preprints are not peer-reviewed
- Verify findings with caution
- Check if preprint has been published (CrossRef)
- Note preprint version and date
### arXiv
Access via direct API or WebFetch:
```python
# Example search categories:
# q-bio.QM (Quantitative Methods)
# q-bio.GN (Genomics)
# q-bio.MN (Molecular Networks)
# cs.LG (Machine Learning)
# stat.ML (Machine Learning Statistics)
# Search format: category AND terms
search_query = "cat:q-bio.QM AND ti:\"single cell sequencing\""
```
### Semantic Scholar
Access via direct API (requires API key, or use free tier):
- 200M+ papers across all fields
- Excellent for cross-disciplinary searches
- Provides citation graphs and paper recommendations
- Use for finding highly influential papers
### Specialized Biomedical Databases
Keep literature-review focused on paper search, screening, synthesis, and evidence extraction.
Biomedical database lookup is a separate problem surface; use this section only to record which
external evidence source a paper cites or which database a review protocol should query.
- **ChEMBL**: chemical bioactivity evidence cited by drug-discovery papers
- **UniProt**: protein annotations, accessions, and protein evidence cited by papers
- **KEGG**: pathway and gene annotations cited by biology papers
- **COSMIC**: cancer mutation evidence cited by oncology papers
- **AlphaFold**: predicted protein structures cited by structural-biology papers
- **PDB**: experimental structures cited by structural-biology papers
### Citation Chaining
Expand search via citation networks:
1. **Forward citations** (papers citing key papers):
- Use Google Scholar "Cited by"
- Use Semantic Scholar or OpenAlex APIs
- Identifies newer research building on seminal work
2. **Backward citations** (references from key papers):
- Extract references from included papers
- Identify highly cited foundational work
- Find papers cited by multiple included studies
## Citation Style Guide
Detailed formatting guidelines are in `references/citation_styles.md`. Quick reference:
### APA (7th Edition)
- In-text: (Smith et al., 2023)
- Reference: Smith, J. D., Johnson, M. L., & Williams, K. R. (2023). Title. *Journal*, *22*(4), 301-318. https://doi.org/10.xxx/yyy
### Nature
- In-text: Superscript numbers^1,2^
- Reference: Smith, J. D., Johnson, M. L. & Williams, K. R. Title. *Nat. Rev. Drug Discov.* **22**, 301-318 (2023).
### Vancouver
- In-text: Superscript numbers^1,2^
- Reference: Smith JD, Johnson ML, Williams KR. Title. Nat Rev Drug Discov. 2023;22(4):301-18.
**Always verify citations** with verify_citations.py before finalizing.
### Prioritizing High-Impact Papers (CRITICAL)
**Always prioritize influential, highly-cited papers from reputable authors and top venues.** Quality matters more than quantity in literature reviews.
#### Citation Count Thresholds
Use citation counts to identify the most impactful papers:
| Paper Age | Citation Threshold | Classification |
|-----------|-------------------|----------------|
| 0-3 years | 20+ citations | Noteworthy |
| 0-3 years | 100+ citations | Highly Influential |
| 3-7 years | 100+ citations | Significant |
| 3-7 years | 500+ citations | Landmark Paper |
| 7+ years | 500+ citations | Seminal Work |
| 7+ years | 1000+ citations | Foundational |
#### Journal and Venue Tiers
Prioritize papers from higher-tier venues:
- **Tier 1 (Always Prefer):** Nature, Science, Cell, NEJM, Lancet, JAMA, PNAS, Nature Medicine, Nature Biotechnology
- **Tier 2 (Strong Preference):** High-impact specialized journals (IF>10), top conferences (NeurIPS, ICML for ML/AI)
- **Tier 3 (Include When Relevant):** Respected specialized journals (IF 5-10)
- **Tier 4 (Use Sparingly):** Lower-impact peer-reviewed venues
#### Author Reputation Assessment
Prefer papers from:
- **Senior researchers** with high h-index (>40 in established fields)
- **Leading research groups** at recognized institutions (Harvard, Stanford, MIT, Oxford, etc.)
- **Authors with multiple Tier-1 publications** in the relevant field
- **Researchers with recognized expertise** (awards, editorial positions, society fellows)
#### Identifying Seminal Papers
For any topic, identify foundational work by:
1. **High citation count** (typically 500+ for papers 5+ years old)
2. **Frequently cited by other included studies** (appears in many reference lists)
3. **Published in Tier-1 venues** (Nature, Science, Cell family)
4. **Written by field pioneers** (often cited as establishing concepts)
## Best Practices
### Search Strategy
1. **Use multiple databases** (minimum 3): Ensures comprehensive coverage
2. **Include preprint servers**: Captures latest unpublished findings
3. **Document everything**: Search strings, dates, result counts for reproducibility
4. **Test and refine**: Run pilot searches, review results, adjust search terms
5. **Sort by citations**: When available, sort search results by citation count to surface influential work first
### Screening and Selection
1. **Use multiple databases** (minimum 3): Ensures comprehensive coverage
2. **Include preprint servers**: Captures latest unpublished findings
3. **Document everything**: Search strings, dates, result counts for reproducibility
4. **Test and refine**: Run pilot searches, review results, adjust search terms
### Screening and Selection
1. **Use clear criteria**: Document inclusion/exclusion criteria before screening
2. **Screen systematically**: Title → Abstract → Full text
3. **Document exclusions**: Record reasons for excluding studies
4. **Consider dual screening**: For systematic reviews, have two reviewers screen independently
### Synthesis
1. **Organize thematically**: Group by themes, NOT by individual studies
2. **Synthesize across studies**: Compare, contrast, identify patterns
3. **Be critical**: Evaluate quality and consistency of evidence
4. **Identify gaps**: Note what's missing or understudied
### Quality and Reproducibility
1. **Assess study quality**: Use appropriate quality assessment tools
2. **Verify all citations**: Run verify_citations.py script
3. **Document methodology**: Provide enough detail for others to reproduce
4. **Follow guidelines**: Use PRISMA for systematic reviews
### Writing
1. **Be objective**: Present evidence fairly, acknowledge limitations
2. **Be systematic**: Follow structured template
3. **Be specific**: Include numbers, statistics, effect sizes where available
4. **Be clear**: Use clear headings, logical flow, thematic organization
## Common Pitfalls to Avoid
1. **Single database search**: Misses relevant papers; always search multiple databases
2. **No search documentation**: Makes review irreproducible; document all searches
3. **Study-by-study summary**: Lacks synthesis; organize thematically instead
4. **Unverified citations**: Leads to errors; always run verify_citations.py
5. **Too broad search**: Yields thousands of irrelevant results; refine with specific terms
6. **Too narrow search**: Misses relevant papers; include synonyms and related terms
7. **Ignoring preprints**: Misses latest findings; include bioRxiv, medRxiv, arXiv
8. **No quality assessment**: Treats all evidence equally; assess and report quality
9. **Publication bias**: Only positive results published; note potential bias
10. **Outdated search**: Field evolves rapidly; clearly state search date
## Example Workflow
Complete workflow for a biomedical literature review:
```bash
# 1. Create review document from template
cp assets/review_template.md crispr_sickle_cell_review.md
# 2. Search multiple literature sources
# - Use PubMed for curated biomedical literature
# - Include bioRxiv/medRxiv preprints when recent non-peer-reviewed evidence matters
# - Use direct API access for arXiv, Semantic Scholar
# - Export results in JSON format
# 3. Aggregate and process results
python scripts/search_databases.py combined_results.json \
--deduplicate \
--rank citations \
--year-start 2015 \
--year-end 2024 \
--format markdown \
--output search_results.md \
--summary
# 4. Screen results and extract data
# - Manually screen titles, abstracts, full texts
# - Extract key data into the review document
# - Organize by themes
# 5. Write the review following template structure
# - Introduction with clear objectives
# - Detailed methodology section
# - Results organized thematically
# - Critical discussion
# - Clear conclusions
# 6. Verify all citations
python scripts/verify_citations.py crispr_sickle_cell_review.md
# Review the citation report
cat crispr_sickle_cell_review_citation_report.json
# Fix any failed citations and re-verify
python scripts/verify_citations.py crispr_sickle_cell_review.md
# 7. Generate professional PDF
python scripts/generate_pdf.py crispr_sickle_cell_review.md \
--citation-style nature \
--output crispr_sickle_cell_review.pdf
# 8. Review final PDF and markdown outputs
```
## Resources
### Bundled Resources
**Scripts:**
- `scripts/verify_citations.py`: Verify DOIs and generate formatted citations
- `scripts/generate_pdf.py`: Convert markdown to professional PDF
- `scripts/search_databases.py`: Process, deduplicate, and format search results
**References:**
- `references/citation_styles.md`: Detailed citation formatting guide (APA, Nature, Vancouver, Chicago, IEEE)
- `references/database_strategies.md`: Comprehensive database search strategies
**Assets:**
- `assets/review_template.md`: Complete literature review template with all sections
### External Resources
**Guidelines:**
- PRISMA (Systematic Reviews): http://www.prisma-statement.org/
- Cochrane Handbook: https://training.cochrane.org/handbook
- AMSTAR 2 (Review Quality): https://amstar.ca/
**Tools:**
- MeSH Browser: https://meshb.nlm.nih.gov/search
- PubMed Advanced Search: https://pubmed.ncbi.nlm.nih.gov/advanced/
- Boolean Search Guide: https://www.ncbi.nlm.nih.gov/books/NBK3827/
**Citation Styles:**
- APA Style: https://apastyle.apa.org/
- Nature Portfolio: https://www.nature.com/nature-portfolio/editorial-policies/reporting-standards
- NLM/Vancouver: https://www.nlm.nih.gov/bsd/uniform_requirements.html
## Dependencies
### Required Python Packages
```bash
pip install requests # For citation verification
```
### Required System Tools
```bash
# For PDF generation
brew install pandoc # macOS
apt-get install pandoc # Linux
# For LaTeX (PDF generation)
brew install --cask mactex # macOS
apt-get install texlive-xetex # Linux
```
Check dependencies:
```bash
python scripts/generate_pdf.py --check-deps
```
## Summary
This literature-review skill provides:
1. **Systematic methodology** following academic best practices
2. **Multi-source literature coverage** across domain-appropriate databases
3. **Citation verification** ensuring accuracy and credibility
4. **Professional output** in markdown and PDF formats
5. **Comprehensive guidance** covering the entire review process
6. **Quality assurance** with verification and validation tools
7. **Reproducibility** through detailed documentation requirements
Conduct thorough, rigorous literature reviews that meet academic standards and provide comprehensive synthesis of current knowledge in any domain.
SKILL.md
| name | literature-review |
|---|---|
| description | Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.). This skill should be used when conducting systematic literature reviews, meta-analyses, research synthesis, or comprehensive literature searches across biomedical, scientific, and technical domains. Creates professionally formatted markdown documents and PDFs with verified citations in multiple citation styles (APA, Nature, Vancouver, etc.). |
| allowed-tools | Read Write Edit Bash |
| license | MIT license |
| metadata | { "skill-author": "K-Dense Inc." } |
文献综述
概述
按照严格的学术方法进行系统、全面的文献综述。检索多个文献来源,按主题综合整理发现,验证引用的准确性,并生成专业的 Markdown 和 PDF 输出文档。
何时使用此技能
在以下情况下使用此技能:
- 为研究或发表进行系统文献综述
- 综合特定主题的多个来源的当前知识
- 进行荟萃分析或范围综述
- 撰写研究论文或学位论文的文献综述部分
- 调查某个研究领域的最新技术进展
- 识别研究空白和未来方向
- 需要经过验证的引用和专业格式
核心工作流程
文献综述遵循结构化的多阶段工作流程:
阶段1:规划与范围界定
-
定义研究问题:对于临床/生物医学综述,使用 PICO 框架(Population 人群, Intervention 干预, Comparison 比较, Outcome 结局)
- 示例:“CRISPR-Cas9(I)治疗镰状细胞病(P)与标准治疗(C)相比,效果如何?”
-
确定范围和目标:
- 定义清晰、具体的研究问题
- 确定综述类型(叙述性、系统性、范围综述、荟萃分析)
- 设定边界(时间范围、地理范围、研究类型)
-
制定检索策略:
- 从研究问题中识别 2-4 个主要概念
- 列出每个概念的同义词、缩写和相关术语
- 规划布尔运算符(AND、OR、NOT)以组合术语
- 选择至少 3 个互补数据库
-
设定纳入/排除标准:
- 日期范围(例如,最近 10 年:2015-2024)
- 语言(通常为英语,或指定多语言)
- 出版物类型(同行评审、预印本、综述)
- 研究设计(随机对照试验、观察性研究、体外实验等)
- 清晰地记录所有标准
阶段2:系统文献检索
-
多数据库检索:
根据领域选择合适的数据库:
生物医学与生命科学:
- 使用 PubMed/PMC 检索生物医学文献
- 当综述范围需要最新的非同行评议证据时,包括 bioRxiv/medRxiv 预印本
- 在相关情况下,使用专门的生物医学数据库获取通路、结构、变异和靶点证据
通用科学文献:
- 通过直接 API 搜索 arXiv(物理、数学、计算机科学、定量生物学等领域的预印本)
- 通过 API 搜索 Semantic Scholar(2 亿多篇论文,跨学科)
- 使用 Google Scholar 进行全面覆盖(手动或谨慎抓取)
专业数据库:
- 在需要时包括蛋白质结构证据
- 在需要时包括癌症基因组证据
- 在需要时包括人口、经济或健康统计数据
- 根据领域适当使用专业数据库
-
记录检索参数:
markdown## 检索策略 ### 数据库:PubMed - **检索日期**:2024-10-25 - **日期范围**:2015-01-01 至 2024-10-25 - **检索式**:("CRISPR"[Title] OR "Cas9"[Title]) AND ("sickle cell"[MeSH] OR "SCD"[Title/Abstract]) AND 2015:2024[Publication Date]
text- **结果**:247 篇文献对每个检索的数据库重复以上步骤。
-
导出和汇总结果:
- 从每个数据库以 JSON 格式导出结果
- 将所有结果合并到一个文件中
- 使用
scripts/search_databases.py进行后处理:bashpython search_databases.py combined_results.json \ --deduplicate \ --format markdown \ --output aggregated_results.md
阶段3:筛选与选择
-
去重:
bashpython search_databases.py results.json --deduplicate --output unique_results.json- 通过 DOI(主要)或标题(备用)删除重复项
- 记录删除的重复项数量
-
标题筛选:
- 根据纳入/排除标准检查所有标题
- 排除明显不相关的研究
- 记录此阶段排除的数量
-
摘要筛选:
- 阅读剩余研究的摘要
- 严格应用纳入/排除标准
- 记录排除原因
-
全文筛选:
- 获取剩余研究的全文
- 根据所有标准进行详细审查
- 记录具体的排除原因
- 记录最终纳入的研究数量
-
创建 PRISMA 流程图:
text初始检索:n = X ├─ 去重后:n = Y ├─ 标题筛选后:n = Z ├─ 摘要筛选后:n = A └─ 纳入综述:n = B
阶段4:数据提取与质量评估
-
从每项纳入的研究中提取关键数据:
- 研究元数据(作者、年份、期刊、DOI)
- 研究设计和方法
- 样本量和人群特征
- 关键发现和结果
- 作者指出的局限性
- 资金来源和利益冲突
-
评估研究质量:
- 对于 RCT:使用 Cochrane 偏倚风险工具
- 对于观察性研究:使用 Newcastle-Ottawa 量表
- 对于系统综述:使用 AMSTAR 2
- 对每项研究进行评级:高质量、中等质量、低质量或极低质量
- 考虑排除极低质量的研究
-
按主题组织:
- 识别跨研究的 3-5 个主要主题
- 按主题对研究进行分组(研究可能出现在多个主题中)
- 记录模式、共识和争议
阶段5:综合与分析
-
从模板创建综述文档:
bashcp assets/review_template.md my_literature_review.md -
撰写主题综合(而非逐项总结研究):
- 按主题或研究问题组织结果部分
- 在每个主题内综合多项研究的发现
- 比较和对比不同的方法和结果
- 识别共识领域和争议点
- 突出最有力的证据
示例结构:
markdown#### 3.3.1 主题:CRISPR 递送方法 针对治疗性基因编辑已研究了多种递送方法。病毒载体(AAV)在 15 项研究^1-15^中使用, 并显示出高转导效率(65-85%),但引起了免疫原性问题^3,7,12^。相比之下,脂质纳米颗粒 表现出较低效率(40-60%),但改善了安全性概况^16-23^。 -
批判性分析:
- 评估整个研究的方法学优势和局限性
- 评估证据的质量和一致性
- 识别知识空白和方法学空白
- 指出需要未来研究的领域
-
撰写讨论部分:
- 在更广泛的背景下解释发现
- 讨论临床、实践或研究意义
- 承认综述本身的局限性
- 如果适用,与以前的综述进行比较
- 提出具体的未来研究方向
阶段6:引文验证
关键:在最终提交之前,必须验证所有引文的准确性。
-
验证所有 DOI:
bashpython scripts/verify_citations.py my_literature_review.md此脚本:
- 从文档中提取所有 DOI
- 验证每个 DOI 是否正确解析
- 从 CrossRef 检索元数据
- 生成验证报告
- 输出格式正确的引文
-
查看验证报告:
- 检查任何失败的 DOI
- 验证作者姓名、标题和出版细节是否匹配
- 更正原始文档中的任何错误
- 重新运行验证,直到所有引文通过
-
一致地格式化引文:
- 选择一种引用风格并在全文中使用(参见
references/citation_styles.md) - 常见风格:APA、Nature、Vancouver、Chicago、IEEE
- 使用验证脚本输出正确格式化引文
- 确保正文引用与参考文献列表格式匹配
- 选择一种引用风格并在全文中使用(参见
阶段7:文档生成
-
生成 PDF:
bashpython scripts/generate_pdf.py my_literature_review.md \ --citation-style apa \ --output my_review.pdf选项:
--citation-style:apa、nature、chicago、vancouver、ieee--no-toc:禁用目录--no-numbers:禁用章节编号--check-deps:检查是否安装了 pandoc/xelatex
-
检查最终输出:
- 检查 PDF 格式和布局
- 验证所有部分是否存在
- 确保引文显示正确
- 检查图表是否正确显示
- 验证目录的准确性
-
质量检查清单:
- 所有 DOI 已使用 verify_citations.py 验证
- 引文格式一致
- 包含 PRISMA 流程图(对于系统综述)
- 检索方法已充分记录
- 纳入/排除标准已明确说明
- 结果按主题组织(而非逐项研究)
- 质量评估已完成
- 局限性已承认
- 参考文献完整且准确
- PDF 生成无错误
数据库特定检索指南
PubMed / PubMed Central
通过 PubMed、PMC 或兼容 Entrez 的 API 访问:
# 检索 PubMed
pubmed search "CRISPR gene editing" --limit 100
# 使用过滤器检索
# 使用 PubMed 高级检索构建器构建复杂查询
# 然后通过 PubMed/Entrez 兼容接口执行检索技巧:
- 使用 MeSH 术语:
"sickle cell disease"[MeSH] - 字段标签:
[Title]、[Title/Abstract]、[Author] - 日期过滤器:
2020:2024[Publication Date] - 布尔运算符:AND、OR、NOT
- 参见 MeSH 浏览器:https://meshb.nlm.nih.gov/search
bioRxiv / medRxiv
在综述工作流程中使用预印本服务器作为文献来源:
search biorxiv "CRISPR sickle cell" --limit 50重要注意事项:
- 预印本未经同行评审
- 谨慎验证发现
- 通过 CrossRef 检查预印本是否已发表
- 注明预印本版本和日期
arXiv
通过直接 API 或 WebFetch 访问:
# 示例检索类别:
# q-bio.QM (定量方法)
# q-bio.GN (基因组学)
# q-bio.MN (分子网络)
# cs.LG (机器学习)
# stat.ML (机器学习统计)
# 检索格式:类别 AND 术语
search_query = "cat:q-bio.QM AND ti:\"single cell sequencing\""Semantic Scholar
通过直接 API 访问(需要 API 密钥,或使用免费层):
- 涵盖所有领域的 2 亿多篇论文
- 非常适合跨学科检索
- 提供引用图和论文推荐
- 用于寻找高影响力论文
专门生物医学数据库
将文献综述的重点放在论文检索、筛选、综合和证据提取上。 生物医学数据库查找是一个独立的问题域;本节仅用于记录某篇论文引用了哪个外部证据源, 或综述方案应查询哪个数据库。
- ChEMBL:药物发现论文引用的化学生物活性证据
- UniProt:论文引用的蛋白质注释、登录号和蛋白质证据
- KEGG:生物学论文引用的通路和基因注释
- COSMIC:肿瘤学论文引用的癌症突变证据
- AlphaFold:结构生物学论文引用的预测蛋白质结构
- PDB:结构生物学论文引用的实验结构
引用链
通过引用网络扩展检索:
-
正向引用(引用了关键论文的论文):
- 使用 Google Scholar 的“被引”
- 使用 Semantic Scholar 或 OpenAlex API
- 识别在开创性工作基础上发展的较新研究
-
反向引用(来自关键论文的参考文献):
- 从纳入的论文中提取参考文献
- 识别被高引的基础性工作
- 找到被多个纳入研究引用的论文
引用风格指南
详细格式指南见 references/citation_styles.md。快速参考:
APA(第7版)
- 正文中:(Smith et al., 2023)
- 参考文献:Smith, J. D., Johnson, M. L., & Williams, K. R. (2023). Title. Journal, 22(4), 301-318. https://doi.org/10.xxx/yyy
Nature
- 正文中:上标数字^1,2^
- 参考文献:Smith, J. D., Johnson, M. L. & Williams, K. R. Title. Nat. Rev. Drug Discov. 22, 301-318 (2023).
Vancouver
- 正文中:上标数字^1,2^
- 参考文献:Smith JD, Johnson ML, Williams KR. Title. Nat Rev Drug Discov. 2023;22(4):301-18.
在最终定稿前,始终使用 verify_citations.py 验证引文。
优先考虑高影响力论文(关键)
**始终优先考虑来自知名作者和顶级期刊的有影响力、高被引论文。**质量比数量在文献综述中更重要。
引用次数阈值
使用引用次数来确定最具影响力的论文:
| 论文年龄 | 引用阈值 | 分类 |
|---|---|---|
| 0-3年 | 20+ 引用 | 值得注意 |
| 0-3年 | 100+ 引用 | 高度有影响力 |
| 3-7年 | 100+ 引用 | 显著 |
| 3-7年 | 500+ 引用 | 里程碑论文 |
| 7年以上 | 500+ 引用 | 开创性工作 |
| 7年以上 | 1000+ 引用 | 基础性 |
期刊和会议等级
优先选择来自更高层次期刊的论文:
- 等级1(总是优先): Nature、Science、Cell、NEJM、Lancet、JAMA、PNAS、Nature Medicine、Nature Biotechnology
- 等级2(强烈偏好): 高影响因子专业期刊(IF>10),顶级会议(如 NeurIPS、ICML 适用于 ML/AI)
- 等级3(相关时纳入): 知名专业期刊(IF 5-10)
- 等级4(谨慎使用): 较低影响因子的同行评审期刊
作者声誉评估
优先选择来自以下作者的文章:
- 资深研究人员,h 指数高(在成熟领域>40)
- 知名研究机构的领导研究小组(哈佛、斯坦福、MIT、牛津等)
- 在相关领域有多篇一级期刊发表的作者
- 具有公认专业知识的研究人员(奖项、编辑职位、学会会士)
识别开创性论文
对于任何主题,通过以下方法识别基础性工作:
- 高引用次数(对于 5 年以上论文,通常 500+)
- 被其他纳入研究频繁引用(出现在许多参考文献列表中)
- 发表在一级期刊(Nature、Science、Cell 系列)
- 由领域先驱撰写(经常被引用来确立概念)
最佳实践
检索策略
- 使用多个数据库(至少3个):确保全面覆盖
- 包括预印本服务器:捕获最新的未发表发现
- 记录所有内容:检索式、日期、结果计数,以实现可重复性
- 测试和完善:运行试点检索,查看结果,调整检索词
- 按引用数排序:可用时,按引用数排序检索结果,首先显示有影响力的工作
筛选与选择
- 使用明确的标准:在筛选前记录纳入/排除标准
- 系统筛选:标题 → 摘要 → 全文
- 记录排除:记录排除研究的原因
- 考虑双重筛选:对于系统综述,由两位综述者独立筛选
综合
- 按主题组织:按主题分组,而非按个别研究
- 跨研究综合:比较、对比、识别模式
- 批判性:评估证据的质量和一致性
- 识别空白:指出缺失或研究不足的地方
质量与可重复性
- 评估研究质量:使用适当的质量评估工具
- 验证所有引文:运行 verify_citations.py 脚本
- 记录方法:提供足够的细节供他人重复
- 遵循指南:对系统综述遵循 PRISMA
写作
- 客观:公平呈现证据,承认局限性
- 系统:遵循结构化模板
- 具体:在可获取的情况下,包括数字、统计、效应量
- 清晰:使用清晰的标题、逻辑流程和主题组织
避免的常见陷阱
- 单一数据库检索:会遗漏相关论文;始终检索多个数据库
- 没有检索文档:使综述不可重复;记录所有检索
- 逐项研究总结:缺乏综合;改为按主题组织
- 未经验证的引文:导致错误;始终运行 verify_citations.py
- 过于宽泛的检索:产生数千条不相关的结果;使用特定术语加以完善
- 过于狭窄的检索:遗漏相关论文;包括同义词和相关术语
- 忽略预印本:错过最新发现;包括 bioRxiv、medRxiv、arXiv
- 没有质量评估:同等对待所有证据;评估并报告质量
- 发表偏倚:只有阳性结果发表;注意潜在偏倚
- 过时的检索:领域发展迅速;明确说明检索日期
示例工作流程
生物医学文献综述的完整工作流程:
# 1. 从模板创建综述文档
cp assets/review_template.md crispr_sickle_cell_review.md
# 2. 检索多个文献来源
# - 使用 PubMed 获取精选的生物医学文献
# - 当最近的未同行评议证据重要时,包括 bioRxiv/medRxiv 预印本
# - 使用直接 API 访问 arXiv、Semantic Scholar
# - 以 JSON 格式导出结果
# 3. 汇总和处理结果
python scripts/search_databases.py combined_results.json \
--deduplicate \
--rank citations \
--year-start 2015 \
--year-end 2024 \
--format markdown \
--output search_results.md \
--summary
# 4. 筛选结果并提取数据
# - 手动筛选标题、摘要、全文
# - 将关键数据提取到综述文档中
# - 按主题组织
# 5. 按照模板结构撰写综述
# - 引言包含明确目标
# - 详细的方法学部分
# - 按主题组织的结果
# - 批判性讨论
# - 清晰的结论
# 6. 验证所有引文
python scripts/verify_citations.py crispr_sickle_cell_review.md
# 查看引文报告
cat crispr_sickle_cell_review_citation_report.json
# 修复任何失败的引文并重新验证
python scripts/verify_citations.py crispr_sickle_cell_review.md
# 7. 生成专业 PDF
python scripts/generate_pdf.py crispr_sickle_cell_review.md \
--citation-style nature \
--output crispr_sickle_cell_review.pdf
# 8. 审查最终 PDF 和 Markdown 输出资源
捆绑资源
脚本:
scripts/verify_citations.py:验证 DOI 并生成格式化引文scripts/generate_pdf.py:将 Markdown 转换为专业 PDFscripts/search_databases.py:处理、去重和格式化检索结果
参考:
references/citation_styles.md:详细的引用格式指南(APA、Nature、Vancouver、Chicago、IEEE)references/database_strategies.md:全面的数据库检索策略
素材:
assets/review_template.md:包含所有部分的完整文献综述模板
外部资源
指南:
- PRISMA(系统综述):http://www.prisma-statement.org/
- Cochrane Handbook:https://training.cochrane.org/handbook
- AMSTAR 2(综述质量):https://amstar.ca/
工具:
- MeSH 浏览器:https://meshb.nlm.nih.gov/search
- PubMed 高级检索:https://pubmed.ncbi.nlm.nih.gov/advanced/
- 布尔检索指南:https://www.ncbi.nlm.nih.gov/books/NBK3827/
引用风格:
- APA 风格:https://apastyle.apa.org/
- Nature Portfolio:https://www.nature.com/nature-portfolio/editorial-policies/reporting-standards
- NLM/Vancouver:https://www.nlm.nih.gov/bsd/uniform_requirements.html
依赖项
所需的 Python 包
pip install requests # 用于引文验证所需的系统工具
# 用于 PDF 生成
brew install pandoc # macOS
apt-get install pandoc # Linux
# 用于 LaTeX(PDF 生成)
brew install --cask mactex # macOS
apt-get install texlive-xetex # Linux检查依赖:
python scripts/generate_pdf.py --check-deps总结
此文献综述技能提供:
- 遵循学术最佳实践的系统方法
- 跨领域适当数据库的多源文献覆盖
- 确保准确性和可信度的引文验证
- 以 Markdown 和 PDF 格式的专业输出
- 覆盖整个综述过程的全面指导
- 带有验证和校验工具的质量保证
- 通过详细文档要求实现的可重复性
在任何领域开展符合学术标准的全面、严谨的文献综述,并提供对当前知识的全面综合。