科研技能库/系统性文献综述
文献检索
未发现用户侧风险

系统性文献综述

使用多个学术数据库(PubMed、arXiv、bioRxiv、Semantic Scholar 等)进行全面、系统的文献综述。此技能应在进行系统性文献综述、荟萃分析、研究综合或涵盖生物医学、科学和技术领域的全面文献检索时使用。可创建具有多种引文格式(APA、Nature、Vancouver 等)的专业格式 Markdown 文档和 PDF。

文件预览

9 个文件
assets
references
scripts
SKILL.md
27.4 KB · 可预览
---
name: literature-review
description: Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.). This skill should be used when conducting systematic literature reviews, meta-analyses, research synthesis, or comprehensive literature searches across biomedical, scientific, and technical domains. Creates professionally formatted markdown documents and PDFs with verified citations in multiple citation styles (APA, Nature, Vancouver, etc.).
allowed-tools: Read Write Edit Bash
license: MIT license
metadata:
  version: "1.0"
  skill-author: K-Dense Inc.
---

# Literature Review

## Overview

Conduct systematic, comprehensive literature reviews following rigorous academic methodology. Search multiple literature databases, synthesize findings thematically, verify all citations for accuracy, and generate professional output documents in markdown and PDF formats.

This skill uses the **parallel-web skill** (`parallel-cli search`) as the primary web search tool for broad academic literature discovery, supplemented by specialized database access skills (gget, bioservices, datacommons-client). It provides specialized tools for citation verification, result aggregation, and document generation.

## When to Use This Skill

Use this skill when:
- Conducting a systematic literature review for research or publication
- Synthesizing current knowledge on a specific topic across multiple sources
- Performing meta-analysis or scoping reviews
- Writing the literature review section of a research paper or thesis
- Investigating the state of the art in a research domain
- Identifying research gaps and future directions
- Requiring verified citations and professional formatting

## Visual Enhancement with Scientific Schematics

**⚠️ MANDATORY: Every literature review MUST include at least 1-2 AI-generated figures using the scientific-schematics skill.**

This is not optional. Literature reviews without visual elements are incomplete. Before finalizing any document:
1. Generate at minimum ONE schematic or diagram (e.g., PRISMA flow diagram for systematic reviews)
2. Prefer 2-3 figures for comprehensive reviews (search strategy flowchart, thematic synthesis diagram, conceptual framework)

**How to generate figures:**
- Use the **scientific-schematics** skill to generate AI-powered publication-quality diagrams
- Simply describe your desired diagram in natural language
- Nano Banana Pro will automatically generate, review, and refine the schematic

**How to generate schematics:**
```bash
python scripts/generate_schematic.py "your diagram description" -o figures/output.png
```

The AI will automatically:
- Create publication-quality images with proper formatting
- Review and refine through multiple iterations
- Ensure accessibility (colorblind-friendly, high contrast)
- Save outputs in the figures/ directory

**When to add schematics:**
- PRISMA flow diagrams for systematic reviews
- Literature search strategy flowcharts
- Thematic synthesis diagrams
- Research gap visualization maps
- Citation network diagrams
- Conceptual framework illustrations
- Any complex concept that benefits from visualization

For detailed guidance on creating schematics, refer to the scientific-schematics skill documentation.

---

## Core Workflow

Literature reviews follow a structured, multi-phase workflow:

### Phase 1: Planning and Scoping

1. **Define Research Question**: Use PICO framework (Population, Intervention, Comparison, Outcome) for clinical/biomedical reviews
   - Example: "What is the efficacy of CRISPR-Cas9 (I) for treating sickle cell disease (P) compared to standard care (C)?"

2. **Establish Scope and Objectives**:
   - Define clear, specific research questions
   - Determine review type (narrative, systematic, scoping, meta-analysis)
   - Set boundaries (time period, geographic scope, study types)

3. **Develop Search Strategy**:
   - Identify 2-4 main concepts from research question
   - List synonyms, abbreviations, and related terms for each concept
   - Plan Boolean operators (AND, OR, NOT) to combine terms
   - Select minimum 3 complementary databases
   - **Use the parallel-web skill (`parallel-cli search`) for initial scoping** to quickly gauge the landscape before formal database searches

4. **Set Inclusion/Exclusion Criteria**:
   - Date range (e.g., last 10 years: 2015-2024)
   - Language (typically English, or specify multilingual)
   - Publication types (peer-reviewed, preprints, reviews)
   - Study designs (RCTs, observational, in vitro, etc.)
   - Document all criteria clearly

### Phase 2: Systematic Literature Search

1. **Multi-Database Search**:

   Select databases appropriate for the domain. **Always start with parallel-web for broad academic coverage**, then supplement with domain-specific databases.

   **Web-Based Academic Search (parallel-web skill — START HERE):**
   - Use `parallel-cli search` with academic domain filtering for broad scholarly coverage
   - Run two searches: academic-focused + general to catch all relevant sources
   ```bash
   # Academic-focused search across scholarly sources
   parallel-cli search "your research topic" -q "keyword1" -q "keyword2" \
     --json --max-results 10 --excerpt-max-chars-total 27000 \
     --include-domains "scholar.google.com,arxiv.org,pubmed.ncbi.nlm.nih.gov,semanticscholar.org,biorxiv.org,medrxiv.org,ncbi.nlm.nih.gov,nature.com,science.org,ieee.org,acm.org,springer.com,wiley.com,cell.com,pnas.org,nih.gov" \
     -o sources/litreview_<topic>-academic.json

   # General search for supplementary sources
   parallel-cli search "your research topic" -q "keyword1" -q "keyword2" \
     --json --max-results 10 --excerpt-max-chars-total 27000 \
     -o sources/litreview_<topic>-general.json
   ```
   - Use `parallel-cli extract` to fetch full content from specific paper URLs or PDFs found in search results
   ```bash
   parallel-cli extract "https://arxiv.org/abs/XXXX.XXXXX" --json
   ```

   **Biomedical & Life Sciences:**
   - Use `gget` skill: `gget search pubmed "search terms"` for PubMed/PMC
   - Use `gget` skill: `gget search biorxiv "search terms"` for preprints
   - Use `bioservices` skill for ChEMBL, KEGG, UniProt, etc.

   **General Scientific Literature:**
   - Search arXiv via direct API (preprints in physics, math, CS, q-bio)
   - Search Semantic Scholar via API (200M+ papers, cross-disciplinary)
   - Use Google Scholar for comprehensive coverage (manual or careful scraping)

   **Specialized Databases:**
   - Use `gget alphafold` for protein structures
   - Use `gget cosmic` for cancer genomics
   - Use `datacommons-client` for demographic/statistical data
   - Use specialized databases as appropriate for the domain

2. **Document Search Parameters**:
   ```markdown
   ## Search Strategy

   ### Database: PubMed
   - **Date searched**: 2024-10-25
   - **Date range**: 2015-01-01 to 2024-10-25
   - **Search string**:
     ```
     ("CRISPR"[Title] OR "Cas9"[Title])
     AND ("sickle cell"[MeSH] OR "SCD"[Title/Abstract])
     AND 2015:2024[Publication Date]
     ```
   - **Results**: 247 articles
   ```

   Repeat for each database searched.

3. **Export and Aggregate Results**:
   - Export results in JSON format from each database
   - Combine all results into a single file
   - Use `scripts/search_databases.py` for post-processing:
     ```bash
     python search_databases.py combined_results.json \
       --deduplicate \
       --format markdown \
       --output aggregated_results.md
     ```

### Phase 3: Screening and Selection

1. **Deduplication**:
   ```bash
   python search_databases.py results.json --deduplicate --output unique_results.json
   ```
   - Removes duplicates by DOI (primary) or title (fallback)
   - Document number of duplicates removed

2. **Title Screening**:
   - Review all titles against inclusion/exclusion criteria
   - Exclude obviously irrelevant studies
   - Document number excluded at this stage

3. **Abstract Screening**:
   - Read abstracts of remaining studies
   - Apply inclusion/exclusion criteria rigorously
   - Document reasons for exclusion

4. **Full-Text Screening**:
   - Obtain full texts of remaining studies
   - Conduct detailed review against all criteria
   - Document specific reasons for exclusion
   - Record final number of included studies

5. **Create PRISMA Flow Diagram**:
   ```
   Initial search: n = X
   ├─ After deduplication: n = Y
   ├─ After title screening: n = Z
   ├─ After abstract screening: n = A
   └─ Included in review: n = B
   ```

### Phase 4: Data Extraction and Quality Assessment

1. **Extract Key Data** from each included study:
   - Study metadata (authors, year, journal, DOI)
   - Study design and methods
   - Sample size and population characteristics
   - Key findings and results
   - Limitations noted by authors
   - Funding sources and conflicts of interest

2. **Assess Study Quality**:
   - **For RCTs**: Use Cochrane Risk of Bias tool
   - **For observational studies**: Use Newcastle-Ottawa Scale
   - **For systematic reviews**: Use AMSTAR 2
   - Rate each study: High, Moderate, Low, or Very Low quality
   - Consider excluding very low-quality studies

3. **Organize by Themes**:
   - Identify 3-5 major themes across studies
   - Group studies by theme (studies may appear in multiple themes)
   - Note patterns, consensus, and controversies

### Phase 5: Synthesis and Analysis

1. **Create Review Document** from template:
   ```bash
   cp assets/review_template.md my_literature_review.md
   ```

2. **Write Thematic Synthesis** (NOT study-by-study summaries):
   - Organize Results section by themes or research questions
   - Synthesize findings across multiple studies within each theme
   - Compare and contrast different approaches and results
   - Identify consensus areas and points of controversy
   - Highlight the strongest evidence

   Example structure:
   ```markdown
   #### 3.3.1 Theme: CRISPR Delivery Methods

   Multiple delivery approaches have been investigated for therapeutic
   gene editing. Viral vectors (AAV) were used in 15 studies^1-15^ and
   showed high transduction efficiency (65-85%) but raised immunogenicity
   concerns^3,7,12^. In contrast, lipid nanoparticles demonstrated lower
   efficiency (40-60%) but improved safety profiles^16-23^.
   ```

3. **Critical Analysis**:
   - Evaluate methodological strengths and limitations across studies
   - Assess quality and consistency of evidence
   - Identify knowledge gaps and methodological gaps
   - Note areas requiring future research

4. **Write Discussion**:
   - Interpret findings in broader context
   - Discuss clinical, practical, or research implications
   - Acknowledge limitations of the review itself
   - Compare with previous reviews if applicable
   - Propose specific future research directions

### Phase 6: Citation Verification

**CRITICAL**: All citations must be verified for accuracy before final submission.

1. **Verify All DOIs**:
   ```bash
   python scripts/verify_citations.py my_literature_review.md
   ```

   This script:
   - Extracts all DOIs from the document
   - Verifies each DOI resolves correctly
   - Retrieves metadata from CrossRef
   - Generates verification report
   - Outputs properly formatted citations

2. **Review Verification Report**:
   - Check for any failed DOIs
   - Verify author names, titles, and publication details match
   - Correct any errors in the original document
   - Re-run verification until all citations pass

3. **Format Citations Consistently**:
   - Choose one citation style and use throughout (see `references/citation_styles.md`)
   - Common styles: APA, Nature, Vancouver, Chicago, IEEE
   - Use verification script output to format citations correctly
   - Ensure in-text citations match reference list format

### Phase 7: Document Generation

1. **Generate PDF**:
   ```bash
   python scripts/generate_pdf.py my_literature_review.md \
     --citation-style apa \
     --output my_review.pdf
   ```

   Options:
   - `--citation-style`: apa, nature, chicago, vancouver, ieee
   - `--no-toc`: Disable table of contents
   - `--no-numbers`: Disable section numbering
   - `--check-deps`: Check if pandoc/xelatex are installed

2. **Review Final Output**:
   - Check PDF formatting and layout
   - Verify all sections are present
   - Ensure citations render correctly
   - Check that figures/tables appear properly
   - Verify table of contents is accurate

3. **Quality Checklist**:
   - [ ] All DOIs verified with verify_citations.py
   - [ ] Citations formatted consistently
   - [ ] PRISMA flow diagram included (for systematic reviews)
   - [ ] Search methodology fully documented
   - [ ] Inclusion/exclusion criteria clearly stated
   - [ ] Results organized thematically (not study-by-study)
   - [ ] Quality assessment completed
   - [ ] Limitations acknowledged
   - [ ] References complete and accurate
   - [ ] PDF generates without errors

## Database-Specific Search Guidance

### PubMed / PubMed Central

Access via `gget` skill:
```bash
# Search PubMed
gget search pubmed "CRISPR gene editing" -l 100

# Search with filters
# Use PubMed Advanced Search Builder to construct complex queries
# Then execute via gget or direct Entrez API
```

**Search tips**:
- Use MeSH terms: `"sickle cell disease"[MeSH]`
- Field tags: `[Title]`, `[Title/Abstract]`, `[Author]`
- Date filters: `2020:2024[Publication Date]`
- Boolean operators: AND, OR, NOT
- See MeSH browser: https://meshb.nlm.nih.gov/search

### bioRxiv / medRxiv

Access via `gget` skill:
```bash
gget search biorxiv "CRISPR sickle cell" -l 50
```

**Important considerations**:
- Preprints are not peer-reviewed
- Verify findings with caution
- Check if preprint has been published (CrossRef)
- Note preprint version and date

### arXiv

Access via direct API or WebFetch:
```python
# Example search categories:
# q-bio.QM (Quantitative Methods)
# q-bio.GN (Genomics)
# q-bio.MN (Molecular Networks)
# cs.LG (Machine Learning)
# stat.ML (Machine Learning Statistics)

# Search format: category AND terms
search_query = "cat:q-bio.QM AND ti:\"single cell sequencing\""
```

### Semantic Scholar

Access via direct API (requires API key, or use free tier):
- 200M+ papers across all fields
- Excellent for cross-disciplinary searches
- Provides citation graphs and paper recommendations
- Use for finding highly influential papers

### Specialized Biomedical Databases

Use appropriate skills:
- **ChEMBL**: `bioservices` skill for chemical bioactivity
- **UniProt**: `gget` or `bioservices` skill for protein information
- **KEGG**: `bioservices` skill for pathways and genes
- **COSMIC**: `gget` skill for cancer mutations
- **AlphaFold**: `gget alphafold` for protein structures
- **PDB**: `gget` or direct API for experimental structures

### Citation Chaining

Expand search via citation networks:

1. **Forward citations** (papers citing key papers):
   - Use `parallel-cli search` to find papers citing a specific work:
     ```bash
     parallel-cli search "papers citing [Author et al. Year] [paper title]" \
       -q "citing" -q "[key author]" \
       --json --max-results 10 --excerpt-max-chars-total 27000 \
       --include-domains "scholar.google.com,semanticscholar.org,arxiv.org,pubmed.ncbi.nlm.nih.gov" \
       -o sources/litreview_forward_citations.json
     ```
   - Use Google Scholar "Cited by"
   - Use Semantic Scholar or OpenAlex APIs
   - Identifies newer research building on seminal work

2. **Backward citations** (references from key papers):
   - Use `parallel-cli extract` to fetch full text of key papers and extract their reference lists:
     ```bash
     parallel-cli extract "https://doi.org/10.xxxx/yyyy" --json
     ```
   - Extract references from included papers
   - Identify highly cited foundational work
   - Find papers cited by multiple included studies

## Citation Style Guide

Detailed formatting guidelines are in `references/citation_styles.md`. Quick reference:

### APA (7th Edition)
- In-text: (Smith et al., 2023)
- Reference: Smith, J. D., Johnson, M. L., & Williams, K. R. (2023). Title. *Journal*, *22*(4), 301-318. https://doi.org/10.xxx/yyy

### Nature
- In-text: Superscript numbers^1,2^
- Reference: Smith, J. D., Johnson, M. L. & Williams, K. R. Title. *Nat. Rev. Drug Discov.* **22**, 301-318 (2023).

### Vancouver
- In-text: Superscript numbers^1,2^
- Reference: Smith JD, Johnson ML, Williams KR. Title. Nat Rev Drug Discov. 2023;22(4):301-18.

**Always verify citations** with verify_citations.py before finalizing.

### Prioritizing High-Impact Papers (CRITICAL)

**Always prioritize influential, highly-cited papers from reputable authors and top venues.** Quality matters more than quantity in literature reviews.

#### Citation Count Thresholds

Use citation counts to identify the most impactful papers:

| Paper Age | Citation Threshold | Classification |
|-----------|-------------------|----------------|
| 0-3 years | 20+ citations | Noteworthy |
| 0-3 years | 100+ citations | Highly Influential |
| 3-7 years | 100+ citations | Significant |
| 3-7 years | 500+ citations | Landmark Paper |
| 7+ years | 500+ citations | Seminal Work |
| 7+ years | 1000+ citations | Foundational |

#### Journal and Venue Tiers

Prioritize papers from higher-tier venues:

- **Tier 1 (Always Prefer):** Nature, Science, Cell, NEJM, Lancet, JAMA, PNAS, Nature Medicine, Nature Biotechnology
- **Tier 2 (Strong Preference):** High-impact specialized journals (IF>10), top conferences (NeurIPS, ICML for ML/AI)
- **Tier 3 (Include When Relevant):** Respected specialized journals (IF 5-10)
- **Tier 4 (Use Sparingly):** Lower-impact peer-reviewed venues

#### Author Reputation Assessment

Prefer papers from:
- **Senior researchers** with high h-index (>40 in established fields)
- **Leading research groups** at recognized institutions (Harvard, Stanford, MIT, Oxford, etc.)
- **Authors with multiple Tier-1 publications** in the relevant field
- **Researchers with recognized expertise** (awards, editorial positions, society fellows)

#### Identifying Seminal Papers

For any topic, identify foundational work by:
1. **High citation count** (typically 500+ for papers 5+ years old)
2. **Frequently cited by other included studies** (appears in many reference lists)
3. **Published in Tier-1 venues** (Nature, Science, Cell family)
4. **Written by field pioneers** (often cited as establishing concepts)

## Best Practices

### Search Strategy
1. **Start with parallel-web**: Use `parallel-cli search` with academic domains for initial broad coverage before querying specialized databases
2. **Use multiple databases** (minimum 3): Ensures comprehensive coverage — parallel-web counts as one source
3. **Include preprint servers**: Captures latest unpublished findings
4. **Document everything**: Search strings, dates, result counts for reproducibility — save all parallel-cli output to `sources/`
5. **Test and refine**: Run pilot searches, review results, adjust search terms
6. **Sort by citations**: When available, sort search results by citation count to surface influential work first
7. **Use parallel-cli extract**: Fetch full content from promising URLs found during search to verify relevance before full-text screening

### Screening and Selection
1. **Use multiple databases** (minimum 3): Ensures comprehensive coverage
2. **Include preprint servers**: Captures latest unpublished findings
3. **Document everything**: Search strings, dates, result counts for reproducibility
4. **Test and refine**: Run pilot searches, review results, adjust search terms

### Screening and Selection
1. **Use clear criteria**: Document inclusion/exclusion criteria before screening
2. **Screen systematically**: Title → Abstract → Full text
3. **Document exclusions**: Record reasons for excluding studies
4. **Consider dual screening**: For systematic reviews, have two reviewers screen independently

### Synthesis
1. **Organize thematically**: Group by themes, NOT by individual studies
2. **Synthesize across studies**: Compare, contrast, identify patterns
3. **Be critical**: Evaluate quality and consistency of evidence
4. **Identify gaps**: Note what's missing or understudied

### Quality and Reproducibility
1. **Assess study quality**: Use appropriate quality assessment tools
2. **Verify all citations**: Run verify_citations.py script
3. **Document methodology**: Provide enough detail for others to reproduce
4. **Follow guidelines**: Use PRISMA for systematic reviews

### Writing
1. **Be objective**: Present evidence fairly, acknowledge limitations
2. **Be systematic**: Follow structured template
3. **Be specific**: Include numbers, statistics, effect sizes where available
4. **Be clear**: Use clear headings, logical flow, thematic organization

## Common Pitfalls to Avoid

1. **Single database search**: Misses relevant papers; always search multiple databases
2. **No search documentation**: Makes review irreproducible; document all searches
3. **Study-by-study summary**: Lacks synthesis; organize thematically instead
4. **Unverified citations**: Leads to errors; always run verify_citations.py
5. **Too broad search**: Yields thousands of irrelevant results; refine with specific terms
6. **Too narrow search**: Misses relevant papers; include synonyms and related terms
7. **Ignoring preprints**: Misses latest findings; include bioRxiv, medRxiv, arXiv
8. **No quality assessment**: Treats all evidence equally; assess and report quality
9. **Publication bias**: Only positive results published; note potential bias
10. **Outdated search**: Field evolves rapidly; clearly state search date

## Example Workflow

Complete workflow for a biomedical literature review:

```bash
# 1. Create review document from template
cp assets/review_template.md crispr_sickle_cell_review.md

# 2. Start with parallel-web for broad academic search
parallel-cli search "CRISPR Cas9 sickle cell disease gene therapy efficacy" \
  -q "CRISPR" -q "sickle cell" -q "gene therapy" \
  --json --max-results 10 --excerpt-max-chars-total 27000 \
  --include-domains "scholar.google.com,arxiv.org,pubmed.ncbi.nlm.nih.gov,semanticscholar.org,biorxiv.org,nature.com,science.org,cell.com,pnas.org,nih.gov" \
  -o sources/litreview_crispr_scd-academic.json

parallel-cli search "CRISPR sickle cell disease clinical trials treatment" \
  -q "CRISPR" -q "sickle cell" \
  --json --max-results 10 --excerpt-max-chars-total 27000 \
  -o sources/litreview_crispr_scd-general.json

# 3. Search specialized databases using appropriate skills
# - Use gget skill for PubMed, bioRxiv
# - Use direct API access for arXiv, Semantic Scholar
# - Export results in JSON format

# 4. Aggregate and process results (combine parallel-cli + database results)
python scripts/search_databases.py combined_results.json \
  --deduplicate \
  --rank citations \
  --year-start 2015 \
  --year-end 2024 \
  --format markdown \
  --output search_results.md \
  --summary

# 5. Screen results and extract data
# - Use parallel-cli extract to fetch full content from promising URLs
# - Manually screen titles, abstracts, full texts
# - Extract key data into the review document
# - Organize by themes

# 6. Write the review following template structure
# - Introduction with clear objectives
# - Detailed methodology section
# - Results organized thematically
# - Critical discussion
# - Clear conclusions

# 7. Verify all citations
python scripts/verify_citations.py crispr_sickle_cell_review.md

# Review the citation report
cat crispr_sickle_cell_review_citation_report.json

# Fix any failed citations and re-verify
python scripts/verify_citations.py crispr_sickle_cell_review.md

# 8. Generate professional PDF
python scripts/generate_pdf.py crispr_sickle_cell_review.md \
  --citation-style nature \
  --output crispr_sickle_cell_review.pdf

# 9. Review final PDF and markdown outputs
```

## Integration with Other Skills

This skill works seamlessly with other scientific skills:

### Web Search & Extraction (parallel-web skill — PRIMARY)
- **parallel-cli search**: Broad academic and general web search with domain filtering — use for initial scoping, finding papers, citation chaining, and supplementary searches
- **parallel-cli extract**: Fetch full content from paper URLs, journal websites, and preprint servers — use for reading abstracts, extracting reference lists, and verifying paper details
- **parallel-cli search --include-domains**: Academic-focused search across scholarly domains (arxiv.org, pubmed, nature.com, etc.)

### Database Access Skills
- **gget**: PubMed, bioRxiv, COSMIC, AlphaFold, Ensembl, UniProt
- **bioservices**: ChEMBL, KEGG, Reactome, UniProt, PubChem
- **datacommons-client**: Demographics, economics, health statistics

### Analysis Skills
- **pydeseq2**: RNA-seq differential expression (for methods sections)
- **scanpy**: Single-cell analysis (for methods sections)
- **anndata**: Single-cell data (for methods sections)
- **biopython**: Sequence analysis (for background sections)

### Visualization Skills
- **matplotlib**: Generate figures and plots for review
- **seaborn**: Statistical visualizations

### Writing Skills
- **brand-guidelines**: Apply institutional branding to PDF
- **internal-comms**: Adapt review for different audiences

## Resources

### Bundled Resources

**Scripts:**
- `scripts/verify_citations.py`: Verify DOIs and generate formatted citations
- `scripts/generate_pdf.py`: Convert markdown to professional PDF
- `scripts/search_databases.py`: Process, deduplicate, and format search results

**References:**
- `references/citation_styles.md`: Detailed citation formatting guide (APA, Nature, Vancouver, Chicago, IEEE)
- `references/database_strategies.md`: Comprehensive database search strategies

**Assets:**
- `assets/review_template.md`: Complete literature review template with all sections

### External Resources

**Guidelines:**
- PRISMA (Systematic Reviews): http://www.prisma-statement.org/
- Cochrane Handbook: https://training.cochrane.org/handbook
- AMSTAR 2 (Review Quality): https://amstar.ca/

**Tools:**
- MeSH Browser: https://meshb.nlm.nih.gov/search
- PubMed Advanced Search: https://pubmed.ncbi.nlm.nih.gov/advanced/
- Boolean Search Guide: https://www.ncbi.nlm.nih.gov/books/NBK3827/

**Citation Styles:**
- APA Style: https://apastyle.apa.org/
- Nature Portfolio: https://www.nature.com/nature-portfolio/editorial-policies/reporting-standards
- NLM/Vancouver: https://www.nlm.nih.gov/bsd/uniform_requirements.html

## Dependencies

### Required CLI Tools
```bash
# parallel-cli (PRIMARY — for web search and URL extraction)
curl -fsSL https://parallel.ai/install.sh | bash
# Or: uv tool install "parallel-web-tools[cli]"
# Authenticate: parallel-cli auth
```

### Required Python Packages
```bash
pip install requests  # For citation verification
```

### Required System Tools
```bash
# For PDF generation
brew install pandoc  # macOS
apt-get install pandoc  # Linux

# For LaTeX (PDF generation)
brew install --cask mactex  # macOS
apt-get install texlive-xetex  # Linux
```

Check dependencies:
```bash
python scripts/generate_pdf.py --check-deps
```

## Summary

This literature-review skill provides:

1. **Systematic methodology** following academic best practices
2. **Parallel-web powered search** using `parallel-cli search` for fast, broad academic literature discovery with scholarly domain filtering
3. **Multi-database integration** via existing scientific skills (gget, bioservices, datacommons-client)
4. **Citation verification** ensuring accuracy and credibility
5. **Professional output** in markdown and PDF formats
6. **Comprehensive guidance** covering the entire review process
7. **Quality assurance** with verification and validation tools
8. **Reproducibility** through detailed documentation requirements

Conduct thorough, rigorous literature reviews that meet academic standards and provide comprehensive synthesis of current knowledge in any domain.

SKILL.md

元数据
nameliterature-review
description使用多个学术数据库(PubMed、arXiv、bioRxiv、Semantic Scholar 等)进行全面、系统的文献综述。此技能应在进行系统性文献综述、荟萃分析、研究综合或涵盖生物医学、科学和技术领域的全面文献检索时使用。可创建具有多种引文格式(APA、Nature、Vancouver 等)的专业格式 Markdown 文档和 PDF。
allowed-toolsRead Write Edit Bash
licenseMIT license
metadata{ "version": "1.0", "skill-author": "K-Dense Inc." }

文献综述

概述

按照严谨的学术方法进行系统性、全面的文献综述。搜索多个文献数据库,按主题综合发现,验证所有引文的准确性,并生成专业的 Markdown 和 PDF 输出文档。

此技能使用 parallel-web 技能(parallel-cli search)作为主要的 Web 搜索工具,进行广泛的学术文献发现,并辅以专门的数据库访问技能(gget、bioservices、datacommons-client)。它提供专用的引文验证、结果汇总和文档生成工具。

何时使用此技能

以下情况使用此技能:

  • 进行研究或发表系统性文献综述
  • 综合多个来源的特定主题当前知识
  • 进行荟萃分析或范围界定综述
  • 撰写研究论文或学位论文的文献综述部分
  • 调查某一研究领域的现状
  • 识别研究空白和未来方向
  • 需要经过验证的引文和专业格式排版

使用科学示意图进行视觉增强

⚠️ 强制性要求:每篇文献综述必须包含至少 1-2 张使用 scientific-schematics 技能生成的 AI 图示。

这不是可选项。没有视觉元素的文献综述是不完整的。在最终确定任何文档之前:

  1. 至少生成一张示意图或图表(例如,系统性综述的 PRISMA 流程图)
  2. 对于全面的综述,最好包含 2-3 张图(检索策略流程图、主题综合图、概念框架)

如何生成图示:

  • 使用 scientific-schematics 技能生成 AI 驱动的、适合发表的图表
  • 只需用自然语言描述您想要的图表
  • Nano Banana Pro 将自动生成、审查并优化示意图

生成示意图的命令:

bash
python scripts/generate_schematic.py "your diagram description" -o figures/output.png

AI 将自动:

  • 创建格式正确的、适合发表的图像
  • 通过多次迭代进行审查和优化
  • 确保可访问性(色盲友好、高对比度)
  • 将输出保存在 figures/ 目录中

何时添加示意图:

  • 系统性综述的 PRISMA 流程图
  • 文献检索策略流程图
  • 主题综合图
  • 研究空白可视化地图
  • 引文网络图
  • 概念框架图示
  • 任何受益于可视化的复杂概念

有关创建示意图的详细指导,请参阅 scientific-schematics 技能文档。


核心工作流程

文献综述遵循结构化的多阶段工作流程:

第一阶段:规划和范围界定

  1. 定义研究问题:对于临床/生物医学综述,使用 PICO 框架(人群、干预、比较、结局)

    • 示例:“CRISPR-Cas9(I)用于治疗镰状细胞病(P)与标准治疗(C)相比的疗效如何?”
  2. 确定范围和目标:

    • 定义明确、具体的研究问题
    • 确定综述类型(叙事性、系统性、范围界定、荟萃分析)
    • 设定边界(时间段、地理位置范围、研究类型)
  3. 制定检索策略:

    • 从研究问题中识别 2-4 个主要概念
    • 为每个概念列出同义词、缩写和相关术语
    • 计划布尔运算符(AND、OR、NOT)来组合术语
    • 至少选择 3 个互补的数据库
    • 在正式数据库检索之前,使用 parallel-web 技能(parallel-cli search)进行初步范围界定,以快速了解整体情况
  4. 设定纳入/排除标准:

    • 日期范围(例如,过去 10 年:2015-2024)
    • 语言(通常为英文,或指定多语种)
    • 出版物类型(同行评审、预印本、综述)
    • 研究设计(随机对照试验、观察性研究、体外研究等)
    • 清晰地记录所有标准

第二阶段:系统性文献检索

  1. 多数据库检索:

    选择适合该领域的数据库。始终从 parallel-web 开始进行广泛的学术覆盖,然后辅以特定领域数据库。

    基于网络的学术检索(parallel-web 技能——从此处开始):

    • 使用 parallel-cli search 结合学术领域过滤,进行广泛的学术资源覆盖
    • 执行两次搜索:一次学术聚焦搜索 + 一次通用搜索,以捕获所有相关来源
    bash
    # 跨学术资源的学术聚焦搜索
    parallel-cli search "your research topic" -q "keyword1" -q "keyword2" \
      --json --max-results 10 --excerpt-max-chars-total 27000 \
      --include-domains "scholar.google.com,arxiv.org,pubmed.ncbi.nlm.nih.gov,semanticscholar.org,biorxiv.org,medrxiv.org,ncbi.nlm.nih.gov,nature.com,science.org,ieee.org,acm.org,springer.com,wiley.com,cell.com,pnas.org,nih.gov" \
      -o sources/litreview_<topic>-academic.json
    
    # 获取补充资源的通用搜索
    parallel-cli search "your research topic" -q "keyword1" -q "keyword2" \
      --json --max-results 10 --excerpt-max-chars-total 27000 \
      -o sources/litreview_<topic>-general.json
    • 使用 parallel-cli extract 从搜索结果中找到的特定论文 URL 或 PDF 中获取完整内容
    bash
    parallel-cli extract "https://arxiv.org/abs/XXXX.XXXXX" --json

    生物医学与生命科学:

    • 使用 gget 技能:gget search pubmed "search terms" 用于 PubMed/PMC
    • 使用 gget 技能:gget search biorxiv "search terms" 用于预印本
    • 使用 bioservices 技能用于 ChEMBL、KEGG、UniProt 等

    通用科学文献:

    • 通过直接 API 检索 arXiv(物理、数学、计算机科学、定量生物学的预印本)
    • 通过 API 检索 Semantic Scholar(超过 2 亿篇论文,跨学科)
    • 使用 Google Scholar 进行全面覆盖(手动或谨慎抓取)

    专业数据库:

    • 使用 gget alphafold 获取蛋白质结构
    • 使用 gget cosmic 获取癌症基因组学
    • 使用 datacommons-client 获取人口/统计数据
    • 根据领域使用其他专业数据库
  2. 记录检索参数:

    markdown
    ## 检索策略
    
    ### 数据库:PubMed
    - **检索日期**:2024-10-25
    - **日期范围**:2015-01-01 至 2024-10-25
    - **检索字符串**:

    ("CRISPR"[Title] OR "Cas9"[Title]) AND ("sickle cell"[MeSH] OR "SCD"[Title/Abstract]) AND 2015:2024[Publication Date]

    text
    - **结果**:247 篇文章

    对每个数据库重复上述操作。

  3. 导出和汇总结果:

    • 从每个数据库导出 JSON 格式的结果
    • 将所有结果合并到一个文件中
    • 使用 scripts/search_databases.py 进行后处理:
      bash
      python search_databases.py combined_results.json \
        --deduplicate \
        --format markdown \
        --output aggregated_results.md

第三阶段:筛选与选择

  1. 去重:

    bash
    python search_databases.py results.json --deduplicate --output unique_results.json
    • 根据 DOI(首选)或标题(后备)去除重复记录
    • 记录删除的重复数量
  2. 标题筛选:

    • 对照纳入/排除标准审查所有标题
    • 排除明显不相关的研究
    • 记录此阶段排除的数量
  3. 摘要筛选:

    • 阅读剩余研究的摘要
    • 严格应用纳入/排除标准
    • 记录排除理由
  4. 全文筛选:

    • 获取剩余研究的全文
    • 对照所有标准进行详细审查
    • 记录具体的排除理由
    • 记录最终纳入的研究数量
  5. 创建 PRISMA 流程图:

    text
    初步检索:n = X
    ├─ 去重后:n = Y
    ├─ 标题筛选后:n = Z
    ├─ 摘要筛选后:n = A
    └─ 纳入综述:n = B

第四阶段:数据提取与质量评估

  1. 从每项纳入的研究中提取关键数据:

    • 研究元数据(作者、年份、期刊、DOI)
    • 研究设计与方法
    • 样本量和人群特征
    • 主要发现和结果
    • 作者指出的局限性
    • 资金来源和利益冲突
  2. 评估研究质量:

    • 对于随机对照试验:使用 Cochrane 偏倚风险工具
    • 对于观察性研究:使用纽卡斯尔-渥太华量表
    • 对于系统性综述:使用 AMSTAR 2
    • 为每项研究评级:高、中、低或极低质量
    • 考虑排除极低质量的研究
  3. 按主题组织:

    • 识别跨研究的 3-5 个主要主题
    • 按主题对研究进行分组(研究可能出现在多个主题中)
    • 注意模式、共识和争议

第五阶段:综合与分析

  1. 从模板创建综述文档:

    bash
    cp assets/review_template.md my_literature_review.md
  2. 撰写主题综合(而非逐篇研究总结):

    • 按主题或研究问题组织结果部分
    • 在每个主题内综合多篇研究的发现
    • 比较和对比不同的方法和结果
    • 确定共识领域和争议点
    • 突出最强有力的证据

    示例结构:

    markdown
    #### 3.3.1 主题:CRISPR 递送方法
    
    多种递送方法已被研究用于治疗性基因编辑。病毒载体(AAV)在 15 项研究^1-15^ 中被使用,显示高转导效率(65-85%),但引发了免疫原性担忧^3,7,12^。相比之下,脂质纳米颗粒效率较低(40-60%),但安全性更优^16-23^。
  3. 批判性分析:

    • 评价跨研究的方法学优势和局限性
    • 评估证据的质量和一致性
    • 识别知识空白和方法学空白
    • 指出需要未来研究的领域
  4. 撰写讨论:

    • 在更广泛的背景下解释发现
    • 讨论临床、实践或研究意义
    • 承认综述本身的局限性
    • 如适用,与以往的综述进行比较
    • 提出具体的未来研究方向

第六阶段:引文验证

关键:所有引文在最终提交前必须验证准确性。

  1. 验证所有 DOI:

    bash
    python scripts/verify_citations.py my_literature_review.md

    此脚本:

    • 从文档中提取所有 DOI
    • 验证每个 DOI 是否正确解析
    • 从 CrossRef 检索元数据
    • 生成验证报告
    • 输出格式正确的引文
  2. 审查验证报告:

    • 检查是否有任何失败的 DOI
    • 验证作者姓名、标题和出版详细信息是否匹配
    • 纠正原始文档中的任何错误
    • 重新运行验证直到所有引文通过
  3. 统一引文格式:

    • 选择一种引文样式并在全文中使用(参见 references/citation_styles.md)
    • 常见样式:APA、Nature、Vancouver、Chicago、IEEE
    • 使用验证脚本的输出正确格式化引文
    • 确保文中引文与参考文献列表格式匹配

第七阶段:文档生成

  1. 生成 PDF:

    bash
    python scripts/generate_pdf.py my_literature_review.md \
      --citation-style apa \
      --output my_review.pdf

    选项:

    • --citation-style:apa、nature、chicago、vancouver、ieee
    • --no-toc:禁用目录
    • --no-numbers:禁用章节编号
    • --check-deps:检查 pandoc/xelatex 是否已安装
  2. 审查最终输出:

    • 检查 PDF 格式和布局
    • 验证所有部分是否存在
    • 确保引文正确渲染
    • 检查图形/表格是否正确显示
    • 验证目录准确
  3. 质量检查清单:

    • 所有 DOI 已通过 verify_citations.py 验证
    • 引文格式统一
    • 包含 PRISMA 流程图(针对系统性综述)
    • 检索方法详细记录
    • 纳入/排除标准明确陈述
    • 结果按主题组织(而非逐篇研究)
    • 完成质量评估
    • 局限性已承认
    • 参考文献完整且准确
    • PDF 生成无错误

特定数据库检索指导

PubMed / PubMed Central

通过 gget 技能访问:

bash
# 搜索 PubMed
gget search pubmed "CRISPR gene editing" -l 100

# 使用过滤器搜索
# 使用 PubMed 高级搜索构建器构建复杂查询
# 然后通过 gget 或直接 Entrez API 执行

检索提示:

  • 使用 MeSH 术语:"sickle cell disease"[MeSH]
  • 字段标签:[Title]、[Title/Abstract]、[Author]
  • 日期过滤器:2020:2024[Publication Date]
  • 布尔运算符:AND、OR、NOT
  • 参见 MeSH 浏览器:https://meshb.nlm.nih.gov/search

bioRxiv / medRxiv

通过 gget 技能访问:

bash
gget search biorxiv "CRISPR sickle cell" -l 50

重要注意事项:

  • 预印本未经同行评审
  • 谨慎验证发现
  • 检查预印本是否已发表(CrossRef)
  • 注意预印本版本和日期

arXiv

通过直接 API 或 WebFetch 访问:

python
# 示例搜索类别:
# q-bio.QM(定量方法)
# q-bio.GN(基因组学)
# q-bio.MN(分子网络)
# cs.LG(机器学习)
# stat.ML(机器学习统计学)

# 搜索格式:类别 AND 术语
search_query = "cat:q-bio.QM AND ti:\"single cell sequencing\""

Semantic Scholar

通过直接 API 访问(需要 API 密钥,或使用免费层):

  • 超过 2 亿篇论文,涵盖所有领域
  • 非常适合跨学科搜索
  • 提供引文图和论文推荐
  • 用于查找高影响力论文

专业生物医学数据库

使用适当的技能:

  • ChEMBL:bioservices 技能用于化学生物活性
  • UniProt:gget 或 bioservices 技能用于蛋白质信息
  • KEGG:bioservices 技能用于通路和基因
  • COSMIC:gget 技能用于癌症突变
  • AlphaFold:gget alphafold 用于蛋白质结构
  • PDB:gget 或直接 API 用于实验结构

引文链

通过引文网络扩展搜索:

  1. 前向引文(引用关键论文的论文):

    • 使用 parallel-cli search 查找引用特定工作的论文:
      bash
      parallel-cli search "papers citing [Author et al. Year] [paper title]" \
        -q "citing" -q "[key author]" \
        --json --max-results 10 --excerpt-max-chars-total 27000 \
        --include-domains "scholar.google.com,semanticscholar.org,arxiv.org,pubmed.ncbi.nlm.nih.gov" \
        -o sources/litreview_forward_citations.json
    • 使用 Google Scholar 的“引用”
    • 使用 Semantic Scholar 或 OpenAlex API
    • 识别在开创性工作基础上发展的较新研究
  2. 后向引文(关键论文的参考文献):

    • 使用 parallel-cli extract 获取关键论文全文并提取其参考文献列表:
      bash
      parallel-cli extract "https://doi.org/10.xxxx/yyyy" --json
    • 从纳入的论文中提取参考文献
    • 识别高引用量的基础性工作
    • 查找被多篇纳入研究引用的论文

引文样式指南

详细格式指南见 references/citation_styles.md。快速参考:

APA(第 7 版)

  • 文中:(Smith et al., 2023)
  • 参考文献:Smith, J. D., Johnson, M. L., & Williams, K. R. (2023). Title. Journal, 22(4), 301-318. https://doi.org/10.xxx/yyy

Nature

  • 文中:上标数字^1,2^
  • 参考文献:Smith, J. D., Johnson, M. L. & Williams, K. R. Title. Nat. Rev. Drug Discov. 22, 301-318 (2023).

Vancouver

  • 文中:上标数字^1,2^
  • 参考文献:Smith JD, Johnson ML, Williams KR. Title. Nat Rev Drug Discov. 2023;22(4):301-18.

在最终确定前,始终使用 verify_citations.py 验证引文。

优先考虑高影响力论文(关键)

始终优先考虑来自知名作者和顶级期刊的、有影响力的高引用论文。 在文献综述中,质量比数量更重要。

引用次数阈值

使用引用次数来识别最有影响力的论文:

论文年龄引用阈值分类
0-3 年20+ 引用值得关注
0-3 年100+ 引用高影响力
3-7 年100+ 引用重要
3-7 年500+ 引用里程碑式论文
7+ 年500+ 引用开创性工作
7+ 年1000+ 引用基础性工作

期刊和会议级别

优先选择更高级别期刊的论文:

  • Tier 1(始终优先): Nature、Science、Cell、NEJM、Lancet、JAMA、PNAS、Nature Medicine、Nature Biotechnology
  • Tier 2(强烈偏好): 高影响力专业期刊(IF>10)、顶级会议(机器学习/AI 的 NeurIPS、ICML)
  • Tier 3(相关时包含): 受尊敬的专业期刊(IF 5-10)
  • Tier 4(谨慎使用): 较低影响力的同行评审渠道

作者声誉评估

优先选择以下作者的论文:

  • 资深研究人员 具有高 h 指数(在成熟领域 >40)
  • 领先研究团队 来自知名机构(哈佛大学、斯坦福大学、麻省理工学院、牛津大学等)
  • 在相关领域有多篇 Tier-1 发表的作者
  • 具有公认专业知识的作者(奖项、编辑职务、学会会士)

识别开创性论文

对于任何主题,按以下方式识别基础性工作:

  1. 高引用量(通常 5 年以上论文达到 500+)
  2. 被其他纳入研究频繁引用(出现在许多参考文献列表中)
  3. 发表在 Tier-1 期刊(Nature、Science、Cell 系列)
  4. 由领域先驱撰写(常被引用为概念建立者)

最佳实践

检索策略

  1. 从 parallel-web 开始:在查询专业数据库之前,使用 parallel-cli search 结合学术域名进行初步广泛覆盖
  2. 使用多个数据库(至少 3 个):确保全面覆盖——parallel-web 算作一个来源
  3. 包含预印本服务器:捕获最新的未发表发现
  4. 记录所有内容:检索字符串、日期、结果数量以确保可重复性——将所有 parallel-cli 输出保存到 sources/ 目录
  5. 测试并优化:进行试点检索,审查结果,调整检索词
  6. 按引用量排序:当可用时,按引用量对检索结果排序,优先呈现有影响力的工作
  7. 使用 parallel-cli extract:从搜索中发现的潜在 URL 获取完整内容,以在全文筛选前验证相关性

筛选与选择

  1. 使用明确的标准:在筛选前记录纳入/排除标准
  2. 系统筛选:标题 → 摘要 → 全文
  3. 记录排除项:记录排除研究的理由
  4. 考虑双重筛选:对于系统性综述,由两名评审员独立筛选

综合

  1. 按主题组织:按主题分组,而非逐篇研究
  2. 跨研究综合:比较、对比、识别模式
  3. 批判性评估:评价证据的质量和一致性
  4. 识别空白:指出缺失或研究不足的领域

质量与可重复性

  1. 评估研究质量:使用适当的评估工具
  2. 验证所有引文:运行 verify_citations.py 脚本
  3. 记录方法:提供足够详细的信息供他人重现
  4. 遵循指南:系统性综述使用 PRISMA

写作

  1. 客观:公平呈现证据,承认局限性
  2. 系统:遵循结构化模板
  3. 具体:包含数字、统计数据、效应量(如有)
  4. 清晰:使用清晰的标题、逻辑流程、主题组织

常见陷阱

  1. 单一数据库搜索:遗漏相关论文;始终搜索多个数据库
  2. 无搜索文档:导致综述不可复现;记录所有搜索
  3. 逐篇研究总结:缺乏综合;应改为按主题组织
  4. 未验证引文:导致错误;始终运行 verify_citations.py
  5. 搜索范围过宽:产生大量不相关结果;使用具体术语精炼
  6. 搜索范围过窄:遗漏相关论文;包含同义词和相关术语
  7. 忽略预印本:错过最新发现;包含 bioRxiv、medRxiv、arXiv
  8. 无质量评估:平等对待所有证据;评估并报告质量
  9. 发表偏倚:只有阳性结果被发表;注意潜在偏倚
  10. 过时的检索:领域发展迅速;明确说明检索日期

示例工作流程

生物医学文献综述的完整工作流程:

bash
# 1. 从模板创建综述文档
cp assets/review_template.md crispr_sickle_cell_review.md

# 2. 首先使用 parallel-web 进行广泛的学术搜索
parallel-cli search "CRISPR Cas9 sickle cell disease gene therapy efficacy" \
  -q "CRISPR" -q "sickle cell" -q "gene therapy" \
  --json --max-results 10 --excerpt-max-chars-total 27000 \
  --include-domains "scholar.google.com,arxiv.org,pubmed.ncbi.nlm.nih.gov,semanticscholar.org,biorxiv.org,nature.com,science.org,cell.com,pnas.org,nih.gov" \
  -o sources/litreview_crispr_scd-academic.json

parallel-cli search "CRISPR sickle cell disease clinical trials treatment" \
  -q "CRISPR" -q "sickle cell" \
  --json --max-results 10 --excerpt-max-chars-total 27000 \
  -o sources/litreview_crispr_scd-general.json

# 3. 使用适当的技能搜索专业数据库
# - 使用 gget 技能搜索 PubMed、bioRxiv
# - 使用直接 API 访问 arXiv、Semantic Scholar
# - 以 JSON 格式导出结果

# 4. 汇总和处理结果(合并 parallel-cli + 数据库结果)
python scripts/search_databases.py combined_results.json \
  --deduplicate \
  --rank citations \
  --year-start 2015 \
  --year-end 2024 \
  --format markdown \
  --output search_results.md \
  --summary

# 5. 筛选结果并提取数据
# - 使用 parallel-cli extract 从潜在 URL 获取完整内容
# - 手动筛选标题、摘要、全文
# - 将关键数据提取到综述文档中
# - 按主题组织

# 6. 按照模板结构撰写综述
# - 明确目标的引言
# - 详细的方法学部分
# - 按主题组织的结果
# - 批判性讨论
# - 清晰的结论

# 7. 验证所有引文
python scripts/verify_citations.py crispr_sickle_cell_review.md

# 审查引文报告
cat crispr_sickle_cell_review_citation_report.json

# 修正任何失败的引文并重新验证
python scripts/verify_citations.py crispr_sickle_cell_review.md

# 8. 生成专业 PDF
python scripts/generate_pdf.py crispr_sickle_cell_review.md \
  --citation-style nature \
  --output crispr_sickle_cell_review.pdf

# 9. 审查最终 PDF 和 markdown 输出

与其他技能集成

此技能可与...配合使用。

(剩余部分因原始文档截断而省略)