科研技能库/文献引用管理
知识管理
未发现用户侧风险

文献引用管理

全面的学术研究引用管理。搜索 Google Scholar 和 PubMed 论文,提取准确元数据,验证引用,生成格式正确的 BibTeX 条目。当需要查找论文、验证引用信息、将 DOI 转换为 BibTeX 或确保科学写作中引用准确性时使用此技能。

文件预览

16 个文件
assets
references
scripts
SKILL.md
31.8 KB · 可预览
---
name: citation-management
description: Comprehensive citation management for academic research. Search Google Scholar and PubMed for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries. This skill should be used when you need to find papers, verify citation information, convert DOIs to BibTeX, or ensure reference accuracy in scientific writing.
allowed-tools: Read Write Edit Bash
license: MIT License
metadata:
  version: "1.0"
  skill-author: K-Dense Inc.
---

# Citation Management

## Overview

Manage citations systematically throughout the research and writing process. This skill provides tools and strategies for searching academic databases (Google Scholar, PubMed), extracting accurate metadata from multiple sources (CrossRef, PubMed, arXiv), validating citation information, and generating properly formatted BibTeX entries.

Critical for maintaining citation accuracy, avoiding reference errors, and ensuring reproducible research. Integrates seamlessly with the literature-review skill for comprehensive research workflows.

## When to Use This Skill

Use this skill when:
- Searching for specific papers on Google Scholar or PubMed
- Converting DOIs, PMIDs, or arXiv IDs to properly formatted BibTeX
- Extracting complete metadata for citations (authors, title, journal, year, etc.)
- Validating existing citations for accuracy
- Cleaning and formatting BibTeX files
- Finding highly cited papers in a specific field
- Verifying that citation information matches the actual publication
- Building a bibliography for a manuscript or thesis
- Checking for duplicate citations
- Ensuring consistent citation formatting

## Visual Enhancement with Scientific Schematics

**When creating documents with this skill, always consider adding scientific diagrams and schematics to enhance visual communication.**

If your document does not already contain schematics or diagrams:
- Use the **scientific-schematics** skill to generate AI-powered publication-quality diagrams
- Simply describe your desired diagram in natural language
- Nano Banana Pro will automatically generate, review, and refine the schematic

**For new documents:** Scientific schematics should be generated by default to visually represent key concepts, workflows, architectures, or relationships described in the text.

**How to generate schematics:**
```bash
python scripts/generate_schematic.py "your diagram description" -o figures/output.png
```

The AI will automatically:
- Create publication-quality images with proper formatting
- Review and refine through multiple iterations
- Ensure accessibility (colorblind-friendly, high contrast)
- Save outputs in the figures/ directory

**When to add schematics:**
- Citation workflow diagrams
- Literature search methodology flowcharts
- Reference management system architectures
- Citation style decision trees
- Database integration diagrams
- Any complex concept that benefits from visualization

For detailed guidance on creating schematics, refer to the scientific-schematics skill documentation.

---

## Core Workflow

Citation management follows a systematic process:

### Phase 1: Paper Discovery and Search

**Goal**: Find relevant papers using academic search engines.

#### Google Scholar Search

Google Scholar provides the most comprehensive coverage across disciplines.

**Basic Search**:
```bash
# Search for papers on a topic
python scripts/search_google_scholar.py "CRISPR gene editing" \
  --limit 50 \
  --output results.json

# Search with year filter
python scripts/search_google_scholar.py "machine learning protein folding" \
  --year-start 2020 \
  --year-end 2024 \
  --limit 100 \
  --output ml_proteins.json
```

**Advanced Search Strategies** (see `references/google_scholar_search.md`):
- Use quotation marks for exact phrases: `"deep learning"`
- Search by author: `author:LeCun`
- Search in title: `intitle:"neural networks"`
- Exclude terms: `machine learning -survey`
- Find highly cited papers using sort options
- Filter by date ranges to get recent work

**Best Practices**:
- Use specific, targeted search terms
- Include key technical terms and acronyms
- Filter by recent years for fast-moving fields
- Check "Cited by" to find seminal papers
- Export top results for further analysis

#### PubMed Search

PubMed specializes in biomedical and life sciences literature (35+ million citations).

**Basic Search**:
```bash
# Search PubMed
python scripts/search_pubmed.py "Alzheimer's disease treatment" \
  --limit 100 \
  --output alzheimers.json

# Search with MeSH terms and filters
python scripts/search_pubmed.py \
  --query '"Alzheimer Disease"[MeSH] AND "Drug Therapy"[MeSH]' \
  --date-start 2020 \
  --date-end 2024 \
  --publication-types "Clinical Trial,Review" \
  --output alzheimers_trials.json
```

**Advanced PubMed Queries** (see `references/pubmed_search.md`):
- Use MeSH terms: `"Diabetes Mellitus"[MeSH]`
- Field tags: `"cancer"[Title]`, `"Smith J"[Author]`
- Boolean operators: `AND`, `OR`, `NOT`
- Date filters: `2020:2024[Publication Date]`
- Publication types: `"Review"[Publication Type]`
- Combine with E-utilities API for automation

**Best Practices**:
- Use MeSH Browser to find correct controlled vocabulary
- Construct complex queries in PubMed Advanced Search Builder first
- Include multiple synonyms with OR
- Retrieve PMIDs for easy metadata extraction
- Export to JSON or directly to BibTeX

### Phase 2: Metadata Extraction

**Goal**: Convert paper identifiers (DOI, PMID, arXiv ID) to complete, accurate metadata.

#### Quick DOI to BibTeX Conversion

For single DOIs, use the quick conversion tool:

```bash
# Convert single DOI
python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2

# Convert multiple DOIs from a file
python scripts/doi_to_bibtex.py --input dois.txt --output references.bib

# Different output formats
python scripts/doi_to_bibtex.py 10.1038/nature12345 --format json
```

#### Comprehensive Metadata Extraction

For DOIs, PMIDs, arXiv IDs, or URLs:

```bash
# Extract from DOI
python scripts/extract_metadata.py --doi 10.1038/s41586-021-03819-2

# Extract from PMID
python scripts/extract_metadata.py --pmid 34265844

# Extract from arXiv ID
python scripts/extract_metadata.py --arxiv 2103.14030

# Extract from URL
python scripts/extract_metadata.py --url "https://www.nature.com/articles/s41586-021-03819-2"

# Batch extraction from file (mixed identifiers)
python scripts/extract_metadata.py --input identifiers.txt --output citations.bib
```

**Metadata Sources** (see `references/metadata_extraction.md`):

1. **CrossRef API**: Primary source for DOIs
   - Comprehensive metadata for journal articles
   - Publisher-provided information
   - Includes authors, title, journal, volume, pages, dates
   - Free, no API key required

2. **PubMed E-utilities**: Biomedical literature
   - Official NCBI metadata
   - Includes MeSH terms, abstracts
   - PMID and PMCID identifiers
   - Free, API key recommended for high volume

3. **arXiv API**: Preprints in physics, math, CS, q-bio
   - Complete metadata for preprints
   - Version tracking
   - Author affiliations
   - Free, open access

4. **DataCite API**: Research datasets, software, other resources
   - Metadata for non-traditional scholarly outputs
   - DOIs for datasets and code
   - Free access

**What Gets Extracted**:
- **Required fields**: author, title, year
- **Journal articles**: journal, volume, number, pages, DOI
- **Books**: publisher, ISBN, edition
- **Conference papers**: booktitle, conference location, pages
- **Preprints**: repository (arXiv, bioRxiv), preprint ID
- **Additional**: abstract, keywords, URL

### Phase 3: BibTeX Formatting

**Goal**: Generate clean, properly formatted BibTeX entries.

#### Understanding BibTeX Entry Types

See `references/bibtex_formatting.md` for complete guide.

**Common Entry Types**:
- `@article`: Journal articles (most common)
- `@book`: Books
- `@inproceedings`: Conference papers
- `@incollection`: Book chapters
- `@phdthesis`: Dissertations
- `@misc`: Preprints, software, datasets

**Required Fields by Type**:

```bibtex
@article{citationkey,
  author  = {Last1, First1 and Last2, First2},
  title   = {Article Title},
  journal = {Journal Name},
  year    = {2024},
  volume  = {10},
  number  = {3},
  pages   = {123--145},
  doi     = {10.1234/example}
}

@inproceedings{citationkey,
  author    = {Last, First},
  title     = {Paper Title},
  booktitle = {Conference Name},
  year      = {2024},
  pages     = {1--10}
}

@book{citationkey,
  author    = {Last, First},
  title     = {Book Title},
  publisher = {Publisher Name},
  year      = {2024}
}
```

#### Formatting and Cleaning

Use the formatter to standardize BibTeX files:

```bash
# Format and clean BibTeX file
python scripts/format_bibtex.py references.bib \
  --output formatted_references.bib

# Sort entries by citation key
python scripts/format_bibtex.py references.bib \
  --sort key \
  --output sorted_references.bib

# Sort by year (newest first)
python scripts/format_bibtex.py references.bib \
  --sort year \
  --descending \
  --output sorted_references.bib

# Remove duplicates
python scripts/format_bibtex.py references.bib \
  --deduplicate \
  --output clean_references.bib

# Validate and report issues
python scripts/format_bibtex.py references.bib \
  --validate \
  --report validation_report.txt
```

**Formatting Operations**:
- Standardize field order
- Consistent indentation and spacing
- Proper capitalization in titles (protected with {})
- Standardized author name format
- Consistent citation key format
- Remove unnecessary fields
- Fix common errors (missing commas, braces)

### Phase 4: Citation Validation

**Goal**: Verify all citations are accurate and complete.

#### Comprehensive Validation

```bash
# Validate BibTeX file
python scripts/validate_citations.py references.bib

# Validate and fix common issues
python scripts/validate_citations.py references.bib \
  --auto-fix \
  --output validated_references.bib

# Generate detailed validation report
python scripts/validate_citations.py references.bib \
  --report validation_report.json \
  --verbose
```

**Validation Checks** (see `references/citation_validation.md`):

1. **DOI Verification**:
   - DOI resolves correctly via doi.org
   - Metadata matches between BibTeX and CrossRef
   - No broken or invalid DOIs

2. **Required Fields**:
   - All required fields present for entry type
   - No empty or missing critical information
   - Author names properly formatted

3. **Data Consistency**:
   - Year is valid (4 digits, reasonable range)
   - Volume/number are numeric
   - Pages formatted correctly (e.g., 123--145)
   - URLs are accessible

4. **Duplicate Detection**:
   - Same DOI used multiple times
   - Similar titles (possible duplicates)
   - Same author/year/title combinations

5. **Format Compliance**:
   - Valid BibTeX syntax
   - Proper bracing and quoting
   - Citation keys are unique
   - Special characters handled correctly

**Validation Output**:
```json
{
  "total_entries": 150,
  "valid_entries": 145,
  "errors": [
    {
      "citation_key": "Smith2023",
      "error_type": "missing_field",
      "field": "journal",
      "severity": "high"
    },
    {
      "citation_key": "Jones2022",
      "error_type": "invalid_doi",
      "doi": "10.1234/broken",
      "severity": "high"
    }
  ],
  "warnings": [
    {
      "citation_key": "Brown2021",
      "warning_type": "possible_duplicate",
      "duplicate_of": "Brown2021a",
      "severity": "medium"
    }
  ]
}
```

### Phase 5: Integration with Writing Workflow

#### Building References for Manuscripts

Complete workflow for creating a bibliography:

```bash
# 1. Search for papers on your topic
python scripts/search_pubmed.py \
  '"CRISPR-Cas Systems"[MeSH] AND "Gene Editing"[MeSH]' \
  --date-start 2020 \
  --limit 200 \
  --output crispr_papers.json

# 2. Extract DOIs from search results and convert to BibTeX
python scripts/extract_metadata.py \
  --input crispr_papers.json \
  --output crispr_refs.bib

# 3. Add specific papers by DOI
python scripts/doi_to_bibtex.py 10.1038/nature12345 >> crispr_refs.bib
python scripts/doi_to_bibtex.py 10.1126/science.abcd1234 >> crispr_refs.bib

# 4. Format and clean the BibTeX file
python scripts/format_bibtex.py crispr_refs.bib \
  --deduplicate \
  --sort year \
  --descending \
  --output references.bib

# 5. Validate all citations
python scripts/validate_citations.py references.bib \
  --auto-fix \
  --report validation.json \
  --output final_references.bib

# 6. Review validation report and fix any remaining issues
cat validation.json

# 7. Use in your LaTeX document
# \bibliography{final_references}
```

#### Integration with Literature Review Skill

This skill complements the `literature-review` skill:

**Literature Review Skill** → Systematic search and synthesis
**Citation Management Skill** → Technical citation handling

**Combined Workflow**:
1. Use `literature-review` for comprehensive multi-database search
2. Use `citation-management` to extract and validate all citations
3. Use `literature-review` to synthesize findings thematically
4. Use `citation-management` to verify final bibliography accuracy

```bash
# After completing literature review
# Verify all citations in the review document
python scripts/validate_citations.py my_review_references.bib --report review_validation.json

# Format for specific citation style if needed
python scripts/format_bibtex.py my_review_references.bib \
  --style nature \
  --output formatted_refs.bib
```

## Search Strategies

### Google Scholar Best Practices

**Finding Seminal and High-Impact Papers** (CRITICAL):

Always prioritize papers based on citation count, venue quality, and author reputation:

**Citation Count Thresholds:**
| Paper Age | Citations | Classification |
|-----------|-----------|----------------|
| 0-3 years | 20+ | Noteworthy |
| 0-3 years | 100+ | Highly Influential |
| 3-7 years | 100+ | Significant |
| 3-7 years | 500+ | Landmark Paper |
| 7+ years | 500+ | Seminal Work |
| 7+ years | 1000+ | Foundational |

**Venue Quality Tiers:**
- **Tier 1 (Prefer):** Nature, Science, Cell, NEJM, Lancet, JAMA, PNAS
- **Tier 2 (High Priority):** Impact Factor >10, top conferences (NeurIPS, ICML, ICLR)
- **Tier 3 (Good):** Specialized journals (IF 5-10)
- **Tier 4 (Sparingly):** Lower-impact peer-reviewed venues

**Author Reputation Indicators:**
- Senior researchers with h-index >40
- Multiple publications in Tier-1 venues
- Leadership at recognized institutions
- Awards and editorial positions

**Search Strategies for High-Impact Papers:**
- Sort by citation count (most cited first)
- Look for review articles from Tier-1 journals for overview
- Check "Cited by" for impact assessment and recent follow-up work
- Use citation alerts for tracking new citations to key papers
- Filter by top venues using `source:Nature` or `source:Science`
- Search for papers by known field leaders using `author:LastName`

**Advanced Operators** (full list in `references/google_scholar_search.md`):
```
"exact phrase"           # Exact phrase matching
author:lastname          # Search by author
intitle:keyword          # Search in title only
source:journal           # Search specific journal
-exclude                 # Exclude terms
OR                       # Alternative terms
2020..2024              # Year range
```

**Example Searches**:
```
# Find recent reviews on a topic
"CRISPR" intitle:review 2023..2024

# Find papers by specific author on topic
author:Church "synthetic biology"

# Find highly cited foundational work
"deep learning" 2012..2015 sort:citations

# Exclude surveys and focus on methods
"protein folding" -survey -review intitle:method
```

### PubMed Best Practices

**Using MeSH Terms**:
MeSH (Medical Subject Headings) provides controlled vocabulary for precise searching.

1. **Find MeSH terms** at https://meshb.nlm.nih.gov/search
2. **Use in queries**: `"Diabetes Mellitus, Type 2"[MeSH]`
3. **Combine with keywords** for comprehensive coverage

**Field Tags**:
```
[Title]              # Search in title only
[Title/Abstract]     # Search in title or abstract
[Author]             # Search by author name
[Journal]            # Search specific journal
[Publication Date]   # Date range
[Publication Type]   # Article type
[MeSH]              # MeSH term
```

**Building Complex Queries**:
```bash
# Clinical trials on diabetes treatment published recently
"Diabetes Mellitus, Type 2"[MeSH] AND "Drug Therapy"[MeSH] 
AND "Clinical Trial"[Publication Type] AND 2020:2024[Publication Date]

# Reviews on CRISPR in specific journal
"CRISPR-Cas Systems"[MeSH] AND "Nature"[Journal] AND "Review"[Publication Type]

# Specific author's recent work
"Smith AB"[Author] AND cancer[Title/Abstract] AND 2022:2024[Publication Date]
```

**E-utilities for Automation**:
The scripts use NCBI E-utilities API for programmatic access:
- **ESearch**: Search and retrieve PMIDs
- **EFetch**: Retrieve full metadata
- **ESummary**: Get summary information
- **ELink**: Find related articles

See `references/pubmed_search.md` for complete API documentation.

## Tools and Scripts

### search_google_scholar.py

Search Google Scholar and export results.

**Features**:
- Automated searching with rate limiting
- Pagination support
- Year range filtering
- Export to JSON or BibTeX
- Citation count information

**Usage**:
```bash
# Basic search
python scripts/search_google_scholar.py "quantum computing"

# Advanced search with filters
python scripts/search_google_scholar.py "quantum computing" \
  --year-start 2020 \
  --year-end 2024 \
  --limit 100 \
  --sort-by citations \
  --output quantum_papers.json

# Export directly to BibTeX
python scripts/search_google_scholar.py "machine learning" \
  --limit 50 \
  --format bibtex \
  --output ml_papers.bib
```

### search_pubmed.py

Search PubMed using E-utilities API.

**Features**:
- Complex query support (MeSH, field tags, Boolean)
- Date range filtering
- Publication type filtering
- Batch retrieval with metadata
- Export to JSON or BibTeX

**Usage**:
```bash
# Simple keyword search
python scripts/search_pubmed.py "CRISPR gene editing"

# Complex query with filters
python scripts/search_pubmed.py \
  --query '"CRISPR-Cas Systems"[MeSH] AND "therapeutic"[Title/Abstract]' \
  --date-start 2020-01-01 \
  --date-end 2024-12-31 \
  --publication-types "Clinical Trial,Review" \
  --limit 200 \
  --output crispr_therapeutic.json

# Export to BibTeX
python scripts/search_pubmed.py "Alzheimer's disease" \
  --limit 100 \
  --format bibtex \
  --output alzheimers.bib
```

### extract_metadata.py

Extract complete metadata from paper identifiers.

**Features**:
- Supports DOI, PMID, arXiv ID, URL
- Queries CrossRef, PubMed, arXiv APIs
- Handles multiple identifier types
- Batch processing
- Multiple output formats

**Usage**:
```bash
# Single DOI
python scripts/extract_metadata.py --doi 10.1038/s41586-021-03819-2

# Single PMID
python scripts/extract_metadata.py --pmid 34265844

# Single arXiv ID
python scripts/extract_metadata.py --arxiv 2103.14030

# From URL
python scripts/extract_metadata.py \
  --url "https://www.nature.com/articles/s41586-021-03819-2"

# Batch processing (file with one identifier per line)
python scripts/extract_metadata.py \
  --input paper_ids.txt \
  --output references.bib

# Different output formats
python scripts/extract_metadata.py \
  --doi 10.1038/nature12345 \
  --format json  # or bibtex, yaml
```

### validate_citations.py

Validate BibTeX entries for accuracy and completeness.

**Features**:
- DOI verification via doi.org and CrossRef
- Required field checking
- Duplicate detection
- Format validation
- Auto-fix common issues
- Detailed reporting

**Usage**:
```bash
# Basic validation
python scripts/validate_citations.py references.bib

# With auto-fix
python scripts/validate_citations.py references.bib \
  --auto-fix \
  --output fixed_references.bib

# Detailed validation report
python scripts/validate_citations.py references.bib \
  --report validation_report.json \
  --verbose

# Only check DOIs
python scripts/validate_citations.py references.bib \
  --check-dois-only
```

### format_bibtex.py

Format and clean BibTeX files.

**Features**:
- Standardize formatting
- Sort entries (by key, year, author)
- Remove duplicates
- Validate syntax
- Fix common errors
- Enforce citation key conventions

**Usage**:
```bash
# Basic formatting
python scripts/format_bibtex.py references.bib

# Sort by year (newest first)
python scripts/format_bibtex.py references.bib \
  --sort year \
  --descending \
  --output sorted_refs.bib

# Remove duplicates
python scripts/format_bibtex.py references.bib \
  --deduplicate \
  --output clean_refs.bib

# Complete cleanup
python scripts/format_bibtex.py references.bib \
  --deduplicate \
  --sort year \
  --validate \
  --auto-fix \
  --output final_refs.bib
```

### doi_to_bibtex.py

Quick DOI to BibTeX conversion.

**Features**:
- Fast single DOI conversion
- Batch processing
- Multiple output formats
- Clipboard support

**Usage**:
```bash
# Single DOI
python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2

# Multiple DOIs
python scripts/doi_to_bibtex.py \
  10.1038/nature12345 \
  10.1126/science.abc1234 \
  10.1016/j.cell.2023.01.001

# From file (one DOI per line)
python scripts/doi_to_bibtex.py --input dois.txt --output references.bib

# Copy to clipboard
python scripts/doi_to_bibtex.py 10.1038/nature12345 --clipboard
```

## Best Practices

### Search Strategy

1. **Start broad, then narrow**:
   - Begin with general terms to understand the field
   - Refine with specific keywords and filters
   - Use synonyms and related terms

2. **Use multiple sources**:
   - Google Scholar for comprehensive coverage
   - PubMed for biomedical focus
   - arXiv for preprints
   - Combine results for completeness

3. **Leverage citations**:
   - Check "Cited by" for seminal papers
   - Review references from key papers
   - Use citation networks to discover related work

4. **Document your searches**:
   - Save search queries and dates
   - Record number of results
   - Note any filters or restrictions applied

### Metadata Extraction

1. **Always use DOIs when available**:
   - Most reliable identifier
   - Permanent link to the publication
   - Best metadata source via CrossRef

2. **Verify extracted metadata**:
   - Check author names are correct
   - Verify journal/conference names
   - Confirm publication year
   - Validate page numbers and volume

3. **Handle edge cases**:
   - Preprints: Include repository and ID
   - Preprints later published: Use published version
   - Conference papers: Include conference name and location
   - Book chapters: Include book title and editors

4. **Maintain consistency**:
   - Use consistent author name format
   - Standardize journal abbreviations
   - Use same DOI format (URL preferred)

### BibTeX Quality

1. **Follow conventions**:
   - Use meaningful citation keys (FirstAuthor2024keyword)
   - Protect capitalization in titles with {}
   - Use -- for page ranges (not single dash)
   - Include DOI field for all modern publications

2. **Keep it clean**:
   - Remove unnecessary fields
   - No redundant information
   - Consistent formatting
   - Validate syntax regularly

3. **Organize systematically**:
   - Sort by year or topic
   - Group related papers
   - Use separate files for different projects
   - Merge carefully to avoid duplicates

### Validation

1. **Validate early and often**:
   - Check citations when adding them
   - Validate complete bibliography before submission
   - Re-validate after any manual edits

2. **Fix issues promptly**:
   - Broken DOIs: Find correct identifier
   - Missing fields: Extract from original source
   - Duplicates: Choose best version, remove others
   - Format errors: Use auto-fix when safe

3. **Manual review for critical citations**:
   - Verify key papers cited correctly
   - Check author names match publication
   - Confirm page numbers and volume
   - Ensure URLs are current

## Common Pitfalls to Avoid

1. **Single source bias**: Only using Google Scholar or PubMed
   - **Solution**: Search multiple databases for comprehensive coverage

2. **Accepting metadata blindly**: Not verifying extracted information
   - **Solution**: Spot-check extracted metadata against original sources

3. **Ignoring DOI errors**: Broken or incorrect DOIs in bibliography
   - **Solution**: Run validation before final submission

4. **Inconsistent formatting**: Mixed citation key styles, formatting
   - **Solution**: Use format_bibtex.py to standardize

5. **Duplicate entries**: Same paper cited multiple times with different keys
   - **Solution**: Use duplicate detection in validation

6. **Missing required fields**: Incomplete BibTeX entries
   - **Solution**: Validate and ensure all required fields present

7. **Outdated preprints**: Citing preprint when published version exists
   - **Solution**: Check if preprints have been published, update to journal version

8. **Special character issues**: Broken LaTeX compilation due to characters
   - **Solution**: Use proper escaping or Unicode in BibTeX

9. **No validation before submission**: Submitting with citation errors
   - **Solution**: Always run validation as final check

10. **Manual BibTeX entry**: Typing entries by hand
    - **Solution**: Always extract from metadata sources using scripts

## Example Workflows

### Example 1: Building a Bibliography for a Paper

```bash
# Step 1: Find key papers on your topic
python scripts/search_google_scholar.py "transformer neural networks" \
  --year-start 2017 \
  --limit 50 \
  --output transformers_gs.json

python scripts/search_pubmed.py "deep learning medical imaging" \
  --date-start 2020 \
  --limit 50 \
  --output medical_dl_pm.json

# Step 2: Extract metadata from search results
python scripts/extract_metadata.py \
  --input transformers_gs.json \
  --output transformers.bib

python scripts/extract_metadata.py \
  --input medical_dl_pm.json \
  --output medical.bib

# Step 3: Add specific papers you already know
python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2 >> specific.bib
python scripts/doi_to_bibtex.py 10.1126/science.aam9317 >> specific.bib

# Step 4: Combine all BibTeX files
cat transformers.bib medical.bib specific.bib > combined.bib

# Step 5: Format and deduplicate
python scripts/format_bibtex.py combined.bib \
  --deduplicate \
  --sort year \
  --descending \
  --output formatted.bib

# Step 6: Validate
python scripts/validate_citations.py formatted.bib \
  --auto-fix \
  --report validation.json \
  --output final_references.bib

# Step 7: Review any issues
cat validation.json | grep -A 3 '"errors"'

# Step 8: Use in LaTeX
# \bibliography{final_references}
```

### Example 2: Converting a List of DOIs

```bash
# You have a text file with DOIs (one per line)
# dois.txt contains:
# 10.1038/s41586-021-03819-2
# 10.1126/science.aam9317
# 10.1016/j.cell.2023.01.001

# Convert all to BibTeX
python scripts/doi_to_bibtex.py --input dois.txt --output references.bib

# Validate the result
python scripts/validate_citations.py references.bib --verbose
```

### Example 3: Cleaning an Existing BibTeX File

```bash
# You have a messy BibTeX file from various sources
# Clean it up systematically

# Step 1: Format and standardize
python scripts/format_bibtex.py messy_references.bib \
  --output step1_formatted.bib

# Step 2: Remove duplicates
python scripts/format_bibtex.py step1_formatted.bib \
  --deduplicate \
  --output step2_deduplicated.bib

# Step 3: Validate and auto-fix
python scripts/validate_citations.py step2_deduplicated.bib \
  --auto-fix \
  --output step3_validated.bib

# Step 4: Sort by year
python scripts/format_bibtex.py step3_validated.bib \
  --sort year \
  --descending \
  --output clean_references.bib

# Step 5: Final validation report
python scripts/validate_citations.py clean_references.bib \
  --report final_validation.json \
  --verbose

# Review report
cat final_validation.json
```

### Example 4: Finding and Citing Seminal Papers

```bash
# Find highly cited papers on a topic
python scripts/search_google_scholar.py "AlphaFold protein structure" \
  --year-start 2020 \
  --year-end 2024 \
  --sort-by citations \
  --limit 20 \
  --output alphafold_seminal.json

# Extract the top 10 by citation count
# (script will have included citation counts in JSON)

# Convert to BibTeX
python scripts/extract_metadata.py \
  --input alphafold_seminal.json \
  --output alphafold_refs.bib

# The BibTeX file now contains the most influential papers
```

## Integration with Other Skills

### Literature Review Skill

**Citation Management** provides the technical infrastructure for **Literature Review**:

- **Literature Review**: Multi-database systematic search and synthesis
- **Citation Management**: Metadata extraction and validation

**Combined workflow**:
1. Use literature-review for systematic search methodology
2. Use citation-management to extract and validate citations
3. Use literature-review to synthesize findings
4. Use citation-management to ensure bibliography accuracy

### Scientific Writing Skill

**Citation Management** ensures accurate references for **Scientific Writing**:

- Export validated BibTeX for use in LaTeX manuscripts
- Verify citations match publication standards
- Format references according to journal requirements

### Venue Templates Skill

**Citation Management** works with **Venue Templates** for submission-ready manuscripts:

- Different venues require different citation styles
- Generate properly formatted references
- Validate citations meet venue requirements

## Resources

### Bundled Resources

**References** (in `references/`):
- `google_scholar_search.md`: Complete Google Scholar search guide
- `pubmed_search.md`: PubMed and E-utilities API documentation
- `metadata_extraction.md`: Metadata sources and field requirements
- `citation_validation.md`: Validation criteria and quality checks
- `bibtex_formatting.md`: BibTeX entry types and formatting rules

**Scripts** (in `scripts/`):
- `search_google_scholar.py`: Google Scholar search automation
- `search_pubmed.py`: PubMed E-utilities API client
- `extract_metadata.py`: Universal metadata extractor
- `validate_citations.py`: Citation validation and verification
- `format_bibtex.py`: BibTeX formatter and cleaner
- `doi_to_bibtex.py`: Quick DOI to BibTeX converter

**Assets** (in `assets/`):
- `bibtex_template.bib`: Example BibTeX entries for all types
- `citation_checklist.md`: Quality assurance checklist

### External Resources

**Search Engines**:
- Google Scholar: https://scholar.google.com/
- PubMed: https://pubmed.ncbi.nlm.nih.gov/
- PubMed Advanced Search: https://pubmed.ncbi.nlm.nih.gov/advanced/

**Metadata APIs**:
- CrossRef API: https://api.crossref.org/
- PubMed E-utilities: https://www.ncbi.nlm.nih.gov/books/NBK25501/
- arXiv API: https://arxiv.org/help/api/
- DataCite API: https://api.datacite.org/

**Tools and Validators**:
- MeSH Browser: https://meshb.nlm.nih.gov/search
- DOI Resolver: https://doi.org/
- BibTeX Format: http://www.bibtex.org/Format/

**Citation Styles**:
- BibTeX documentation: http://www.bibtex.org/
- LaTeX bibliography management: https://www.overleaf.com/learn/latex/Bibliography_management

## Dependencies

### Required Python Packages

```bash
# Core dependencies
pip install requests  # HTTP requests for APIs
pip install bibtexparser  # BibTeX parsing and formatting
pip install biopython  # PubMed E-utilities access

# Optional (for Google Scholar)
pip install scholarly  # Google Scholar API wrapper
# or
pip install selenium  # For more robust Scholar scraping
```

### Optional Tools

```bash
# For advanced validation
pip install crossref-commons  # Enhanced CrossRef API access
pip install pylatexenc  # LaTeX special character handling
```

## Summary

The citation-management skill provides:

1. **Comprehensive search capabilities** for Google Scholar and PubMed
2. **Automated metadata extraction** from DOI, PMID, arXiv ID, URLs
3. **Citation validation** with DOI verification and completeness checking
4. **BibTeX formatting** with standardization and cleaning tools
5. **Quality assurance** through validation and reporting
6. **Integration** with scientific writing workflow
7. **Reproducibility** through documented search and extraction methods

Use this skill to maintain accurate, complete citations throughout your research and ensure publication-ready bibliographies.


SKILL.md

元数据
namecitation-management
descriptionComprehensive citation management for academic research. Search Google Scholar and PubMed for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries. This skill should be used when you need to find papers, verify citation information, convert DOIs to BibTeX, or ensure reference accuracy in scientific writing.
allowed-toolsRead Write Edit Bash
licenseMIT License
metadata{ "version": "1.0", "skill-author": "K-Dense Inc." }

引用管理

概述

在研究与写作全流程中系统管理引用。本技能提供搜索学术数据库(Google Scholar、PubMed)、从多源(CrossRef、PubMed、arXiv)提取准确元数据、验证引用信息及生成格式正确 BibTeX 条目的工具与策略。

对维护引用准确性、避免引用错误、确保可复现研究至关重要。与文献综述技能无缝集成,构成完整研究流程。

何时使用本技能

遇到以下情况时使用:

  • 在 Google Scholar 或 PubMed 上搜索特定论文
  • 将 DOI、PMID 或 arXiv ID 转换为格式正确的 BibTeX
  • 提取引用所需的完整元数据(作者、标题、期刊、年份等)
  • 验证已有引文的准确性
  • 清理并格式化 BibTeX 文件
  • 查找某领域的高被引论文
  • 核实引用信息与实际出版物是否一致
  • 为稿件或学位论文构建参考文献列表
  • 检查重复引用
  • 确保引用格式一致性

使用科学示意图增强视觉效果

使用本技能创建文档时,始终考虑添加科学图表与示意图以提升视觉传达效果。

如果文档尚未包含示意图或图表:

  • 使用 scientific-schematics 技能生成 AI 驱动的出版级图表
  • 用自然语言描述所需图表
  • Nano Banana Pro 将自动生成、审阅并优化示意图

对于新文档: 应默认生成科学示意图,以可视化呈现文中描述的关键概念、工作流、架构或关系。

生成示意图的方法:

bash
python scripts/generate_schematic.py "your diagram description" -o figures/output.png

AI 将自动:

  • 创建格式规范的出版级图像
  • 通过多次迭代进行审阅与优化
  • 确保可访问性(色盲友好、高对比度)
  • 将输出保存到 figures/ 目录

何时添加示意图:

  • 引用工作流图
  • 文献检索方法学流程图
  • 参考文献管理系统架构
  • 引用风格决策树
  • 数据库集成图
  • 任何从可视化中受益的复杂概念

关于创建示意图的详细指导,请参阅 scientific-schematics 技能文档。


核心工作流

引用管理遵循系统化流程:

第一阶段:论文发现与搜索

目标:使用学术搜索引擎查找相关论文。

Google Scholar 搜索

Google Scholar 覆盖学科最全面。

基础搜索:

bash
# 按主题搜索论文
python scripts/search_google_scholar.py "CRISPR gene editing" \
  --limit 50 \
  --output results.json

# 带年份过滤的搜索
python scripts/search_google_scholar.py "machine learning protein folding" \
  --year-start 2020 \
  --year-end 2024 \
  --limit 100 \
  --output ml_proteins.json

高级搜索策略(详见 references/google_scholar_search.md):

  • 使用引号进行精确短语匹配:"deep learning"
  • 按作者搜索:author:LeCun
  • 在标题中搜索:intitle:"neural networks"
  • 排除词条:machine learning -survey
  • 使用排序选项查找高被引论文
  • 按日期范围过滤以获取近期成果

最佳实践:

  • 使用具体、有针对性的搜索词
  • 包含关键术语与缩写
  • 对快速发展的领域按最近年份过滤
  • 查看“被引用次数”以发现开创性论文
  • 导出排名靠前的结果以供进一步分析

PubMed 搜索

PubMed 专精生物医学和生命科学文献(3500 万+引文)。

基础搜索:

bash
# 搜索 PubMed
python scripts/search_pubmed.py "Alzheimer's disease treatment" \
  --limit 100 \
  --output alzheimers.json

# 使用 MeSH 词条与过滤器搜索
python scripts/search_pubmed.py \
  --query '"Alzheimer Disease"[MeSH] AND "Drug Therapy"[MeSH]' \
  --date-start 2020 \
  --date-end 2024 \
  --publication-types "Clinical Trial,Review" \
  --output alzheimers_trials.json

高级 PubMed 查询(详见 references/pubmed_search.md):

  • 使用 MeSH 词条:"Diabetes Mellitus"[MeSH]
  • 字段标签:"cancer"[Title]、"Smith J"[Author]
  • 布尔运算符:AND、OR、NOT
  • 日期过滤:2020:2024[Publication Date]
  • 出版物类型:"Review"[Publication Type]
  • 结合 E-utilities API 自动化

最佳实践:

  • 使用 MeSH Browser 查找正确的控制词汇
  • 先在 PubMed Advanced Search Builder 中构建复杂查询
  • 用 OR 连接多个同义词
  • 检索 PMID 以便于元数据提取
  • 导出为 JSON 或直接转为 BibTeX

第二阶段:元数据提取

目标:将论文标识符(DOI、PMID、arXiv ID)转换为完整、准确的元数据。

快速 DOI 转 BibTeX

对于单个 DOI,使用快速转换工具:

bash
# 转换单个 DOI
python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2

# 从文件转换多个 DOI
python scripts/doi_to_bibtex.py --input dois.txt --output references.bib

# 其他输出格式
python scripts/doi_to_bibtex.py 10.1038/nature12345 --format json

综合元数据提取

适用于 DOI、PMID、arXiv ID 或 URL:

bash
# 从 DOI 提取
python scripts/extract_metadata.py --doi 10.1038/s41586-021-03819-2

# 从 PMID 提取
python scripts/extract_metadata.py --pmid 34265844

# 从 arXiv ID 提取
python scripts/extract_metadata.py --arxiv 2103.14030

# 从 URL 提取
python scripts/extract_metadata.py --url "https://www.nature.com/articles/s41586-021-03819-2"

# 从文件批量提取(混合标识符)
python scripts/extract_metadata.py --input identifiers.txt --output citations.bib

元数据来源(详见 references/metadata_extraction.md):

  1. CrossRef API:DOI 主来源

    • 期刊文章全面元数据
    • 出版商提供信息
    • 包含作者、标题、期刊、卷、页码、日期
    • 免费,无需 API 密钥
  2. PubMed E-utilities:生物医学文献

    • 官方 NCBI 元数据
    • 包含 MeSH 词条、摘要
    • PMID 和 PMCID 标识符
    • 免费,高流量时建议使用 API 密钥
  3. arXiv API:物理、数学、CS、q-bio 等领域的预印本

    • 预印本完整元数据
    • 版本跟踪
    • 作者单位
    • 免费、开放获取
  4. DataCite API:研究数据集、软件及其他资源

    • 非传统学术成果的元数据
    • 数据集和代码的 DOI
    • 免费访问

提取内容:

  • 必填字段:作者、标题、年份
  • 期刊文章:期刊、卷、期、页码、DOI
  • 图书:出版商、ISBN、版本
  • 会议论文:会议名称、地点、页码
  • 预印本:存档库(arXiv、bioRxiv)、预印本 ID
  • 附加项:摘要、关键词、URL

第三阶段:BibTeX 格式化

目标:生成干净、格式正确的 BibTeX 条目。

理解 BibTeX 条目类型

完整指南见 references/bibtex_formatting.md。

常见条目类型:

  • @article:期刊文章(最常见)
  • @book:图书
  • @inproceedings:会议论文
  • @incollection:书章节
  • @phdthesis:学位论文
  • @misc:预印本、软件、数据集

各类型必填字段:

bibtex
@article{citationkey,
  author  = {Last1, First1 and Last2, First2},
  title   = {Article Title},
  journal = {Journal Name},
  year    = {2024},
  volume  = {10},
  number  = {3},
  pages   = {123--145},
  doi     = {10.1234/example}
}

@inproceedings{citationkey,
  author    = {Last, First},
  title     = {Paper Title},
  booktitle = {Conference Name},
  year      = {2024},
  pages     = {1--10}
}

@book{citationkey,
  author    = {Last, First},
  title     = {Book Title},
  publisher = {Publisher Name},
  year      = {2024}
}

格式化与清理

使用格式化工具标准化 BibTeX 文件:

bash
# 格式化并清理 BibTeX 文件
python scripts/format_bibtex.py references.bib \
  --output formatted_references.bib

# 按引用键排序
python scripts/format_bibtex.py references.bib \
  --sort key \
  --output sorted_references.bib

# 按年份排序(最新优先)
python scripts/format_bibtex.py references.bib \
  --sort year \
  --descending \
  --output sorted_references.bib

# 移除重复项
python scripts/format_bibtex.py references.bib \
  --deduplicate \
  --output clean_references.bib

# 验证并报告问题
python scripts/format_bibtex.py references.bib \
  --validate \
  --report validation_report.txt

格式化操作:

  • 标准化字段顺序
  • 一致的缩进与间距
  • 标题中正确的大小写保护(用 {})
  • 标准化的作者姓名格式
  • 一致的引用键格式
  • 移除非必要字段
  • 修正常见错误(缺失逗号、大括号)

第四阶段:引用验证

目标:确认所有引用准确且完整。

综合验证

bash
# 验证 BibTeX 文件
python scripts/validate_citations.py references.bib

# 验证并自动修复常见问题
python scripts/validate_citations.py references.bib \
  --auto-fix \
  --output validated_references.bib

# 生成详细验证报告
python scripts/validate_citations.py references.bib \
  --report validation_report.json \
  --verbose

验证检查(详见 references/citation_validation.md):

  1. DOI 核实:

    • DOI 通过 doi.org 正确解析
    • BibTeX 与 CrossRef 元数据匹配
    • 无失效或无效 DOI
  2. 必填字段:

    • 各条目类型所需字段齐全
    • 无空白或缺失关键信息
    • 作者姓名格式正确
  3. 数据一致性:

    • 年份有效(4 位数字、合理范围)
    • 卷/期为数字
    • 页码格式正确(如 123--145)
    • URL 可访问
  4. 重复检测:

    • 同一 DOI 多次使用
    • 相似标题(疑似重复)
    • 相同作者/年份/标题组合
  5. 格式合规:

    • 有效 BibTeX 语法
    • 正确的大括号与引号
    • 引用键唯一
    • 特殊字符正确处理

验证输出示例:

json
{
  "total_entries": 150,
  "valid_entries": 145,
  "errors": [
    {
      "citation_key": "Smith2023",
      "error_type": "missing_field",
      "field": "journal",
      "severity": "high"
    },
    {
      "citation_key": "Jones2022",
      "error_type": "invalid_doi",
      "doi": "10.1234/broken",
      "severity": "high"
    }
  ],
  "warnings": [
    {
      "citation_key": "Brown2021",
      "warning_type": "possible_duplicate",
      "duplicate_of": "Brown2021a",
      "severity": "medium"
    }
  ]
}

第五阶段:与写作工作流集成

为稿件构建参考文献

创建参考文献列表的完整工作流:

bash
# 1. 搜索主题相关论文
python scripts/search_pubmed.py \
  '"CRISPR-Cas Systems"[MeSH] AND "Gene Editing"[MeSH]' \
  --date-start 2020 \
  --limit 200 \
  --output crispr_papers.json

# 2. 从搜索结果提取 DOI 并转换为 BibTeX
python scripts/extract_metadata.py \
  --input crispr_papers.json \
  --output crispr_refs.bib

# 3. 通过 DOI 添加特定论文
python scripts/doi_to_bibtex.py 10.1038/nature12345 >> crispr_refs.bib
python scripts/doi_to_bibtex.py 10.1126/science.abcd1234 >> crispr_refs.bib

# 4. 格式化并清理 BibTeX 文件
python scripts/format_bibtex.py crispr_refs.bib \
  --deduplicate \
  --sort year \
  --descending \
  --output references.bib

# 5. 验证所有引用
python scripts/validate_citations.py references.bib \
  --auto-fix \
  --report validation.json \
  --output final_references.bib

# 6. 查看验证报告并修正剩余问题
cat validation.json

# 7. 在 LaTeX 文档中使用
# \bibliography{final_references}

与文献综述技能集成

本技能与 literature-review 技能互补:

文献综述技能 → 系统搜索与综合 引用管理技能 → 技术性引用处理

组合工作流:

  1. 使用 literature-review 进行全面的多数据库搜索
  2. 使用 citation-management 提取并验证所有引用
  3. 使用 literature-review 按主题综合发现
  4. 使用 citation-management 验证最终参考文献的准确性
bash
# 完成文献综述后
# 验证综述文档中的所有引用
python scripts/validate_citations.py my_review_references.bib --report review_validation.json

# 若需要特定引用风格,可进行格式化
python scripts/format_bibtex.py my_review_references.bib \
  --style nature \
  --output formatted_refs.bib

搜索策略

Google Scholar 最佳实践

寻找开创性与高影响力论文(至关重要):

始终根据引用次数、发表平台质量、作者声誉优先筛选论文:

引用次数阈值:

论文年限引用次数分类
0-3年20+值得关注
0-3年100+极具影响力
3-7年100+显著
3-7年500+里程碑式论文
7年以上500+开创性工作
7年以上1000+奠基性

发表平台质量分级:

  • 第一梯队(优先):Nature, Science, Cell, NEJM, Lancet, JAMA, PNAS
  • 第二梯队(高度优先):影响因子 >10 的期刊,顶级会议(NeurIPS、ICML、ICLR)
  • 第三梯队(良好):专业期刊(IF 5-10)
  • 第四梯队(谨慎):影响力较低的同行评审平台

作者声誉指标:

  • 高年资研究者,h-index >40
  • 在第一梯队平台有多篇发表
  • 在知名机构担任领导职位
  • 获奖与编辑职务

高影响力论文搜索策略:

  • 按引用次数排序(被引最多优先)
  • 查看第一梯队期刊的综述文章以获得概览
  • 检查“被引用次数”以评估影响力及后续跟进工作
  • 使用引用提醒跟踪关键论文的新引用
  • 使用 source:Nature 或 source:Science 过滤顶级平台
  • 使用 author:LastName 搜索领域知名学者的论文

高级操作符(完整列表见 references/google_scholar_search.md):

text
"exact phrase"           # 精确短语匹配
author:lastname          # 按作者搜索
intitle:keyword          # 仅搜索标题
source:journal           # 搜索特定期刊
-exclude                 # 排除词条
OR                       # 备选词条
2020..2024              # 年份范围

搜索示例:

text
# 查找某一主题的近期综述
"CRISPR" intitle:review 2023..2024

# 查找特定作者在某一主题的论文
author:Church "synthetic biology"

# 查找高被引奠基性工作
"deep learning" 2012..2015 sort:citations

# 排除综述,聚焦方法
"protein folding" -survey -review intitle:method

PubMed 最佳实践

使用 MeSH 词条: MeSH(医学主题词表)提供受控词汇以精准搜索。

  1. 在 https://meshb.nlm.nih.gov/search 查找 MeSH 词条
  2. 在查询中使用:"Diabetes Mellitus, Type 2"[MeSH]
  3. 与关键词结合以全面覆盖

字段标签:

text
[Title]              # 仅搜索标题
[Title/Abstract]     # 搜索标题或摘要
[Author]             # 按作者姓名搜索
[Journal]            # 搜索特定期刊
[Publication Date]   # 日期范围
[Publication Type]   # 文章类型
[MeSH]              # MeSH 词条

构建复杂查询:

bash
# 近期发表的糖尿病治疗临床试验
"Diabetes Mellitus, Type 2"[MeSH] AND "Drug Therapy"[MeSH] 
AND "Clinical Trial"[Publication Type] AND 2020:2024[Publication Date]

# 特定期刊上关于 CRISPR 的综述
"CRISPR-Cas Systems"[MeSH] AND "Nature"[Journal] AND "Review"[Publication Type]

# 某作者的近期研究
"Smith AB"[Author] AND cancer[Title/Abstract] AND 2022:2024[Publication Date]

用于自动化的 E-utilities: 脚本使用 NCBI E-utilities API 进行程序化访问:

  • ESearch:搜索并检索 PMID
  • EFetch:获取完整元数据
  • ESummary:获取摘要信息
  • ELink:查找相关文章

完整 API 文档见 references/pubmed_search.md。

工具与脚本

search_google_scholar.py

搜索 Google Scholar 并导出结果。

功能:

  • 带速率限制的自动搜索
  • 分页支持
  • 年份范围过滤
  • 导出为 JSON 或 BibTeX
  • 引用次数信息

用法:

bash
# 基础搜索
python scripts/search_google_scholar.py "quantum computing"

# 带过滤器的高级搜索
python scripts/search_google_scholar.py "quantum computing" \
  --year-start 2020 \
  --year-end 2024 \
  --limit 100 \
  --sort-by citations \
  --output quantum_papers.json

# 直接导出为 BibTeX
python scripts/search_google_scholar.py "machine learning" \
  --limit 50 \
  --format bibtex \
  --output ml_papers.bib

search_pubmed.py

使用 E-utilities API 搜索 PubMed。

功能:

  • 复杂查询支持(MeSH、字段标签、布尔运算符)
  • 日期范围过滤
  • 出版物类型过滤
  • 批量检索元数据
  • 导出为 JSON 或 BibTeX

用法:

bash
# 简单关键词搜索
python scripts/search_pubmed.py "CRISPR gene editing"

# 带过滤器的复杂查询
python scripts/search_pubmed.py \
  --query '"CRISPR-Cas Systems"[MeSH] AND "therapeutic"[Title/Abstract]' \
  --date-start 2020-01-01 \
  --date-end 2024-12-31 \
  --publication-types "Clinical Trial,Review" \
  --limit 200 \
  --output crispr_therapeutic.json

# 导出为 BibTeX
python scripts/search_pubmed.py "Alzheimer's disease" \
  --limit 100 \
  --format bibtex \
  --output alzheimers.bib

extract_metadata.py

从论文标识符提取完整元数据。

功能:

  • 支持 DOI、PMID、arXiv ID、URL
  • 查询 CrossRef、PubMed、arXiv API
  • 处理多种标识符类型
  • 批量处理
  • 多种输出格式

用法:

bash
# 单个 DOI
python scripts/extract_metadata.py --doi 10.1038/s41586-021-03819-2

# 单个 PMID
python scripts/extract_metadata.py --pmid 34265844

# 单个 arXiv ID
python scripts/extract_metadata.py --arxiv 2103.14030

# 从 URL
python scripts/extract_metadata.py \
  --url "https://www.nature.com/articles/s41586-021-03819-2"

# 批量处理(文件每行一个标识符)
python scripts/extract_metadata.py \
  --input paper_ids.txt \
  --output references.bib

# 不同输出格式
python scripts/extract_metadata.py \
  --doi 10.1038/nature12345 \
  --format json  # 或 bibtex、yaml

validate_citations.py

验证 BibTeX 条目的准确性与完整性。

功能:

  • 通过 doi.org 和 CrossRef 核实 DOI
  • 必填字段检查
  • 重复检测
  • 格式验证
  • 自动修复常见问题
  • 详细报告

用法:

bash
# 基础验证
python scripts/validate_citations.py references.bib

# 自动修复
python scripts/validate_citations.py references.bib \
  --auto-fix \
  --output fixed_references.bib

# 详细验证报告
python scripts/validate_citations.py references.bib \
  --report validation_report.json \
  --verbose

# 仅检查 DOI
python scripts/validate_citations.py references.bib \
  --check-dois-only

format_bibtex.py

格式化并清理 BibTeX 文件。

功能:

  • 标准化格式
  • 排序条目(按键名、年份、作者)
  • 移除重复
  • 验证语法
  • 修正常见错误
  • 强制执行引用键命名规范

用法:

bash
# 基础格式化
python scripts/format_bibtex.py references.bib

# 按年份排序(最新优先)
python scripts/format_bibtex.py references.bib \
  --sort year \
  --descending \
  --output sorted_refs.bib

# 移除重复
python scripts/format_bibtex.py references.bib \
  --deduplicate \
  --output clean_refs.bib

# 完整清理
python scripts/format_bibtex.py references.bib \
  --deduplicate \
  --sort year \
  --validate \
  --auto-fix \
  --output final_refs.bib

doi_to_bibtex.py

快速将 DOI 转换为 BibTeX。

功能:

  • 高速单个 DOI 转换
  • 批量处理
  • 多种输出格式
  • 剪贴板支持

用法:

bash
# 单个 DOI
python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2

# 多个 DOI
python scripts/doi_to_bibtex.py \
  10.1038/nature12345 \
  10.1126/science.abc1234 \
  10.1016/j.cell.2023.01.001

# 从文件读取(每行一个 DOI)
python scripts/doi_to_bibtex.py --input dois.txt --output references.bib

# 复制到剪贴板
python scripts/doi_to_bibtex.py 10.1038/nature12345 --clipboard

最佳实践

搜索策略

  1. 先宽泛后精准:

    • 从通用词条入手了解领域
    • 用特定关键词和过滤器精炼
    • 使用同义词和相关术语
  2. 使用多源:

    • Google Scholar 覆盖全面
    • PubMed 聚焦生物医学
    • arXiv 获取预印本
    • 组合结果以确保完整性
  3. 利用引用:

    • 查看“被引用次数”发现开创性论文
    • 查阅关键论文的参考文献
    • 利用引用网络发现相关研究
  4. 记录搜索过程:

    • 保存搜索查询与日期
    • 记录结果数量
    • 注明应用的过滤器或限制

元数据提取

  1. 有 DOI 时务必使用:

    • 最可靠的标识符
    • 出版物永久链接
    • 通过 CrossRef 获取最佳元数据
  2. 核实提取的元数据:

    • 检查作者姓名是否正确
    • 核实期刊/会议名称
    • 确认出版年份
    • 验证页码与卷号
  3. 处理边缘情况:

    • 预印本:包含存档库与 ID
    • 已出版的预印本:使用正式出版版本
    • 会议论文:包含会议名称与地点
    • 书章节:包含书名与编辑
  4. 保持一致性:

    • 使用一致的作者姓名格式
    • 标准化期刊缩写
    • 使用相同的 DOI 格式(推荐 URL 格式)

BibTeX 质量

  1. 遵循惯例:

    • 使用有意义引用键(如 FirstAuthor2024keyword)
    • 用 {} 保护标题中的大小写
    • 使用 -- 表示页码范围(非单破折号)
    • 所有现代出版物包含 DOI 字段
  2. 保持整洁:

    • 移除非必要字段
    • 无冗余信息
    • 格式一致
    • 定期验证语法
  3. 系统组织:

    • 按年份或主题排序
    • 将相关论文分组
    • 不同项目使用独立文件
    • 合并时谨慎避免重复

验证

  1. 尽早并经常验证:

    • 添加引用时即检查
    • 投稿前验证完整参考文献
    • 任何手工编辑后重新验证
  2. 及时修复问题:

    • DOI 失效:查找正确标识符
    • 缺失字段:从原始来源提取
    • 重复条目:选择最佳版本,移除其他
    • 格式错误:安全时使用自动修复
  3. 手工审核关键引用:

    • 验证关键论文信息无误
    • 检查标题准确性
    • 确认期刊/会议名称
    • 作者列表正确