科研技能库/数据可视化
图表可视化
未发现用户侧风险

数据可视化

创建有效数据可视化所需的图表选择指南、Python 可视化代码模式、设计原则和无障碍性考量。

文件预览

1 个文件
SKILL.md
9.2 KB · 可预览
---
name: visualization
description: "Chart selection guidance, Python visualization code patterns, design principles, and accessibility considerations for creating effective data visualizations."
---

# Data Visualization Skill

Chart selection guidance, Python visualization code patterns, design principles, and accessibility considerations for creating effective data visualizations.

**Design defaults:** See `skills/GENERAL-KNOWLEDGE-WORKER/design-foundations/SKILL.md` for the canonical chart color sequence and default palette.

## Chart Selection Guide

### Choose by Data Relationship

| What You're Showing | Best Chart | Alternatives |
|---|---|---|
| **Trend over time** | Line chart | Area chart (if showing cumulative or composition) |
| **Comparison across categories** | Vertical bar chart | Horizontal bar (many categories), lollipop chart |
| **Ranking** | Horizontal bar chart | Dot plot, slope chart (comparing two periods) |
| **Part-to-whole composition** | Stacked bar chart | Treemap (hierarchical), waffle chart |
| **Composition over time** | Stacked area chart | 100% stacked bar (for proportion focus) |
| **Distribution** | Histogram | Box plot (comparing groups), violin plot, strip plot |
| **Correlation (2 variables)** | Scatter plot | Bubble chart (add 3rd variable as size) |
| **Correlation (many variables)** | Heatmap (correlation matrix) | Pair plot |
| **Geographic patterns** | Choropleth map | Bubble map, hex map |
| **Flow / process** | Sankey diagram | Funnel chart (sequential stages) |
| **Relationship network** | Network graph | Chord diagram |
| **Performance vs. target** | Bullet chart | Gauge (single KPI only) |
| **Multiple KPIs at once** | Small multiples | Dashboard with separate charts |

### When NOT to Use Certain Charts

- **Pie charts**: Avoid unless <6 categories and exact proportions matter less than rough comparison. Humans are bad at comparing angles. Use bar charts instead.
- **3D charts**: Never. They distort perception and add no information.
- **Dual-axis charts**: Use cautiously. They can mislead by implying correlation. Clearly label both axes if used.
- **Stacked bar (many categories)**: Hard to compare middle segments. Use small multiples or grouped bars instead.
- **Donut charts**: Slightly better than pie charts but same fundamental issues. Use for single KPI display at most.

## Python Visualization Code Patterns

### Setup and Style

```python
import matplotlib.pyplot as plt
import matplotlib.ticker as mticker
import seaborn as sns
import pandas as pd
import numpy as np

# Professional style setup
plt.style.use('seaborn-v0_8-whitegrid')
plt.rcParams.update({
    'figure.figsize': (10, 6),
    'figure.dpi': 150,
    'font.size': 11,
    'axes.titlesize': 14,
    'axes.titleweight': 'bold',
    'axes.labelsize': 11,
    'xtick.labelsize': 10,
    'ytick.labelsize': 10,
    'legend.fontsize': 10,
    'figure.titlesize': 16,
})

# Default categorical palette — from skills/GENERAL-KNOWLEDGE-WORKER/design-foundations/SKILL.md
# See skills/GENERAL-KNOWLEDGE-WORKER/design-foundations/SKILL.md for full chart color guidance
PALETTE_CATEGORICAL = ['#20808D', '#A84B2F', '#1B474D', '#BCE2E7', '#944454', '#FFC553', '#848456', '#6E522B']
PALETTE_SEQUENTIAL = 'YlOrRd'
PALETTE_DIVERGING = 'RdBu_r'
```

### Line Chart (Time Series)

```python
fig, ax = plt.subplots(figsize=(10, 6))

for label, group in df.groupby('category'):
    ax.plot(group['date'], group['value'], label=label, linewidth=2)

ax.set_title('Metric Trend by Category', fontweight='bold')
ax.set_xlabel('Date')
ax.set_ylabel('Value')
ax.legend(loc='upper left', frameon=True)
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)

# Format dates on x-axis
fig.autofmt_xdate()

plt.tight_layout()
plt.savefig('trend_chart.png', dpi=150, bbox_inches='tight')
```

### Bar Chart (Comparison)

```python
fig, ax = plt.subplots(figsize=(10, 6))

# Sort by value for easy reading
df_sorted = df.sort_values('metric', ascending=True)

bars = ax.barh(df_sorted['category'], df_sorted['metric'], color=PALETTE_CATEGORICAL[0])

# Add value labels
for bar in bars:
    width = bar.get_width()
    ax.text(width + 0.5, bar.get_y() + bar.get_height()/2,
            f'{width:,.0f}', ha='left', va='center', fontsize=10)

ax.set_title('Metric by Category (Ranked)', fontweight='bold')
ax.set_xlabel('Metric Value')
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)

plt.tight_layout()
plt.savefig('bar_chart.png', dpi=150, bbox_inches='tight')
```

### Histogram (Distribution)

```python
fig, ax = plt.subplots(figsize=(10, 6))

ax.hist(df['value'], bins=30, color=PALETTE_CATEGORICAL[0], edgecolor='white', alpha=0.8)

# Add mean and median lines
mean_val = df['value'].mean()
median_val = df['value'].median()
ax.axvline(mean_val, color='red', linestyle='--', linewidth=1.5, label=f'Mean: {mean_val:,.1f}')
ax.axvline(median_val, color='green', linestyle='--', linewidth=1.5, label=f'Median: {median_val:,.1f}')

ax.set_title('Distribution of Values', fontweight='bold')
ax.set_xlabel('Value')
ax.set_ylabel('Frequency')
ax.legend()
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)

plt.tight_layout()
plt.savefig('histogram.png', dpi=150, bbox_inches='tight')
```

### Heatmap

```python
fig, ax = plt.subplots(figsize=(10, 8))

# Pivot data for heatmap format
pivot = df.pivot_table(index='row_dim', columns='col_dim', values='metric', aggfunc='sum')

sns.heatmap(pivot, annot=True, fmt=',.0f', cmap='YlOrRd',
            linewidths=0.5, ax=ax, cbar_kws={'label': 'Metric Value'})

ax.set_title('Metric by Row Dimension and Column Dimension', fontweight='bold')
ax.set_xlabel('Column Dimension')
ax.set_ylabel('Row Dimension')

plt.tight_layout()
plt.savefig('heatmap.png', dpi=150, bbox_inches='tight')
```

### Small Multiples

```python
categories = df['category'].unique()
n_cats = len(categories)
n_cols = min(3, n_cats)
n_rows = (n_cats + n_cols - 1) // n_cols

fig, axes = plt.subplots(n_rows, n_cols, figsize=(5*n_cols, 4*n_rows), sharex=True, sharey=True)
axes = axes.flatten() if n_cats > 1 else [axes]

for i, cat in enumerate(categories):
    ax = axes[i]
    subset = df[df['category'] == cat]
    ax.plot(subset['date'], subset['value'], color=PALETTE_CATEGORICAL[i % len(PALETTE_CATEGORICAL)])
    ax.set_title(cat, fontsize=12)
    ax.spines['top'].set_visible(False)
    ax.spines['right'].set_visible(False)

# Hide empty subplots
for j in range(i+1, len(axes)):
    axes[j].set_visible(False)

fig.suptitle('Trends by Category', fontsize=14, fontweight='bold', y=1.02)
plt.tight_layout()
plt.savefig('small_multiples.png', dpi=150, bbox_inches='tight')
```

### Number Formatting Helpers

```python
def format_number(val, fmt='number'):
    """Format numbers for chart labels: 'number', 'currency', or 'percent'."""
    if fmt == 'percent':
        return f'{val:.1f}%'
    prefix = '$' if fmt == 'currency' else ''
    if abs(val) >= 1e9: return f'{prefix}{val/1e9:.1f}B'
    if abs(val) >= 1e6: return f'{prefix}{val/1e6:.1f}M'
    if abs(val) >= 1e3: return f'{prefix}{val/1e3:.1f}K'
    return f'{prefix}{val:,.0f}'

# Usage with axis formatter
ax.yaxis.set_major_formatter(mticker.FuncFormatter(lambda x, p: format_number(x, 'currency')))
```

### Interactive Charts with Plotly

```python
import plotly.express as px
import plotly.graph_objects as go

# Simple interactive line chart
fig = px.line(df, x='date', y='value', color='category',
              title='Interactive Metric Trend',
              labels={'value': 'Metric Value', 'date': 'Date'})
fig.update_layout(hovermode='x unified')
fig.write_html('interactive_chart.html')
fig.show()

# Interactive scatter with hover data
fig = px.scatter(df, x='metric_a', y='metric_b', color='category',
                 size='size_metric', hover_data=['name', 'detail_field'],
                 title='Correlation Analysis')
fig.show()
```

## Design Principles

Design-foundations covers color theory, data-ink ratio, labeling, and accessibility rules. Below adds chart-specific guidance not in those files.

- **Highlight the story**: Bright accent for the key insight; grey everything else
- **Titles state insights**: "Revenue grew 23% YoY" not "Revenue by Month". Subtitle adds date range, filters, source
- **Axis labels**: Never rotated 90°. Shorten or wrap. Data labels on key points only, not every bar
- **Sort meaningfully**: By value (not alphabetically) unless natural order exists (months, stages)
- **Aspect ratio**: Time series wider than tall (3:1 to 2:1); comparisons squarer
- **Bar charts start at zero**: Always. Line charts can have non-zero baselines when range matters
- **Consistent scales across panels**: Same axis range when comparing multiple charts
- **Show uncertainty**: Error bars, confidence intervals, or ranges when data is uncertain

## Accessibility

Design-foundations covers color independence and contrast rules. Python-specific additions:

- Use `sns.color_palette("colorblind")` as an alternative colorblind-safe palette
- Add pattern fills (`hatch` in matplotlib) or different line styles alongside color
- Include alt text describing the chart's key finding; provide data table alternative
- Test: does the chart work in B&W? Text readable at standard zoom?

**Before sharing:** series distinguishable without color, title states the insight, axes labeled with units, legend clear, data source noted.

SKILL.md

元数据
namedata-visualization
description创建有效数据可视化所需的图表选择指南、Python 可视化代码模式、设计原则和无障碍性考量。

数据可视化技能

创建有效数据可视化所需的图表选择指南、Python 可视化代码模式、设计原则和无障碍性考量。

设计默认值: 规范图表颜色序列和默认调色板,详见 skills/GENERAL-KNOWLEDGE-WORKER/design-foundations/SKILL.md。

图表选择指南

按数据关系选择

展示内容最佳图表备选方案
随时间变化的趋势折线图面积图(展示累计或构成)
跨类别比较垂直柱状图水平柱状图(类别多时)、棒棒糖图
排名水平柱状图点图、斜率图(比较两个时期)
部分与整体构成堆叠柱状图树状图(层次结构)、华夫饼图
随时间变化的构成堆叠面积图100% 堆叠柱状图(关注比例)
分布直方图箱线图(比较分组)、小提琴图、散列图
相关性(2 个变量)散点图气泡图(将第3变量映射为大小)
相关性(多变量)热力图(相关矩阵)成对关系图
地理模式等值区域图气泡地图、六边形地图
流程 / 过程桑基图漏斗图(有序阶段)
关系网络网络图和弦图
绩效 vs. 目标子弹图仪表盘(仅单个 KPI)
同时展示多个 KPI小多组图含独立图表的仪表盘

何时不应使用特定图表

  • 饼图:除非类别少于 6 个且精确比例不如粗略比较重要,否则避免使用。人类不擅长比较角度。改用柱状图。
  • 3D 图表:绝不使用。它们扭曲感知且不增加信息。
  • 双轴图表:谨慎使用。可能通过暗示相关性而产生误导。若使用,请清楚标注两个坐标轴。
  • 多类别堆叠柱状图:难以比较中间段。改用小多组图或分组柱状图。
  • 环形图:比饼图稍好但存在相同根本问题。最多用于单一 KPI 展示。

Python 可视化代码模式

设置和样式

python
import matplotlib.pyplot as plt
import matplotlib.ticker as mticker
import seaborn as sns
import pandas as pd
import numpy as np

# 专业样式设置
plt.style.use('seaborn-v0_8-whitegrid')
plt.rcParams.update({
    'figure.figsize': (10, 6),
    'figure.dpi': 150,
    'font.size': 11,
    'axes.titlesize': 14,
    'axes.titleweight': 'bold',
    'axes.labelsize': 11,
    'xtick.labelsize': 10,
    'ytick.labelsize': 10,
    'legend.fontsize': 10,
    'figure.titlesize': 16,
})

# 默认分类调色板 — 来自 skills/GENERAL-KNOWLEDGE-WORKER/design-foundations/SKILL.md
# 完整图表颜色指导可见 skills/GENERAL-KNOWLEDGE-WORKER/design-foundations/SKILL.md
PALETTE_CATEGORICAL = ['#20808D', '#A84B2F', '#1B474D', '#BCE2E7', '#944454', '#FFC553', '#848456', '#6E522B']
PALETTE_SEQUENTIAL = 'YlOrRd'
PALETTE_DIVERGING = 'RdBu_r'

折线图(时间序列)

python
fig, ax = plt.subplots(figsize=(10, 6))

for label, group in df.groupby('category'):
    ax.plot(group['date'], group['value'], label=label, linewidth=2)

ax.set_title('Metric Trend by Category', fontweight='bold')
ax.set_xlabel('Date')
ax.set_ylabel('Value')
ax.legend(loc='upper left', frameon=True)
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)

# 格式化 x 轴日期
fig.autofmt_xdate()

plt.tight_layout()
plt.savefig('trend_chart.png', dpi=150, bbox_inches='tight')

柱状图(比较)

python
fig, ax = plt.subplots(figsize=(10, 6))

# 按值排序以便阅读
df_sorted = df.sort_values('metric', ascending=True)

bars = ax.barh(df_sorted['category'], df_sorted['metric'], color=PALETTE_CATEGORICAL[0])

# 添加数值标签
for bar in bars:
    width = bar.get_width()
    ax.text(width + 0.5, bar.get_y() + bar.get_height()/2,
            f'{width:,.0f}', ha='left', va='center', fontsize=10)

ax.set_title('Metric by Category (Ranked)', fontweight='bold')
ax.set_xlabel('Metric Value')
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)

plt.tight_layout()
plt.savefig('bar_chart.png', dpi=150, bbox_inches='tight')

直方图(分布)

python
fig, ax = plt.subplots(figsize=(10, 6))

ax.hist(df['value'], bins=30, color=PALETTE_CATEGORICAL[0], edgecolor='white', alpha=0.8)

# 添加均值和中位数线
mean_val = df['value'].mean()
median_val = df['value'].median()
ax.axvline(mean_val, color='red', linestyle='--', linewidth=1.5, label=f'Mean: {mean_val:,.1f}')
ax.axvline(median_val, color='green', linestyle='--', linewidth=1.5, label=f'Median: {median_val:,.1f}')

ax.set_title('Distribution of Values', fontweight='bold')
ax.set_xlabel('Value')
ax.set_ylabel('Frequency')
ax.legend()
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)

plt.tight_layout()
plt.savefig('histogram.png', dpi=150, bbox_inches='tight')

热力图

python
fig, ax = plt.subplots(figsize=(10, 8))

# 将数据透视成热力图格式
pivot = df.pivot_table(index='row_dim', columns='col_dim', values='metric', aggfunc='sum')

sns.heatmap(pivot, annot=True, fmt=',.0f', cmap='YlOrRd',
            linewidths=0.5, ax=ax, cbar_kws={'label': 'Metric Value'})

ax.set_title('Metric by Row Dimension and Column Dimension', fontweight='bold')
ax.set_xlabel('Column Dimension')
ax.set_ylabel('Row Dimension')

plt.tight_layout()
plt.savefig('heatmap.png', dpi=150, bbox_inches='tight')

小多组图

python
categories = df['category'].unique()
n_cats = len(categories)
n_cols = min(3, n_cats)
n_rows = (n_cats + n_cols - 1) // n_cols

fig, axes = plt.subplots(n_rows, n_cols, figsize=(5*n_cols, 4*n_rows), sharex=True, sharey=True)
axes = axes.flatten() if n_cats > 1 else [axes]

for i, cat in enumerate(categories):
    ax = axes[i]
    subset = df[df['category'] == cat]
    ax.plot(subset['date'], subset['value'], color=PALETTE_CATEGORICAL[i % len(PALETTE_CATEGORICAL)])
    ax.set_title(cat, fontsize=12)
    ax.spines['top'].set_visible(False)
    ax.spines['right'].set_visible(False)

# 隐藏空白子图
for j in range(i+1, len(axes)):
    axes[j].set_visible(False)

fig.suptitle('Trends by Category', fontsize=14, fontweight='bold', y=1.02)
plt.tight_layout()
plt.savefig('small_multiples.png', dpi=150, bbox_inches='tight')

数字格式化辅助函数

python
def format_number(val, fmt='number'):
    """为图表标签格式化数字:'number'、'currency' 或 'percent'。"""
    if fmt == 'percent':
        return f'{val:.1f}%'
    prefix = '$' if fmt == 'currency' else ''
    if abs(val) >= 1e9: return f'{prefix}{val/1e9:.1f}B'
    if abs(val) >= 1e6: return f'{prefix}{val/1e6:.1f}M'
    if abs(val) >= 1e3: return f'{prefix}{val/1e3:.1f}K'
    return f'{prefix}{val:,.0f}'

# 与坐标轴格式化器配合使用
ax.yaxis.set_major_formatter(mticker.FuncFormatter(lambda x, p: format_number(x, 'currency')))

使用 Plotly 创建交互式图表

python
import plotly.express as px
import plotly.graph_objects as go

# 简单交互式折线图
fig = px.line(df, x='date', y='value', color='category',
              title='Interactive Metric Trend',
              labels={'value': 'Metric Value', 'date': 'Date'})
fig.update_layout(hovermode='x unified')
fig.write_html('interactive_chart.html')
fig.show()

# 带悬停数据的交互式散点图
fig = px.scatter(df, x='metric_a', y='metric_b', color='category',
                 size='size_metric', hover_data=['name', 'detail_field'],
                 title='Correlation Analysis')
fig.show()

设计原则

设计基础部分涵盖色彩理论、数据墨水比、标注和无障碍性规则。以下补充图表特定指导,不在那些文件中。

  • 突出故事:用亮色强调关键洞察;其余元素用灰色
  • 标题陈述洞察:如“收入同比增长 23%”,而非“月度收入”。副标题添加日期范围、筛选条件和数据来源
  • 坐标轴标签:绝不旋转 90 度。缩短或换行。仅对关键数据点添加数据标签,而非每个柱形
  • 有意义的排序:按数值排序(而非字母顺序),除非存在自然顺序(月份、阶段)
  • 纵横比:时间序列宽大于高(3:1 至 2:1);比较图表更接近方形
  • 柱状图从零开始:始终如此。当范围重要时,折线图可使用非零基线
  • 多面板一致比例:比较多个图表时使用相同的坐标轴范围
  • 展示不确定性:当数据不确定时,添加误差线、置信区间或范围

无障碍性

设计基础涵盖颜色独立性和对比度规则。Python 特定补充:

  • 使用 sns.color_palette("colorblind") 作为备选的色盲安全调色板
  • 在颜色之外添加图案填充(matplotlib 的 hatch)或不同线型
  • 包含描述图表关键发现的替代文本;提供数据表备选
  • 测试:图表在黑白下是否可用?标准缩放比例下文字是否可读?

分享前检查: 系列在不依赖颜色的情况下可区分,标题陈述见解,坐标轴标注单位,图例清晰,注明数据源。