数据分析
未发现用户侧风险
图表数据提取器
从图表或图形的图像中提取像素级数据,并生成结构化数据表。当需要从图表图像中提取数据、从图形中转录数字、将图表数字化或将数据截图转换为表格时使用。生成带有提取值、置信度水平和重建图表源的结构化表格。建议使用 Claude Opus 4.7 或更高版本以获得可靠的图表数据提取。
文件预览
1 个文件
SKILL.md
4.1 KB · 可预览
--- name: chart-data-extractor description: "Extract pixel-level data from an image of a chart or graph and produce a structured data table. Use when asked to extract data from a chart image, transcribe numbers from a graph, digitise a chart, or turn a screenshot of data into a table. Produces a structured table with extracted values, confidence levels, and a reconstructed chart source. Best used with Claude Opus 4.7 or newer for reliable chart data extraction." --- # Chart Data Extractor Skill Extracts data from images of charts and graphs — bar charts, line charts, pie charts, scatter plots, and tables in images — producing a structured data table that can be used in spreadsheets or rebuilt in any charting tool. Built to leverage Opus 4.7 pixel-level image analysis capabilities. ## Required Inputs Ask the user for these if not provided: - **The chart image** (upload a screenshot or image file) - **Chart type** (if ambiguous — bar / line / pie / scatter / other) - **What matters most** (approximate trends / precise values / specific data points / categorisation) - **Known axis values** (optional — if the user knows the max/min values to anchor the extraction) ## Output Structure ### 1. Chart Identification | Attribute | Value | |---|---| | Chart type | [Bar / Line / Pie / Scatter / Area / Other] | | Chart title (if visible) | [Title text] | | X-axis label | [Label + unit] | | Y-axis label | [Label + unit] | | Number of series | N | | Legend categories | [List] | | Data period (if time-based) | [Start — End] | ### 2. Extracted Data Table | [X axis] | [Series 1] | [Series 2] | ... | |---|---|---|---| | [Value] | [Value] | [Value] | | ### 3. Confidence Levels For each data point or series, flag confidence: - **High confidence:** data points where the value is clearly readable against gridlines or labels - **Medium confidence:** data points where the value is interpolated between gridlines - **Low confidence:** data points where the value is ambiguous or overlaps with other elements Low-confidence points should be explicitly listed — not silently included in the main table. ### 4. Notable Observations Observations that the data itself reveals: - Peak value: [Value, when, in which series] - Lowest value: [Value, when, in which series] - Largest delta between series: [Details] - Any anomalies or outliers visible in the chart ### 5. Reconstructed Source CSV format for direct use: ```csv [x_axis],[series_1],[series_2] [value],[value],[value] ``` ### 6. Assumptions and Caveats - Grid resolution: [How precisely values could be read — e.g. "Y-axis has major gridlines every 10 units, minor every 2"] - Interpolation used: [Any values that required estimating between gridlines] - Unclear data: [Anything in the chart that could not be read reliably] - Axis scale: [Linear/logarithmic/etc — note if not obvious] ### 7. Follow-up Options Ask the user which of these they want: - Rebuild the chart in a specified format (Excel formula, Python matplotlib, D3, etc.) - Produce a narrative description of what the chart shows - Compare this data against another chart or source - Flag potentially misleading visual choices in the original (truncated axes, misleading scales, etc.) ## Quality Checks - [ ] Every extracted number specifies which series it belongs to - [ ] Confidence levels are explicit for ambiguous points - [ ] Low-confidence values are flagged separately, not silently included - [ ] Assumptions about axis scale and interpolation are stated - [ ] CSV output is clean and directly usable ## Example Trigger Phrases - "Extract the data from this chart" - "Transcribe the numbers in this graph" - "Turn this chart image into a spreadsheet" - "Digitise this chart so I can rebuild it" - "What are the exact values in this bar chart?" ## Why This Works Better on Opus 4.7 Earlier models struggled with pixel-level data transcription from charts, often hallucinating values or misreading gridline positions. Opus 4.7 uses a higher image resolution (2576px vs 1568px) with coordinates mapping 1:1 to pixels, making chart data extraction reliable for practical use.
SKILL.md
元数据
| name | 图表数据提取器 |
|---|---|
| description | 从图表或图形的图像中提取像素级数据,并生成结构化数据表。当需要从图表图像中提取数据、从图形中转录数字、将图表数字化或将数据截图转换为表格时使用。生成带有提取值、置信度水平和重建图表源的结构化表格。建议使用 Claude Opus 4.7 或更高版本以获得可靠的图表数据提取。 |
图表数据提取器技能
从图表和图形的图像中提取数据——包括条形图、折线图、饼图、散点图以及图像中的表格——生成可在电子表格中使用或可在任何图表工具中重建的结构化数据表。该技能基于 Opus 4.7 的像素级图像分析能力构建。
必需的输入
如果未提供,请向用户询问以下内容:
- 图表图像(上传截图或图像文件)
- 图表类型(如果模棱两可——条形图 / 折线图 / 饼图 / 散点图 / 其他)
- 关注重点(近似趋势 / 精确数值 / 特定数据点 / 分类)
- 已知坐标轴值(可选——如果用户知道最大/最小值以锚定提取)
输出结构
1. 图表识别
| 属性 | 值 |
|---|---|
| 图表类型 | [条形图 / 折线图 / 饼图 / 散点图 / 面积图 / 其他] |
| 图表标题(若可见) | [标题文本] |
| X轴标签 | [标签 + 单位] |
| Y轴标签 | [标签 + 单位] |
| 系列数量 | N |
| 图例类别 | [列表] |
| 数据周期(若为时间序列) | [起始 — 结束] |
2. 提取的数据表
| [X轴] | [系列1] | [系列2] | ... |
|---|---|---|---|
| [值] | [值] | [值] |
3. 置信度水平
为每个数据点或系列标记置信度:
- 高置信度: 数据点的值可根据网格线或标签清晰读取
- 中置信度: 数据点的值在网格线之间插值得到
- 低置信度: 数据点的值模糊或与其他元素重叠
低置信度点应明确列出——不应默默包含在主表中。
4. 显著观察
数据本身揭示的观察:
- 峰值: [值,时刻,哪个系列]
- 最低值: [值,时刻,哪个系列]
- 系列间最大差异: [详情]
- 图表中可见的任何异常或离群值
5. 重建源
可直接使用的 CSV 格式:
csv
[x轴],[系列1],[系列2]
[值],[值],[值]6. 假设与注意事项
- 网格分辨率: [值可读的精度——例如“Y轴每隔10个单位有主网格线,每隔2个单位有次网格线”]
- 使用的插值: [任何需要在网格线间估算的值]
- 不清晰数据: [图表中无法可靠读出的任何内容]
- 轴刻度: [线性/对数/等等——若不明显需注明]
7. 后续选项
询问用户需要哪些后续操作:
- 按指定格式重建图表(Excel 公式、Python matplotlib、D3 等)
- 生成图表内容的叙述性描述
- 将此数据与另一图表或来源进行比较
- 标记原始图表中可能具有误导性的视觉选择(截断轴、误导性刻度等)
质量检查
- 每个提取的数字都注明了所属系列
- 对模糊点明确标注了置信度
- 低置信度值单独标记,未默默包含
- 明确说明轴刻度及插值相关的假设
- CSV 输出清晰且可直接使用
示例触发短语
- “从这个图表中提取数据”
- “转录该图形中的数字”
- “将此图表图像转换为电子表格”
- “数字化该图表,以便我重建它”
- “这个条形图的精确值是多少?”
为何在 Opus 4.7 上效果更好
早期模型在从图表中转录像素级数据时表现挣扎,经常产生幻觉值或误读网格线位置。Opus 4.7 使用更高图像分辨率(2576px vs 1568px),坐标与像素一一映射,使图表数据提取在实际应用中变得可靠。