科研技能库/实验审计
实验设计
未发现用户侧风险

实验审计

在声明实验结果前审计实验完整性。使用跨模型审查(GPT-5.5)检查是否存在伪造基准真实值、分数归一化欺诈、phantom结果和范围不足等问题。当用户说出“审计实验”、“检查实验完整性”、“审计结果”、“实验诚实度”时使用,或在实验完成准备撰写声明前使用。

文件预览

1 个文件
SKILL.md
9.5 KB · 可预览
---
name: experiment-audit
description: "Audit experiment integrity before claiming results. Uses cross-model review (GPT-5.5) to check for fake ground truth, score normalization fraud, phantom results, and insufficient scope. Use when user says \"审计实验\", \"check experiment integrity\", \"audit results\", \"实验诚实度\", or after experiments complete before writing claims."
argument-hint: [experiment-dir-or-results-path]
allowed-tools: Bash(*), Read, Write, Edit, Grep, Glob
---

# Experiment Audit: Cross-Model Integrity Verification

Audit experiment integrity for: **$ARGUMENTS**

## Why This Exists

LLM agents can produce fraudulent experimental results through:
1. **Fake ground truth** — creating synthetic "reference" from model outputs, then reporting high agreement as performance
2. **Score normalization** — dividing metrics by the model's own max to get 0.99+
3. **Phantom results** — claiming numbers from files that don't exist or functions never called
4. **Insufficient scope** — reporting 2-scene pilots as "comprehensive evaluation"

These are NOT intentional deception — they are failure modes of optimizing agents that lack integrity constraints. This skill adds that constraint.

## Core Principle

**The executor (Claude) collects file paths. The reviewer (GPT-5.5) reads code and judges integrity. The executor does NOT participate in integrity judgment.**

This follows `shared-references/reviewer-independence.md` and `shared-references/experiment-integrity.md`.

## Constants

- **REVIEWER_BACKEND = `codex`** — Default: Codex reviewer agent (`spawn_agent`, xhigh). Override with `— reviewer: oracle-pro` for GPT-5.5 Pro via Oracle MCP. See `shared-references/reviewer-routing.md`.

## Workflow

### Step 1: Collect Artifacts (Executor — Claude)

Locate and list these files WITHOUT reading or summarizing their content:

```
Scan project directory for:
1. Evaluation scripts:    *eval*.py, *metric*.py, *test*.py, *benchmark*.py
2. Result files:          *.json, *.csv in results/, outputs/, logs/
3. Ground truth paths:    look in eval scripts for data loading (dataset paths, GT references)
4. Experiment tracker:    EXPERIMENT_TRACKER.md, EXPERIMENT_LOG.md
5. Paper claims:          NARRATIVE_REPORT.md, paper/sections/*.tex, PAPER_PLAN.md
6. Config files:          *.yaml, *.toml, *.json configs with metric definitions
```

**DO NOT summarize, interpret, or explain any file content.** Only collect paths.

### Step 2: Send to Reviewer (GPT-5.5 via Codex MCP)

Pass ONLY file paths and the audit checklist to the reviewer. The reviewer reads everything directly.

```text
spawn_agent:
  model: gpt-5.5
  reasoning_effort: xhigh
  message: |
    You are an experiment integrity auditor. Read ALL files listed below
    and check for the following fraud patterns.

    Files to read:
    - Evaluation scripts: [list paths]
    - Result files: [list paths]
    - Experiment tracker: [list paths]
    - Paper claims: [list paths]
    - Config files: [list paths]

    ## Audit Checklist

    ### A. Ground Truth Provenance
    For each evaluation script:
    1. Where does "ground truth" / "reference" / "target" come from?
    2. Is it loaded from the DATASET, or generated/derived from MODEL OUTPUTS?
    3. If derived: is it explicitly labeled as proxy evaluation?
    4. Are official eval scripts used when available for this benchmark?
    FAIL if: GT is derived from model outputs without explicit proxy labeling.

    ### B. Score Normalization
    For each metric computation:
    1. Is any metric divided by max/min/mean of the model's OWN output?
    2. Are raw scores reported alongside any normalized scores?
    3. Are any scores suspiciously close to 1.0 or 100%?
    FAIL if: Normalization denominator comes from prediction statistics.

    ### C. Result File Existence
    For each claim in the paper/narrative:
    1. Does the referenced result file actually exist?
    2. Does the claimed metric key exist in that file?
    3. Does the claimed NUMBER match what's in the file?
    4. Is the experiment tracker status DONE (not TODO/IN_PROGRESS)?
    FAIL if: Claimed results reference nonexistent files or mismatched numbers.

    ### D. Dead Code Detection
    For each metric function defined in eval scripts:
    1. Is it actually CALLED in any evaluation pipeline?
    2. Does its output appear in any result file?
    WARN if: Metric functions exist but are never called.

    ### E. Scope Assessment
    1. How many scenes/datasets/configurations were actually tested?
    2. How many seeds/runs per configuration?
    3. Does the paper use words like "comprehensive", "extensive", "robust"?
    4. Is the actual scope sufficient for those claims?
    WARN if: Scope language exceeds actual evidence.

    ### F. Evaluation Type Classification
    Classify each evaluation as:
    - real_gt: uses dataset-provided ground truth
    - synthetic_proxy: uses model-generated reference
    - self_supervised_proxy: no GT by design
    - simulation_only: simulated environment
    - human_eval: human judges

    ## Output Format

    For each check (A-F), report:
    - Status: PASS | WARN | FAIL
    - Evidence: exact file:line references
    - Details: what specifically was found

    Overall verdict: PASS | WARN | FAIL
    
    Be thorough. Read every eval script line by line.
```

### Step 3: Parse and Write Report (Executor — Claude)

Parse the reviewer's response and write `EXPERIMENT_AUDIT.md`:

```markdown
# Experiment Audit Report

**Date**: [today]
**Auditor**: GPT-5.5 xhigh (cross-model, read-only)
**Project**: [project name]

## Overall Verdict: [PASS | WARN | FAIL]

## Integrity Status: [pass | warn | fail]

## Checks

### A. Ground Truth Provenance: [PASS|WARN|FAIL]
[details + file:line evidence]

### B. Score Normalization: [PASS|WARN|FAIL]
[details]

### C. Result File Existence: [PASS|WARN|FAIL]
[details]

### D. Dead Code Detection: [PASS|WARN|FAIL]
[details]

### E. Scope Assessment: [PASS|WARN|FAIL]
[details]

### F. Evaluation Type: [real_gt | synthetic_proxy | ...]
[classification + evidence]

## Action Items
- [specific fixes if WARN or FAIL]

## Claim Impact
- Claim 1: [supported | needs qualifier | unsupported]
- Claim 2: ...
```

Also write `EXPERIMENT_AUDIT.json` for machine consumption:

```json
{
  "date": "2026-04-10",
  "auditor": "gpt-5.5-xhigh",
  "overall_verdict": "warn",
  "integrity_status": "warn",
  "checks": {
    "gt_provenance": {"status": "pass", "details": "..."},
    "score_normalization": {"status": "warn", "details": "..."},
    "result_existence": {"status": "pass", "details": "..."},
    "dead_code": {"status": "pass", "details": "..."},
    "scope": {"status": "warn", "details": "..."},
    "eval_type": "real_gt"
  },
  "claims": [
    {"id": "C1", "impact": "supported"},
    {"id": "C2", "impact": "needs_qualifier"}
  ]
}
```

### Step 4: Print Summary

```
🔬 Experiment Audit Complete

  GT Provenance:      ✅ PASS — real dataset GT used
  Score Normalization: ⚠️ WARN — boundary metric uses self-reference
  Result Existence:    ✅ PASS — all files exist, numbers match
  Dead Code:           ✅ PASS — all metric functions called
  Scope:               ⚠️ WARN — 2 scenes, paper says "comprehensive"

  Overall: ⚠️ WARN
  
  See EXPERIMENT_AUDIT.md for details.
```

## Integration with Other Skills

### Automatic in /research-pipeline (advisory, never blocks)

When integrated into the pipeline, this skill runs automatically after `/experiment-bridge` and before `/auto-review-loop`:

```
/experiment-bridge → results ready
    ↓
/experiment-audit (automatic, advisory)
    ├── PASS  → continue normally
    ├── WARN  → print ⚠️ warning, continue, tag claims as [INTEGRITY: WARN]
    └── FAIL  → print 🔴 alert, continue, tag claims as [INTEGRITY CONCERN]
    ↓
/auto-review-loop → proceeds with integrity tags visible to reviewer
```

**Never blocks the pipeline.** Even on FAIL, the pipeline continues — but claims carry visible integrity tags.

### Read by /result-to-claim (if exists)

```
if EXPERIMENT_AUDIT.json exists:
    read integrity_status
    attach to verdict: {claim_supported: "yes", integrity_status: "warn"}
    if integrity_status == "fail":
        downgrade verdict display: "yes [INTEGRITY CONCERN]"
else:
    verdict as normal, integrity_status = "unavailable"
    mark as "provisional — no integrity audit"
```

### Read by /paper-write (if exists)

```
if EXPERIMENT_AUDIT.json exists AND integrity_status == "fail":
    add footnote to affected claims: "Note: integrity audit flagged concerns with this evaluation"
```

## Key Rules

- **Reviewer independence**: executor collects paths, reviewer judges. Period.
- **Never block**: warn loudly, never halt the pipeline.
- **File-as-switch**: no EXPERIMENT_AUDIT.md = skill was never run = zero impact on existing behavior.
- **Cross-model**: the reviewer MUST be a different model family from the executor.
- **Honest about limits**: the audit catches common patterns, not all possible fraud. It is a safety net, not a guarantee.

## Acknowledgements

Motivated by community-reported integrity issues (#57, #131) where executor agents created fake ground truth and self-normalized scores.

## Review Tracing

After each reviewer agent call, save the trace following `shared-references/review-tracing.md` (Policy C — forensic; never silently skip). Use `save_trace.sh` (resolved per the chain in `shared-references/integration-contract.md` §2) or write files directly to `.aris/traces/<skill>/<date>_run<NN>/`. Respect the `--- trace:` parameter (default: `full`).

SKILL.md

元数据
nameexperiment-audit
description在声明实验结果前审计实验完整性。使用跨模型审查(GPT-5.5)检查是否存在伪造基准真实值、分数归一化欺诈、phantom结果和范围不足等问题。当用户说出“审计实验”、“检查实验完整性”、“审计结果”、“实验诚实度”时使用,或在实验完成准备撰写声明前使用。
argument-hint["experiment-dir-or-results-path"]
allowed-toolsBash(*), Read, Write, Edit, Grep, Glob

实验审计:跨模型完整性验证

为以下内容审计实验完整性:$ARGUMENTS

为什么存在

LLM代理可能产生欺诈性实验结果,通过:

  1. 伪造基准真实值 — 从模型输出中创建合成“参考”,然后报告高一致性作为性能。
  2. 分数归一化 — 将指标除以模型自己的最大值,得到0.99+。
  3. Phantom结果 — 声称的数字来自不存在的文件或从未调用的函数。
  4. 范围不足 — 将2个场景的试点报告为“全面评估”。

这些并非有意欺骗——它们是缺乏完整性约束的优化代理的失败模式。本技能添加该约束。

核心原则

执行者(Claude)收集文件路径。审查者(GPT-5.5)阅读代码并判断完整性。执行者不参与完整性判断。

这遵循 shared-references/reviewer-independence.md 和 shared-references/experiment-integrity.md。

常量

  • REVIEWER_BACKEND = codex — 默认:Codex审查代理(spawn_agent,xhigh)。使用 — reviewer: oracle-pro 覆盖为 Oracle MCP 的 GPT-5.5 Pro。请参见 shared-references/reviewer-routing.md。

工作流程

步骤1:收集产物(执行者 — Claude)

查找并列出这些文件,但不要阅读或总结其内容:

text
扫描项目目录,查找:
1. 评估脚本:    *eval*.py, *metric*.py, *test*.py, *benchmark*.py
2. 结果文件:    results/、outputs/、logs/ 中的 *.json、*.csv
3. 基准真实值路径:    在评估脚本中查找数据加载(数据集路径、GT引用)
4. 实验跟踪器:    EXPERIMENT_TRACKER.md, EXPERIMENT_LOG.md
5. 论文声明:    NARRATIVE_REPORT.md, paper/sections/*.tex, PAPER_PLAN.md
6. 配置文件:    *.yaml, *.toml, *.json 配置,含指标定义

不要总结、解释或说明任何文件内容。 仅收集路径。

步骤2:发送给审查者(通过Codex MCP的GPT-5.5)

仅将文件路径和审计检查清单传递给审查者。审查者直接读取所有内容。

text
spawn_agent:
  model: gpt-5.5
  reasoning_effort: xhigh
  message: |
    您是一位实验完整性审计员。阅读下面列出的所有文件,并检查以下欺诈模式。

    要读取的文件:
    - 评估脚本:[列出路径]
    - 结果文件:[列出路径]
    - 实验跟踪器:[列出路径]
    - 论文声明:[列出路径]
    - 配置文件:[列出路径]

    ## 审计检查清单

    ### A. 基准真实值来源
    对于每个评估脚本:
    1. “基准真实值”/“参考”/“目标”来自哪里?
    2. 是从数据集加载,还是从模型输出生成/派生?
    3. 如果派生:是否明确标记为代理评估?
    4. 在适用时是否使用了该基准的官方评估脚本?
    失败条件:基准真实值派生自模型输出,且没有明确的代理标记。

    ### B. 分数归一化
    对于每个指标计算:
    1. 是否有指标除以模型自身输出的最大值/最小值/均值?
    2. 是否在报告任何归一化分数时也报告了原始分数?
    3. 是否有任何分数异常接近1.0或100%?
    失败条件:归一化的分母来自预测统计量。

    ### C. 结果文件存在性
    对于论文/叙述中的每个声明:
    1. 引用的结果文件是否实际存在?
    2. 声明的指标键是否存在于该文件?
    3. 声明的数字是否与文件中的一致?
    4. 实验跟踪器状态是否为DONE(而非TODO/IN_PROGRESS)?
    失败条件:声明的结果引用了不存在的文件或数字不匹配。

    ### D. 死代码检测
    对于评估脚本中定义的每个指标函数:
    1. 它是否在任何评估流水线中被实际调用?
    2. 其输出是否出现在任何结果文件中?
    警告条件:存在从未被调用的指标函数。

    ### E. 范围评估
    1. 实际测试了多少场景/数据集/配置?
    2. 每种配置的种子/运行次数?
    3. 论文是否使用了“全面”、“广泛”等词语?
    4. 实际范围是否足以支撑这些声明?
    警告条件:范围用语超出实际证据。

    ### F. 评估类型分类
    将每个评估分类为:
    - real_gt:使用数据集提供的基准真实值
    - synthetic_proxy:使用模型生成的参考
    - self_supervised_proxy:设计上没有基准真实值
    - simulation_only:模拟环境
    - human_eval:人类评审

    ## 输出格式

    对于每项检查(A-F),报告:
    - 状态:PASS | WARN | FAIL
    - 证据:精确的文件:行号引用
    - 详情:具体发现

    总体结论:PASS | WARN | FAIL
    
    请彻底检查。逐行阅读每个评估脚本。

步骤3:解析并撰写报告(执行者 — Claude)

解析审查者的响应,并写入 EXPERIMENT_AUDIT.md:

markdown
# 实验审计报告

**日期**:[今天]
**审计员**:GPT-5.5 xhigh(跨模型,只读)
**项目**:[项目名称]

## 总体结论:[PASS | WARN | FAIL]

## 完整性状态:[pass | warn | fail]

## 检查项

### A. 基准真实值来源:[PASS|WARN|FAIL]
[详情 + 文件:行号证据]

### B. 分数归一化:[PASS|WARN|FAIL]
[详情]

### C. 结果文件存在性:[PASS|WARN|FAIL]
[详情]

### D. 死代码检测:[PASS|WARN|FAIL]
[详情]

### E. 范围评估:[PASS|WARN|FAIL]
[详情]

### F. 评估类型:[real_gt | synthetic_proxy | ...]
[分类 + 证据]

## 行动项
- [若WARN或FAIL的具体修复]

## 声明影响
- 声明1:[supported | needs qualifier | unsupported]
- 声明2:...

同时写入 EXPERIMENT_AUDIT.json 供程序使用:

json
{
  "date": "2026-04-10",
  "auditor": "gpt-5.5-xhigh",
  "overall_verdict": "warn",
  "integrity_status": "warn",
  "checks": {
    "gt_provenance": {"status": "pass", "details": "..."},
    "score_normalization": {"status": "warn", "details": "..."},
    "result_existence": {"status": "pass", "details": "..."},
    "dead_code": {"status": "pass", "details": "..."},
    "scope": {"status": "warn", "details": "..."},
    "eval_type": "real_gt"
  },
  "claims": [
    {"id": "C1", "impact": "supported"},
    {"id": "C2", "impact": "needs_qualifier"}
  ]
}

步骤4:打印摘要

text
🔬 实验审计完成

  基准真实值来源:      ✅ PASS — 使用了真实数据集基准真实值
  分数归一化:          ⚠️ WARN — 边界度量使用了自引用
  结果文件存在性:      ✅ PASS — 所有文件存在,数字匹配
  死代码:              ✅ PASS — 所有指标函数都被调用
  范围:                ⚠️ WARN — 2个场景,论文声称“全面”

  总体:⚠️ WARN
  
  详见 EXPERIMENT_AUDIT.md。

与其他技能的集成

在 /research-pipeline 中自动执行(建议性,永不阻塞)

当集成到流水线时,本技能在 /experiment-bridge 之后、/auto-review-loop 之前自动运行:

text
/experiment-bridge → 结果就绪
    ↓
/experiment-audit(自动,建议性)
    ├── PASS  → 正常继续
    ├── WARN  → 打印 ⚠️ 警告,继续,将声明标记为 [INTEGRITY: WARN]
    └── FAIL  → 打印 🔴 警报,继续,将声明标记为 [INTEGRITY CONCERN]
    ↓
/auto-review-loop → 继续,完整性标签对审查者可见

永不阻塞流水线。 即使 FAIL,流水线也继续——但声明带有可见的完整性标签。

被 /result-to-claim 读取(若存在)

text
if EXPERIMENT_AUDIT.json 存在:
    读取 integrity_status
    附加到判决:{claim_supported: "yes", integrity_status: "warn"}
    if integrity_status == "fail":
        降级判决显示:"yes [INTEGRITY CONCERN]"
else:
    判决如常,integrity_status = "unavailable"
    标记为“暂定 — 无完整性审计”

被 /paper-write 读取(若存在)

text
if EXPERIMENT_AUDIT.json 存在 AND integrity_status == "fail":
    为受影响的声明添加脚注:“注意:完整性审计对此评估标识了顾虑”

关键规则

  • 审查者独立性:执行者收集路径,审查者判断。就这样。
  • 永不阻塞:大声警告,绝不停止流水线。
  • 文件作为开关:没有 EXPERIMENT_AUDIT.md = 技能从未运行 = 对现有行为零影响。
  • 跨模型:审查者必须是与执行者不同的模型系列。
  • 诚实地对待限制:审计能捕获常见模式,不是所有可能的欺诈。它是安全网,不是保证。

致谢

受社区报告的完整性问题(#57,#131)启发,当时执行代理创建了虚假基准真实值并自行归一化分数。

审查追踪

每次审查代理调用后,按照 shared-references/review-tracing.md 保存追踪(策略 C — 取证;绝不要静默跳过)。使用 save_trace.sh(按 shared-references/integration-contract.md §2 中的链解析)或将文件直接写入 .aris/traces/<skill>/<date>_run<NN>/。遵循 --- trace: 参数(默认:full)。