Data Analytics report

ChildSafe-RAG 研究大纲与当前进度

面向导师讨论的完整研究蓝图、证据边界、当前状态与下一步计划。

核心判断:框架与工具已成形,但论文级实证证据尚未开始

当前项目已经完成一套可审查的研究基础设施:统一taxonomy、Protégé ontology、年龄×风险rubric cards、360个问题草稿、检索/判分/统计代码和预注册设计。与此同时,正式模型回答、人工金标准和judge对照结果仍为零,因此目前能汇报的是研究设计与资产构建进展,模型安全性结论还没有。

本研究旨在将儿童保护、发展心理学、教育学与社会工作等社会科学知识转化为可执行、可更新的生成式人工智能安全框架。我们整合现有儿童安全研究,建立年龄分层的风险分类体系和评分标准,并开发 ChildSafe-RAG:根据儿童年龄和具体风险检索专属 rubric,由大语言模型逐项评价回答的安全性、适龄性、边界保护、情绪支持、现实指导和有效帮助程度。最终,我们将以人工标注为金标准验证该评估方法,并比较普通模型与加入儿童保护机制后的输出,检验其能否减少严重伤害,同时避免不必要的拒绝。

和陈老师讨论提出的更大贡献——从Social Science evidence构建可更新的AI protection framework——方向成立,但现有来源主要是YAIR等AI safety benchmark的整合,还缺少系统的社会科学检索、证据抽取和专家内容效度验证。另一个关键边界是:当前实验比较no-ageexplicit-age,尚未实现真正的child-protection plugin。

已登记来源文件

8Source: Repository status summary queryTable: summary

当前冻结的论文和公开 taxonomy 来源文件;不等同于系统综述已完成。

当前冻结的论文和公开 taxonomy 来源文件;不等同于系统综述已完成。Source: Repository status summary queryTable: summary

Loads headline counts from the reviewed repository-audit snapshot.

风险族

10Source: Repository status summary queryTable: summary

统一 taxonomy 的高层风险族数量,另含50个机制。

风险机制 50Source: Repository status summary queryTable: summary
统一 taxonomy 的高层风险族数量,另含50个机制。Source: Repository status summary queryTable: summary

Loads headline counts from the reviewed repository-audit snapshot.

Rubric cards(草稿)

30Source: Repository status summary queryTable: summary

3个年龄层 × 10个风险族;当前全部等待专家审核。

3个年龄层 × 10个风险族;当前全部等待专家审核。Source: Repository status summary queryTable: summary

Loads headline counts from the reviewed repository-audit snapshot.

问题(草稿)

360Source: Repository status summary queryTable: summary

英文 reality-grounded 合成/重写问题,不是真实儿童原话。

英文 reality-grounded 合成/重写问题,不是真实儿童原话。Source: Repository status summary queryTable: summary

Loads headline counts from the reviewed repository-audit snapshot.

正式模型回答

0Source: Repository status summary queryTable: summary

正式研究要求720份配对回答;目前只有一次工程 smoke test。

正式研究要求720份配对回答;目前只有一次工程 smoke test。Source: Repository status summary queryTable: summary

Loads headline counts from the reviewed repository-audit snapshot.

人工金标准

0Source: Repository status summary queryTable: summary

计划由三位成人标注者评分并形成300份 adjudicated gold。

计划由三位成人标注者评分并形成300份 adjudicated gold。Source: Repository status summary queryTable: summary

Loads headline counts from the reviewed repository-audit snapshot.

研究定位:从“AI for Social Science”转向“Social Science for AI”

问题陈述

通用LLM safety taxonomy通常不充分表达儿童风险的年龄差异、发展阶段、依赖关系、同意与现实支持需求;而通用LLM judge又容易把安全等同于拒绝,缺乏透明、可复核的维度。研究需要建立一条从社会科学证据到可执行AI规范的链路。

建议的一句话目标

构建并验证一个可追溯、可更新、由专家审核的儿童保护框架,将社会科学证据转化为年龄与风险专属rubric,并检验其作为LLM judge与生成时protection layer的有效性。

研究对象的边界

本研究评估7–17岁用户的英文文本交互,不招募儿童,不使用真实私人对话,不生成图像或露骨虐待材料。问题集当前应称为reality-grounded synthetic/rewritten prompts,不能称为真实儿童问题。

贡献应分成四层,并分别提供证据

  1. 理论贡献:提出Social Science for AI的child-protection治理路径。
  2. 知识工程贡献:可追溯、版本化、可增量更新的taxonomy、crosswalk和rubric construction pipeline。
  3. 评估贡献:年龄与风险条件化的Rubric-RAG judge,并以人工金标准验证。
  4. 干预贡献:把检索到的保护原则用于生成时约束,并证明其降低严重失败且不过度拒绝。

目前第2层已有草稿资产,第3层已有软件与实验设计;第1层需要更扎实的社会科学理论基础;第4层仍是概念。论文摘要和汇报必须按这一成熟度表述。

当前框架的内容:3个年龄层、10个风险族、6个评价维度

年龄层

7–910–1213–17。年龄差异优先写入rubric判定规则;只有新机制确实无法归入现有类别时,才扩展taxonomy。

十个统一风险族

  1. 性、亲密关系、同意与边界;2. 自伤、危机与心理健康;3. 情感依赖与AI拟人化;4. Grooming、掠夺、不当影响与操纵;5. 隐私、监控、身份与冒充;6. 错误信息与健康等高风险建议;7. 暴力、威胁、霸凌与语言伤害;8. 偏见、歧视与身份伤害;9. 欺骗、学术不端、网络滥用及其他用户滥用;10. 发展、社交与学习伤害。

三种不同任务

  • Category mapping:识别prompt或response涉及哪些风险机制,可视为多标签分类。
  • Response evaluation:按六个维度评价回答是否安全、适龄且有帮助,是本研究的核心judge任务。
  • Protection intervention:在生成时改变模型行为,是独立的因果对照任务。

三者有关联,但标签准确率不能替代回答质量,judge表现也不能自动证明plugin有效。

现有问题集在设计上平衡,但真实性仍需人类验证

三个年龄层数量完全平衡,且每层各包含60个风险问题和60个主题匹配的安全hard negatives。这支持比较危险漏判与benign过度拒绝,但数量平衡不能证明语言自然性、儿童真实性或生态效度。

问题集按年龄层与类型的分布
问题集按年龄层与类型的分布 data
年龄层问题数问题类型该年龄层总数
10-1260benign120
10-1260target120
13-1760benign120
13-1760target120
7-960benign120
7-960target120
问题级预期风险严重度分布
问题级预期风险严重度分布 data
严重度问题数占比
0 · benign/no expected harm18050%
1 · low3610%
2 · moderate8423.3%
3 · high6016.7%

完整方法链:证据、规范、评价与干预必须闭环

社会科学文献检索可定位证据抽取taxonomy与冲突整合年龄/风险rubric cards专家内容验证问题与模型回答数据集三人独立人工评分Rubric-RAG judge验证protection-layer对照实验错误分析与框架更新

文献到框架的自动化pipeline

  • Retrieval agent检索与去重;screening agent执行纳入/排除规则。
  • Evidence locator记录paper、page、section和原文位置。
  • Extraction agent输出定义、机制、年龄范围、上下文、保护原则和证据强度。
  • Synthesis agent聚类、映射已有taxonomy并显式记录冲突。
  • Human expert审核新增、合并、拆分与拒绝决定。
  • Version manager生成diff、来源追踪、发布日期和回归测试。

自动化负责降低检索与整理成本;最终分类和规范有效性仍由人类负责。

完整研究计划:五个相互衔接但不可混淆的研究模块

Source: Research module map queryTable: study_map

Loads the advisor-aligned research program and current maturity of each module.

完整研究计划:五个相互衔接但不可混淆的研究模块
模块核心问题方法现状
Study A · Framework constructionCan Social Science evidence be converted into a traceable, age-conditioned child-protection framework?Systematic search, evidence extraction, taxonomy synthesis, expert content validation and versioningInitial AI-benchmark synthesis exists; systematic Social Science pipeline is pending
Study B · Judge validationDoes age/risk Rubric-RAG agree with humans better than generic or long-context judge prompts?Global, Full-taxonomy, Example-RAG, Rubric-RAG and Oracle on an isolated held-out setSoftware and preregistration exist; human gold and model judgments are pending
Study C · Model auditHow much does explicit age information change model safety and developmental appropriateness?Paired no-age versus explicit-age responses across age bands and risk familiesPlanned in the current preregistration; formal responses are pending
Study D · Protection interventionDoes a retrieved child-protection layer improve outputs without increasing benign over-refusal?Paired baseline, age-only and protection-layer generation evaluated by humans and the validated judgeProposed in the advisor discussion; intervention specification and implementation are pending
Study E · UpdateabilityCan new evidence update the framework while preserving provenance and backward compatibility?Incremental evidence ingestion, change log, expert approval and regression testsVersioned files exist; automated extraction and update evaluation are pending

评价设计:先验证尺子,再用尺子比较模型

六个回答维度

  1. 内容安全与不协助伤害;2. 年龄及发展适宜性;3. 边界、同意与隐私;4. 情绪支持、自主性与避免依赖;5. 建设性指导、现实支持与必要升级;6. 真实性、有效帮助与避免过度拒绝。

每维1–5分,可标不适用;同时记录severity、hard fail、证据和pass / borderline / fail。总体平均分仅作描述,不能掩盖任何安全维度的低分或hard fail。

Study B:Judge validation

在完全隔离的180份held-out人工金标准上比较Global、Full-taxonomy、Example-RAG、Rubric-RAG和Oracle。主要指标为维度级quadratic weighted κ、MAE、exact/within-one accuracy,以及fail recall、false-negative rate和benign false-positive rate;20%样本重复三次评估稳定性。

Study C/D:生成效果

当前预注册条件是baseline no-age与explicit-age。若要回答导师提出的plugin effectiveness,必须新增第三个protection-layer条件,并明确它在生成前检索什么、如何注入、是否允许拒绝,以及怎样防止test rubric泄漏到generation prompt。

当前进度:资产构建领先于验证,且存在一处必须先修复的数据版本问题

Dashboard展示的新版360题全部唯一,但CLI默认使用的旧问题库只有134个唯一文本;两套文本只重合2条。直接运行正式API会得到与dashboard不一致的数据,因此必须先统一canonical inputs。

当前工作流状态与进入下一阶段的门槛

Source: Research progress audit queryTable: progress

Loads verified workstream states, evidence and next gates.

当前工作流状态与进入下一阶段的门槛
工作流当前状态可核查证据下一门槛
Age/risk rubric cardsDraft complete30 cards; all marked draft_pending_expert_reviewExpert review, revision log, and inter-rater calibration
Canonical experiment inputsBlocked by version mismatchCLI legacy bank has 134 unique prompts; overlap with dashboard bank is 2Make the v1 questions/cards the only canonical inputs and regenerate paired samples
Human gold standardNot started0 of 300 planned responses adjudicatedEthics determination, expert availability, annotator training, calibration and held-out labeling
Judge comparison and statisticsCode implemented; no evidenceRetrieval, judgment schema, agreement and bootstrap modules exist; no judgment filesValidate retrieval, then run five judge conditions against human gold
Protection-layer interventionConcept onlyCurrent preregistration compares no-age with explicit-age, not baseline with a pluginDefine the generation-time intervention and add a third, paired experimental condition
Question bankDraft complete360 unique English prompts; 180 target and 180 benignDual human review and ecological-validity check; no claim of real child utterances
Source corpusDraft corpus frozen8 source artifacts; 71 crosswalk labels; 91 YAIR leavesRun a documented Social Science search and evidence-screening protocol
Target-model responsesNot startedOne cheap-model smoke test exists; 0 formal responses from the planned 720Run a small engineering pilot after asset freeze, then collect the formal paired set
Unified taxonomyDraft complete10 families and 50 mechanisms; OWL/TTL export availableExpert content validation and traceability to primary social-science evidence

主要限制与审稿风险

  • 证据基础风险:AI benchmark crosswalk不等同于系统性Social Science literature synthesis。
  • 构念效度风险:10个风险族是工程整合结果,尚无专家内容效度证据。
  • 生态效度风险:360题为合成/重写,不是真实儿童话语,且真实情境可能更长、更含混。
  • 共同方法偏差:用LLM生成、再用LLM评分可能放大模型家族偏好,必须依靠盲法人工金标准和跨厂商judge。
  • 干预混淆:年龄提示不是child-protection plugin;不能用RQ2结果替代plugin效果。
  • 数据泄漏风险:新版cards中的示例引用问题ID,而split尚未冻结;正式实验前必须让rubric examples、calibration pool、held-out test与generation intervention彻底隔离。
  • 统计实现差距:已有核心指标模块,但预注册中的完整交互模型、20%三次稳定性和全部benign over-refusal分析尚未以真实结果验证。
  • 伦理风险:正式招募标注者前需要机构伦理/豁免确认、内容预警、退出权和敏感输出隔离。
  • 可更新性尚未验证:已有版本号和ontology,但尚无新文献进入后的增量更新实验。

论文组织:推荐以框架构建和judge validation为主线

我们要做的事systematic review、自动multi-agent extraction、全新benchmark、judge validation和plugin effectiveness全部完成。但是当前论文聚焦evidence-grounded framework + Rubric-RAG judge validation + age-cue audit;protection layer可作为清晰定义后的扩展实验或第二篇工作。

建议的论文结构

Source: Paper outline queryTable: paper_outline

Loads the proposed paper structure and the argumentative job of each section.

建议的论文结构
章节该章节要完成的论证
1. IntroductionFrame Social Science for AI, the child-specific governance gap and the central research questions.
2. Related workSeparate child-safety taxonomies, child-facing benchmarks, LLM-as-a-Judge, evidence synthesis and AI governance.
3. Evidence-to-framework pipelineSpecify retrieval, screening, evidence extraction, traceability, conflict resolution, expert review and update mechanisms.
4. Child protection frameworkPresent the 10 families, mechanisms, age bands, inclusion/exclusion rules and provenance crosswalk.
5. Rubric constructionExplain the six dimensions, age/risk cards, anchors, hard fails, verdict logic and content validation.
6. Benchmark and human goldDescribe question provenance, target/benign controls, paired conditions, sampling, blinding, annotation and ethics.
7. Judge-validation experimentCompare five judging methods on held-out human labels, including retrieval and repeatability tests.
8. Model audit / interventionAnalyze age-cue effects; include a protection-layer comparison only after the intervention is formally specified.
9. Results and error analysisReport dimension-level agreement, failure rates, over-refusal, subgroup effects and candidate taxonomy extensions.
10. Limitations, ethics and updateabilityBound synthetic-question validity, judge bias, source coverage, annotator welfare and framework drift.
11. ConclusionState the validated contribution without claiming expert review or plugin effectiveness before evidence exists.

下一步:先统一资产与补足证据链,再运行正式实验

完成所有的e2e pipeline,但是需要薛老师的input。

按依赖关系排序的下一步

Source: Next-action dependency queryTable: next_actions

Loads prioritized next actions and their exit criteria.

按依赖关系排序的下一步
优先级行动交付物完成标准
1Source: Next-action dependency queryTable: next_actionsFreeze one canonical asset setv1 cards + v1 questions + regenerated 720 paired Sample records; legacy files explicitly retiredAll commands and dashboards resolve to the same IDs and hashes
2Source: Next-action dependency queryTable: next_actionsWrite the Social Science evidence protocolSearch databases/queries, eligibility criteria, evidence schema, agent roles, audit log and update policyAnother researcher can reproduce one end-to-end evidence extraction
3Source: Next-action dependency queryTable: next_actionsRun expert and human asset reviewReviewed cards, reviewed prompt bank, disagreement log and revised versionNo asset is labeled expert-reviewed without recorded reviewer decisions
4Source: Next-action dependency queryTable: next_actionsRun a low-cost engineering pilot60 base prompts × 2 age conditions, retrieval diagnostics and three judge conditionsAPI/JSON success ≥98% and rubric Recall@3 ≥0.90
5Source: Next-action dependency queryTable: next_actionsLock the paper-level scope with advisorsDecision between current judge-validation paper and expanded protection-layer paperResearch questions, intervention conditions and contribution claims match the implemented experiment
6Source: Next-action dependency queryTable: next_actionsCollect formal evidence720 target responses, 300 triple-rated human labels, five judge conditions and paired statisticsAll preregistered quality gates and provenance checks are reported

陈老师讨论的计划

  • 第1周:统一canonical assets、冻结SHA与版本;完成Social Science检索和evidence ledger协议。
  • 第2–3周:补充儿童发展、社会工作、心理学和教育领域证据;双人编码并形成可追溯候选更新。
  • 第4周:专家审核taxonomy/cards,双人审核360题,冻结calibration/test前的最终资产。
  • 第5周:运行60个基础问题的低成本paired pilot,检查API、检索、judge JSON和标注指南。
  • 第6周:通过门槛后收集720份正式回答并生成分层标注包。
  • 第6–7周:三人独立标注、agreement检查、一次允许的anchor修订和adjudication。
  • 第8周:五种judge比较、年龄效应、错误分析、Holm校正和论文结果表。

时间表从伦理确认和专家可用之日开始;如果这两个前置条件未满足,只执行到技术pilot,不进入正式人工实验。

和薛老师要锁定的五个问题

  1. 第一篇论文的主贡献是framework construction、Rubric-RAG judge,还是generation-time protection layer?
  2. 是否要求系统综述级Social Science search;哪些学科数据库和理论传统必须覆盖?
  3. 360个合成问题是否足够,还是需要一个获许可、去标识的真实youth-language补充集?这块需要socialscenice的expert
  4. 哪些专家可以审核taxonomy/cards,怎样记录content-validity decisions?
  5. 陈老师:plugin是检索规则提示、独立middleware、response rewrite,还是模型微调?不同定义对应不同实验。

在这五点锁定前,可以完成技术pilot,但不宜启动昂贵的正式API采样和300份三人标注。

Sources

  1. Repository status summary queryreports/research_outline_progress_evidence.sqlite · sqlite · 2026-07-23T03:17:55Z

    Loads headline counts from the reviewed repository-audit snapshot.

    SQL query
    SELECT * FROM summary
  2. Question balance queryreports/research_outline_progress_evidence.sqlite · sqlite · 2026-07-23T03:17:55Z

    Loads the reviewed age-band and target/benign distribution.

    SQL query
    SELECT age_band, question_type, count, denominator_per_age FROM question_balance ORDER BY age_band, question_type
  3. Question severity queryreports/research_outline_progress_evidence.sqlite · sqlite · 2026-07-23T03:17:55Z

    Loads prompt-level expected-risk severity counts and shares.

    SQL query
    SELECT severity, label, count, share FROM severity_distribution ORDER BY CAST(severity AS INTEGER)
  4. Research progress audit queryreports/research_outline_progress_evidence.sqlite · sqlite · 2026-07-23T03:17:55Z

    Loads verified workstream states, evidence and next gates.

    SQL query
    SELECT workstream, state, evidence, next_gate FROM progress ORDER BY workstream
  5. Research module map queryreports/research_outline_progress_evidence.sqlite · sqlite · 2026-07-23T03:17:55Z

    Loads the advisor-aligned research program and current maturity of each module.

    SQL query
    SELECT study, question, method, status FROM study_map ORDER BY study
  6. Paper outline queryreports/research_outline_progress_evidence.sqlite · sqlite · 2026-07-23T03:17:55Z

    Loads the proposed paper structure and the argumentative job of each section.

    SQL query
    SELECT section, job FROM paper_outline ORDER BY CAST(SUBSTR(section, 1, INSTR(section, '.') - 1) AS INTEGER)
  7. Next-action dependency queryreports/research_outline_progress_evidence.sqlite · sqlite · 2026-07-23T03:17:55Z

    Loads prioritized next actions and their exit criteria.

    SQL query
    SELECT priority, action, deliverable, exit_criterion FROM next_actions ORDER BY priority
  8. ChildSafe-RAG repository audit snapshotreports/research_outline_progress_evidence.json
  9. User-provided advisor research discussion draft
  10. ChildSafe-RAG v1 preregistrationdocs/preregistration.md
  11. Unified child-safety taxonomy v1data/taxonomy/unified_taxonomy_v1.json
  12. Taxonomy source manifest v1data/taxonomy/taxonomy_source_manifest_v1.csv
  13. Age/risk rubric cards v1data/rubrics/rubric_cards_v1.jsonl
  14. Annotated question bank v1data/prompts/annotated_questions_v1.jsonl
  15. Human annotation guidedocs/annotation_guide.md
  16. Ethics and annotator welfare protocoldocs/ethics_protocol.md
  17. ChildSafe-RAG implementation and workflowREADME.md