已登记来源文件
8Source: Repository status summary query
当前冻结的论文和公开 taxonomy 来源文件;不等同于系统综述已完成。
Loads headline counts from the reviewed repository-audit snapshot.
Data Analytics report
面向导师讨论的完整研究蓝图、证据边界、当前状态与下一步计划。
当前项目已经完成一套可审查的研究基础设施:统一taxonomy、Protégé ontology、年龄×风险rubric cards、360个问题草稿、检索/判分/统计代码和预注册设计。与此同时,正式模型回答、人工金标准和judge对照结果仍为零,因此目前能汇报的是研究设计与资产构建进展,模型安全性结论还没有。
本研究旨在将儿童保护、发展心理学、教育学与社会工作等社会科学知识转化为可执行、可更新的生成式人工智能安全框架。我们整合现有儿童安全研究,建立年龄分层的风险分类体系和评分标准,并开发 ChildSafe-RAG:根据儿童年龄和具体风险检索专属 rubric,由大语言模型逐项评价回答的安全性、适龄性、边界保护、情绪支持、现实指导和有效帮助程度。最终,我们将以人工标注为金标准验证该评估方法,并比较普通模型与加入儿童保护机制后的输出,检验其能否减少严重伤害,同时避免不必要的拒绝。
和陈老师讨论提出的更大贡献——从Social Science evidence构建可更新的AI protection framework——方向成立,但现有来源主要是YAIR等AI safety benchmark的整合,还缺少系统的社会科学检索、证据抽取和专家内容效度验证。另一个关键边界是:当前实验比较no-age与explicit-age,尚未实现真正的child-protection plugin。
已登记来源文件
8Source: Repository status summary query
当前冻结的论文和公开 taxonomy 来源文件;不等同于系统综述已完成。
Loads headline counts from the reviewed repository-audit snapshot.
风险族
10Source: Repository status summary query
统一 taxonomy 的高层风险族数量,另含50个机制。
Loads headline counts from the reviewed repository-audit snapshot.
Rubric cards(草稿)
30Source: Repository status summary query
3个年龄层 × 10个风险族;当前全部等待专家审核。
Loads headline counts from the reviewed repository-audit snapshot.
问题(草稿)
360Source: Repository status summary query
英文 reality-grounded 合成/重写问题,不是真实儿童原话。
Loads headline counts from the reviewed repository-audit snapshot.
正式模型回答
0Source: Repository status summary query
正式研究要求720份配对回答;目前只有一次工程 smoke test。
Loads headline counts from the reviewed repository-audit snapshot.
人工金标准
0Source: Repository status summary query
计划由三位成人标注者评分并形成300份 adjudicated gold。
Loads headline counts from the reviewed repository-audit snapshot.
通用LLM safety taxonomy通常不充分表达儿童风险的年龄差异、发展阶段、依赖关系、同意与现实支持需求;而通用LLM judge又容易把安全等同于拒绝,缺乏透明、可复核的维度。研究需要建立一条从社会科学证据到可执行AI规范的链路。
构建并验证一个可追溯、可更新、由专家审核的儿童保护框架,将社会科学证据转化为年龄与风险专属rubric,并检验其作为LLM judge与生成时protection layer的有效性。
本研究评估7–17岁用户的英文文本交互,不招募儿童,不使用真实私人对话,不生成图像或露骨虐待材料。问题集当前应称为reality-grounded synthetic/rewritten prompts,不能称为真实儿童问题。
Social Science for AI的child-protection治理路径。目前第2层已有草稿资产,第3层已有软件与实验设计;第1层需要更扎实的社会科学理论基础;第4层仍是概念。论文摘要和汇报必须按这一成熟度表述。
7–9、10–12、13–17。年龄差异优先写入rubric判定规则;只有新机制确实无法归入现有类别时,才扩展taxonomy。
三者有关联,但标签准确率不能替代回答质量,judge表现也不能自动证明plugin有效。
三个年龄层数量完全平衡,且每层各包含60个风险问题和60个主题匹配的安全hard negatives。这支持比较危险漏判与benign过度拒绝,但数量平衡不能证明语言自然性、儿童真实性或生态效度。
Loads the reviewed age-band and target/benign distribution.
| 年龄层 | 问题数 | 问题类型 | 该年龄层总数 |
|---|---|---|---|
| 10-12 | 60 | benign | 120 |
| 10-12 | 60 | target | 120 |
| 13-17 | 60 | benign | 120 |
| 13-17 | 60 | target | 120 |
| 7-9 | 60 | benign | 120 |
| 7-9 | 60 | target | 120 |
Loads prompt-level expected-risk severity counts and shares.
| 严重度 | 问题数 | 占比 |
|---|---|---|
| 0 · benign/no expected harm | 180 | 50% |
| 1 · low | 36 | 10% |
| 2 · moderate | 84 | 23.3% |
| 3 · high | 60 | 16.7% |
社会科学文献检索 → 可定位证据抽取 → taxonomy与冲突整合 → 年龄/风险rubric cards → 专家内容验证 → 问题与模型回答数据集 → 三人独立人工评分 → Rubric-RAG judge验证 → protection-layer对照实验 → 错误分析与框架更新。
自动化负责降低检索与整理成本;最终分类和规范有效性仍由人类负责。
Loads the advisor-aligned research program and current maturity of each module.
| 模块 | 核心问题 | 方法 | 现状 |
|---|---|---|---|
| Study A · Framework construction | Can Social Science evidence be converted into a traceable, age-conditioned child-protection framework? | Systematic search, evidence extraction, taxonomy synthesis, expert content validation and versioning | Initial AI-benchmark synthesis exists; systematic Social Science pipeline is pending |
| Study B · Judge validation | Does age/risk Rubric-RAG agree with humans better than generic or long-context judge prompts? | Global, Full-taxonomy, Example-RAG, Rubric-RAG and Oracle on an isolated held-out set | Software and preregistration exist; human gold and model judgments are pending |
| Study C · Model audit | How much does explicit age information change model safety and developmental appropriateness? | Paired no-age versus explicit-age responses across age bands and risk families | Planned in the current preregistration; formal responses are pending |
| Study D · Protection intervention | Does a retrieved child-protection layer improve outputs without increasing benign over-refusal? | Paired baseline, age-only and protection-layer generation evaluated by humans and the validated judge | Proposed in the advisor discussion; intervention specification and implementation are pending |
| Study E · Updateability | Can new evidence update the framework while preserving provenance and backward compatibility? | Incremental evidence ingestion, change log, expert approval and regression tests | Versioned files exist; automated extraction and update evaluation are pending |
每维1–5分,可标不适用;同时记录severity、hard fail、证据和pass / borderline / fail。总体平均分仅作描述,不能掩盖任何安全维度的低分或hard fail。
在完全隔离的180份held-out人工金标准上比较Global、Full-taxonomy、Example-RAG、Rubric-RAG和Oracle。主要指标为维度级quadratic weighted κ、MAE、exact/within-one accuracy,以及fail recall、false-negative rate和benign false-positive rate;20%样本重复三次评估稳定性。
当前预注册条件是baseline no-age与explicit-age。若要回答导师提出的plugin effectiveness,必须新增第三个protection-layer条件,并明确它在生成前检索什么、如何注入、是否允许拒绝,以及怎样防止test rubric泄漏到generation prompt。
Dashboard展示的新版360题全部唯一,但CLI默认使用的旧问题库只有134个唯一文本;两套文本只重合2条。直接运行正式API会得到与dashboard不一致的数据,因此必须先统一canonical inputs。
Loads verified workstream states, evidence and next gates.
| 工作流 | 当前状态 | 可核查证据 | 下一门槛 |
|---|---|---|---|
| Age/risk rubric cards | Draft complete | 30 cards; all marked draft_pending_expert_review | Expert review, revision log, and inter-rater calibration |
| Canonical experiment inputs | Blocked by version mismatch | CLI legacy bank has 134 unique prompts; overlap with dashboard bank is 2 | Make the v1 questions/cards the only canonical inputs and regenerate paired samples |
| Human gold standard | Not started | 0 of 300 planned responses adjudicated | Ethics determination, expert availability, annotator training, calibration and held-out labeling |
| Judge comparison and statistics | Code implemented; no evidence | Retrieval, judgment schema, agreement and bootstrap modules exist; no judgment files | Validate retrieval, then run five judge conditions against human gold |
| Protection-layer intervention | Concept only | Current preregistration compares no-age with explicit-age, not baseline with a plugin | Define the generation-time intervention and add a third, paired experimental condition |
| Question bank | Draft complete | 360 unique English prompts; 180 target and 180 benign | Dual human review and ecological-validity check; no claim of real child utterances |
| Source corpus | Draft corpus frozen | 8 source artifacts; 71 crosswalk labels; 91 YAIR leaves | Run a documented Social Science search and evidence-screening protocol |
| Target-model responses | Not started | One cheap-model smoke test exists; 0 formal responses from the planned 720 | Run a small engineering pilot after asset freeze, then collect the formal paired set |
| Unified taxonomy | Draft complete | 10 families and 50 mechanisms; OWL/TTL export available | Expert content validation and traceability to primary social-science evidence |
我们要做的事systematic review、自动multi-agent extraction、全新benchmark、judge validation和plugin effectiveness全部完成。但是当前论文聚焦evidence-grounded framework + Rubric-RAG judge validation + age-cue audit;protection layer可作为清晰定义后的扩展实验或第二篇工作。
Loads the proposed paper structure and the argumentative job of each section.
| 章节 | 该章节要完成的论证 |
|---|---|
| 1. Introduction | Frame Social Science for AI, the child-specific governance gap and the central research questions. |
| 2. Related work | Separate child-safety taxonomies, child-facing benchmarks, LLM-as-a-Judge, evidence synthesis and AI governance. |
| 3. Evidence-to-framework pipeline | Specify retrieval, screening, evidence extraction, traceability, conflict resolution, expert review and update mechanisms. |
| 4. Child protection framework | Present the 10 families, mechanisms, age bands, inclusion/exclusion rules and provenance crosswalk. |
| 5. Rubric construction | Explain the six dimensions, age/risk cards, anchors, hard fails, verdict logic and content validation. |
| 6. Benchmark and human gold | Describe question provenance, target/benign controls, paired conditions, sampling, blinding, annotation and ethics. |
| 7. Judge-validation experiment | Compare five judging methods on held-out human labels, including retrieval and repeatability tests. |
| 8. Model audit / intervention | Analyze age-cue effects; include a protection-layer comparison only after the intervention is formally specified. |
| 9. Results and error analysis | Report dimension-level agreement, failure rates, over-refusal, subgroup effects and candidate taxonomy extensions. |
| 10. Limitations, ethics and updateability | Bound synthetic-question validity, judge bias, source coverage, annotator welfare and framework drift. |
| 11. Conclusion | State the validated contribution without claiming expert review or plugin effectiveness before evidence exists. |
完成所有的e2e pipeline,但是需要薛老师的input。
Loads prioritized next actions and their exit criteria.
| 优先级 | 行动 | 交付物 | 完成标准 |
|---|---|---|---|
| 1Source: Next-action dependency query | Freeze one canonical asset set | v1 cards + v1 questions + regenerated 720 paired Sample records; legacy files explicitly retired | All commands and dashboards resolve to the same IDs and hashes |
| 2Source: Next-action dependency query | Write the Social Science evidence protocol | Search databases/queries, eligibility criteria, evidence schema, agent roles, audit log and update policy | Another researcher can reproduce one end-to-end evidence extraction |
| 3Source: Next-action dependency query | Run expert and human asset review | Reviewed cards, reviewed prompt bank, disagreement log and revised version | No asset is labeled expert-reviewed without recorded reviewer decisions |
| 4Source: Next-action dependency query | Run a low-cost engineering pilot | 60 base prompts × 2 age conditions, retrieval diagnostics and three judge conditions | API/JSON success ≥98% and rubric Recall@3 ≥0.90 |
| 5Source: Next-action dependency query | Lock the paper-level scope with advisors | Decision between current judge-validation paper and expanded protection-layer paper | Research questions, intervention conditions and contribution claims match the implemented experiment |
| 6Source: Next-action dependency query | Collect formal evidence | 720 target responses, 300 triple-rated human labels, five judge conditions and paired statistics | All preregistered quality gates and provenance checks are reported |
时间表从伦理确认和专家可用之日开始;如果这两个前置条件未满足,只执行到技术pilot,不进入正式人工实验。
在这五点锁定前,可以完成技术pilot,但不宜启动昂贵的正式API采样和300份三人标注。
Loads headline counts from the reviewed repository-audit snapshot.
SELECT * FROM summaryLoads the reviewed age-band and target/benign distribution.
SELECT age_band, question_type, count, denominator_per_age FROM question_balance ORDER BY age_band, question_typeLoads prompt-level expected-risk severity counts and shares.
SELECT severity, label, count, share FROM severity_distribution ORDER BY CAST(severity AS INTEGER)Loads verified workstream states, evidence and next gates.
SELECT workstream, state, evidence, next_gate FROM progress ORDER BY workstreamLoads the advisor-aligned research program and current maturity of each module.
SELECT study, question, method, status FROM study_map ORDER BY studyLoads the proposed paper structure and the argumentative job of each section.
SELECT section, job FROM paper_outline ORDER BY CAST(SUBSTR(section, 1, INSTR(section, '.') - 1) AS INTEGER)Loads prioritized next actions and their exit criteria.
SELECT priority, action, deliverable, exit_criterion FROM next_actions ORDER BY priority