chore(repo): initialize reproducible research workspace

This commit is contained in:
Jinotech
2026-09-20 05:17:11 +12:00
commit 1e7cc2a71b
36 changed files with 9169 additions and 0 deletions
+14
View File
@@ -0,0 +1,14 @@
source_id,citation,year,publication_status,evidence_type,domain_and_data,sample_or_scope,main_result,project_use,limitations,supports,does_not_support,doi_or_id,source_url,verification_status,checked_date
L01,Chen et al. Pre-statistical harmonization of behavioral instruments across eight surveys and trials,2021,peer_reviewed,empirical workflow/methods,Behavioral instruments across eight dementia surveys and trials,Eight studies; manual instrument review plus automated raw-data checks,"Comparable-looking items often differed in wording, response options, scoring, or direction and required pre-statistical review",Defines the source-review and crosswalk work required before statistical linking,Different population and constructs; does not test semantic embeddings or cross-national DIF,"Official wording, response options, scoring, and populations must be reviewed before pooling",Semantic similarity alone establishes psychometric equivalence,10.1186/s12874-021-01431-6,https://doi.org/10.1186/s12874-021-01431-6,verified_primary,2026-09-20
L02,Kołczyńska. Combining multiple survey sources: A reproducible workflow and toolbox for survey data harmonization,2022,peer_reviewed,methods/workflow,Four cross-national survey projects; trust items,"ESS, EVS, EQLS and Eurobarometer example","Crosswalk-centered, human-auditable documentation improves reproducibility of ex-post harmonization","Supports machine-readable source crosswalks, recodes, provenance and status tracking",Focuses recoding/documentation rather than latent linking or item semantics,Harmonization decisions and transformations need reusable documentation,A documented crosswalk proves measurement invariance,10.1177/20597991221077923,https://doi.org/10.1177/20597991221077923,verified_primary,2026-09-20
L03,McElroy et al. Using natural language processing to facilitate the harmonisation of mental health questionnaires,2024,peer_reviewed,empirical validation,Five mental-health questionnaires in a UK adult sample,"2,058 participants; 741 item pairs",Sentence-BERT semantic similarity correlated moderately with empirical item correlations and predicted held-out pair correlations with small error,Closest evidence for semantic item matching and response-structure signal,"Adult UK sample, overlapping questionnaires and shared respondents; manual rules still needed; no cross-country DIF or survey-design inference",Text embeddings can help propose harmonization candidates,Embedding similarity verifies psychometric equivalence or transportability,10.1186/s12888-024-05954-2,https://doi.org/10.1186/s12888-024-05954-2,verified_primary,2026-09-20
L04,Ravenda et al. Rethinking psychometrics through LLMs: how item semantics shape measurement and prediction in psychological questionnaires,2025,peer_reviewed,empirical proof-of-concept,"Big Five, DASS-42, GAD-7 and PHQ-9 questionnaire data",Large public questionnaire datasets; proof-of-concept response prediction,Semantic structure predicted empirical correlation patterns and supported prediction of responses to unseen items,Shows semantic representations can encode response-structure information in psychological questionnaires,Cross-cultural and multilingual transport were not established; predictive proof-of-concept is not survey harmonization,Item semantics may explain part of response covariance,"Universal psychometric equivalence, DIF recovery or calibrated cross-national latent scores",10.1038/s41598-025-21289-8,https://doi.org/10.1038/s41598-025-21289-8,verified_primary,2026-09-20
L05,Yancey et al. BERT-IRT: Accelerating Item Piloting with BERT Embeddings and Explainable IRT Models,2024,peer_reviewed_conference,method plus operational evaluation,Duolingo English Test items,High-stakes language assessment item bank; exact proprietary sample details require full-paper extraction,BERT embeddings and engineered features reduced pilot length while maintaining reported criterion validity and reliability,Direct precedent for text features predicting IRT item parameters,"Educational test items differ from suicide-related survey items; does not address cross-country DIF, complex samples or latent phenotype harmonization",Text-derived item features can inform item-parameter estimation,This project's core method is unprecedented or immediately transferable to health surveys,ACL Anthology 2024.bea-1.35,https://aclanthology.org/2024.bea-1.35/,verified_primary,2026-09-20
L06,Chen and Chen. From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings,2026,preprint,method/benchmark,Mathematics and medical-licensure item banks,Two item banks; repeated cross-validation and simulation-based ceilings,Difficulty was more predictable than other parameters; reliability/design ceilings and repeated splits changed interpretation,"Requires uncertainty-aware targets, repeated grouped validation and ceiling analysis for semantic parameter prediction",Preprint; educational/assessment domains; not cross-national mental-health surveys,Parameter-prediction benchmarks need target reliability and design ceilings,Reported RMSE alone establishes useful semantic signal,arXiv:2607.07141,https://arxiv.org/abs/2607.07141,verified_preprint,2026-09-20
L07,Peters et al. Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review,2025,preprint,systematic review,Automated item-difficulty prediction,37 articles through May 2025,"Language models can predict item difficulty in some settings, but studies vary in datasets, splits, targets and metrics",Maps existing text-to-difficulty literature and prevents novelty overclaiming,"Preprint; focuses large-scale assessment rather than health questionnaires, DIF or survey design",Text-based difficulty prediction is an established research area,Reported best-case metrics transfer to this project,arXiv:2509.23486,https://arxiv.org/abs/2509.23486,verified_preprint,2026-09-20
L08,Muthén and Asparouhov. IRT studies of many groups: the alignment method,2014,peer_reviewed,method plus Monte Carlo,Binary knowledge items across many country groups,Two surveys plus simulation,Alignment estimates group factor means/variances without requiring exact invariance and reports parameter non-invariance,Core comparator for many-country measurement invariance and DIF,Requires a prespecified factor structure and adequate linkage; alignment is not proof that all groups share one construct,Approximate invariance can be studied across many groups,Alignment repairs absent empirical connections or identifies a scale from semantics alone,10.3389/fpsyg.2014.00978,https://doi.org/10.3389/fpsyg.2014.00978,verified_primary,2026-09-20
L09,Mansolf et al. Extensions of Multiple-Group Item Response Theory Alignment,2020,peer_reviewed,"method, simulation and application",International psychiatric genomics consortium with disparate item sets and formats,Multiple sites/instruments plus real-data-based simulation,Extended alignment accommodated differing item sets and response categories and recovered parameters in simulation,Closest latent-harmonization comparator for psychiatric phenotypes with nonidentical instruments,Needs specified construct/factor model and empirical connections; population and sampling designs differ from this project,Disparate psychiatric item sets can sometimes be aligned with explicit assumptions,Semantic priors alone create a common scale or eliminate anchor requirements,10.1177/0013164419897307,https://doi.org/10.1177/0013164419897307,verified_primary,2026-09-20
L10,Heinz et al. Item response theory and differential test functioning analysis of the HBSC-Symptom-Checklist across 46 countries,2022,peer_reviewed,cross-national psychometric application,Eight-item adolescent HBSC symptom checklist,"229,906 adolescents across 46 countries",Configural/metric invariance was more defensible than scalar invariance; alignment identified item non-invariance,Demonstrates the scale of cross-country adolescent DIF and consequences for comparisons,"Uses one established common instrument, not different survey tools or unseen-item semantic prediction",Cross-national adolescent comparisons require item-level invariance/DIF checks,A common questionnaire automatically yields scalar comparability,10.1186/s12874-022-01698-3,https://doi.org/10.1186/s12874-022-01698-3,verified_primary,2026-09-20
L11,Savitsky and Williams. Pseudo Bayesian Mixed Models under Informative Sampling,2022,peer_reviewed,"method, simulation and application",Hierarchical models under informative multistage sampling,Simulation plus business-establishment survey example,Weighting only unit likelihood contributions can remain biased when random effects correlate with design; weighting random-effect distributions addresses this setting,Constrains how hierarchical country/survey effects and design weights can be combined,Not an IRT application; requires inclusion-probability information and design assumptions,Complex-sample Bayesian multilevel models need design-aware treatment beyond naive weighted likelihood,Multiplying every likelihood by a weight guarantees correct interval coverage,10.2478/jos-2022-0039,https://doi.org/10.2478/jos-2022-0039,verified_primary,2026-09-20
L12,Wu and Stephenson. Bayesian estimation methods for survey data with potential applications to health disparities research,2024,peer_reviewed_review,narrative methodological review,Bayesian analysis of complex survey data,"Reviews MRP, weighted pseudo-likelihood and synthetic-population approaches",No single Bayesian survey method is universally sufficient; assumptions and target estimands determine the route,Provides the survey-design method map for later model specifications and sensitivity analyses,Review rather than project-specific validation; does not resolve IRT identification,Multiple defensible Bayesian survey strategies exist and must be chosen by estimand/design,Bayesian modeling automatically corrects informative sampling,10.1002/wics.1633,https://doi.org/10.1002/wics.1633,verified_primary,2026-09-20
L13,Wu et al. Statistical harmonization of versions of measures across studies using external data,2025,peer_reviewed,calibration-sample method,Self-rated health and memory measured with different response formats,External calibration sample of 300 participants,A bridge sample answering both versions enabled model-based statistical harmonization with moderate agreement,Shows why bridge data may be necessary when archival surveys lack empirical links,Different constructs and older clinical population; external sample design differs from multi-item IRT,Targeted bridge data can identify transformations unavailable from disconnected archives,Text similarity can replace empirical bridge data without uncertainty,10.1016/j.annepidem.2025.01.002,https://doi.org/10.1016/j.annepidem.2025.01.002,verified_primary,2026-09-20
1 source_id citation year publication_status evidence_type domain_and_data sample_or_scope main_result project_use limitations supports does_not_support doi_or_id source_url verification_status checked_date
2 L01 Chen et al. Pre-statistical harmonization of behavioral instruments across eight surveys and trials 2021 peer_reviewed empirical workflow/methods Behavioral instruments across eight dementia surveys and trials Eight studies; manual instrument review plus automated raw-data checks Comparable-looking items often differed in wording, response options, scoring, or direction and required pre-statistical review Defines the source-review and crosswalk work required before statistical linking Different population and constructs; does not test semantic embeddings or cross-national DIF Official wording, response options, scoring, and populations must be reviewed before pooling Semantic similarity alone establishes psychometric equivalence 10.1186/s12874-021-01431-6 https://doi.org/10.1186/s12874-021-01431-6 verified_primary 2026-09-20
3 L02 Kołczyńska. Combining multiple survey sources: A reproducible workflow and toolbox for survey data harmonization 2022 peer_reviewed methods/workflow Four cross-national survey projects; trust items ESS, EVS, EQLS and Eurobarometer example Crosswalk-centered, human-auditable documentation improves reproducibility of ex-post harmonization Supports machine-readable source crosswalks, recodes, provenance and status tracking Focuses recoding/documentation rather than latent linking or item semantics Harmonization decisions and transformations need reusable documentation A documented crosswalk proves measurement invariance 10.1177/20597991221077923 https://doi.org/10.1177/20597991221077923 verified_primary 2026-09-20
4 L03 McElroy et al. Using natural language processing to facilitate the harmonisation of mental health questionnaires 2024 peer_reviewed empirical validation Five mental-health questionnaires in a UK adult sample 2,058 participants; 741 item pairs Sentence-BERT semantic similarity correlated moderately with empirical item correlations and predicted held-out pair correlations with small error Closest evidence for semantic item matching and response-structure signal Adult UK sample, overlapping questionnaires and shared respondents; manual rules still needed; no cross-country DIF or survey-design inference Text embeddings can help propose harmonization candidates Embedding similarity verifies psychometric equivalence or transportability 10.1186/s12888-024-05954-2 https://doi.org/10.1186/s12888-024-05954-2 verified_primary 2026-09-20
5 L04 Ravenda et al. Rethinking psychometrics through LLMs: how item semantics shape measurement and prediction in psychological questionnaires 2025 peer_reviewed empirical proof-of-concept Big Five, DASS-42, GAD-7 and PHQ-9 questionnaire data Large public questionnaire datasets; proof-of-concept response prediction Semantic structure predicted empirical correlation patterns and supported prediction of responses to unseen items Shows semantic representations can encode response-structure information in psychological questionnaires Cross-cultural and multilingual transport were not established; predictive proof-of-concept is not survey harmonization Item semantics may explain part of response covariance Universal psychometric equivalence, DIF recovery or calibrated cross-national latent scores 10.1038/s41598-025-21289-8 https://doi.org/10.1038/s41598-025-21289-8 verified_primary 2026-09-20
6 L05 Yancey et al. BERT-IRT: Accelerating Item Piloting with BERT Embeddings and Explainable IRT Models 2024 peer_reviewed_conference method plus operational evaluation Duolingo English Test items High-stakes language assessment item bank; exact proprietary sample details require full-paper extraction BERT embeddings and engineered features reduced pilot length while maintaining reported criterion validity and reliability Direct precedent for text features predicting IRT item parameters Educational test items differ from suicide-related survey items; does not address cross-country DIF, complex samples or latent phenotype harmonization Text-derived item features can inform item-parameter estimation This project's core method is unprecedented or immediately transferable to health surveys ACL Anthology 2024.bea-1.35 https://aclanthology.org/2024.bea-1.35/ verified_primary 2026-09-20
7 L06 Chen and Chen. From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings 2026 preprint method/benchmark Mathematics and medical-licensure item banks Two item banks; repeated cross-validation and simulation-based ceilings Difficulty was more predictable than other parameters; reliability/design ceilings and repeated splits changed interpretation Requires uncertainty-aware targets, repeated grouped validation and ceiling analysis for semantic parameter prediction Preprint; educational/assessment domains; not cross-national mental-health surveys Parameter-prediction benchmarks need target reliability and design ceilings Reported RMSE alone establishes useful semantic signal arXiv:2607.07141 https://arxiv.org/abs/2607.07141 verified_preprint 2026-09-20
8 L07 Peters et al. Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review 2025 preprint systematic review Automated item-difficulty prediction 37 articles through May 2025 Language models can predict item difficulty in some settings, but studies vary in datasets, splits, targets and metrics Maps existing text-to-difficulty literature and prevents novelty overclaiming Preprint; focuses large-scale assessment rather than health questionnaires, DIF or survey design Text-based difficulty prediction is an established research area Reported best-case metrics transfer to this project arXiv:2509.23486 https://arxiv.org/abs/2509.23486 verified_preprint 2026-09-20
9 L08 Muthén and Asparouhov. IRT studies of many groups: the alignment method 2014 peer_reviewed method plus Monte Carlo Binary knowledge items across many country groups Two surveys plus simulation Alignment estimates group factor means/variances without requiring exact invariance and reports parameter non-invariance Core comparator for many-country measurement invariance and DIF Requires a prespecified factor structure and adequate linkage; alignment is not proof that all groups share one construct Approximate invariance can be studied across many groups Alignment repairs absent empirical connections or identifies a scale from semantics alone 10.3389/fpsyg.2014.00978 https://doi.org/10.3389/fpsyg.2014.00978 verified_primary 2026-09-20
10 L09 Mansolf et al. Extensions of Multiple-Group Item Response Theory Alignment 2020 peer_reviewed method, simulation and application International psychiatric genomics consortium with disparate item sets and formats Multiple sites/instruments plus real-data-based simulation Extended alignment accommodated differing item sets and response categories and recovered parameters in simulation Closest latent-harmonization comparator for psychiatric phenotypes with nonidentical instruments Needs specified construct/factor model and empirical connections; population and sampling designs differ from this project Disparate psychiatric item sets can sometimes be aligned with explicit assumptions Semantic priors alone create a common scale or eliminate anchor requirements 10.1177/0013164419897307 https://doi.org/10.1177/0013164419897307 verified_primary 2026-09-20
11 L10 Heinz et al. Item response theory and differential test functioning analysis of the HBSC-Symptom-Checklist across 46 countries 2022 peer_reviewed cross-national psychometric application Eight-item adolescent HBSC symptom checklist 229,906 adolescents across 46 countries Configural/metric invariance was more defensible than scalar invariance; alignment identified item non-invariance Demonstrates the scale of cross-country adolescent DIF and consequences for comparisons Uses one established common instrument, not different survey tools or unseen-item semantic prediction Cross-national adolescent comparisons require item-level invariance/DIF checks A common questionnaire automatically yields scalar comparability 10.1186/s12874-022-01698-3 https://doi.org/10.1186/s12874-022-01698-3 verified_primary 2026-09-20
12 L11 Savitsky and Williams. Pseudo Bayesian Mixed Models under Informative Sampling 2022 peer_reviewed method, simulation and application Hierarchical models under informative multistage sampling Simulation plus business-establishment survey example Weighting only unit likelihood contributions can remain biased when random effects correlate with design; weighting random-effect distributions addresses this setting Constrains how hierarchical country/survey effects and design weights can be combined Not an IRT application; requires inclusion-probability information and design assumptions Complex-sample Bayesian multilevel models need design-aware treatment beyond naive weighted likelihood Multiplying every likelihood by a weight guarantees correct interval coverage 10.2478/jos-2022-0039 https://doi.org/10.2478/jos-2022-0039 verified_primary 2026-09-20
13 L12 Wu and Stephenson. Bayesian estimation methods for survey data with potential applications to health disparities research 2024 peer_reviewed_review narrative methodological review Bayesian analysis of complex survey data Reviews MRP, weighted pseudo-likelihood and synthetic-population approaches No single Bayesian survey method is universally sufficient; assumptions and target estimands determine the route Provides the survey-design method map for later model specifications and sensitivity analyses Review rather than project-specific validation; does not resolve IRT identification Multiple defensible Bayesian survey strategies exist and must be chosen by estimand/design Bayesian modeling automatically corrects informative sampling 10.1002/wics.1633 https://doi.org/10.1002/wics.1633 verified_primary 2026-09-20
14 L13 Wu et al. Statistical harmonization of versions of measures across studies using external data 2025 peer_reviewed calibration-sample method Self-rated health and memory measured with different response formats External calibration sample of 300 participants A bridge sample answering both versions enabled model-based statistical harmonization with moderate agreement Shows why bridge data may be necessary when archival surveys lack empirical links Different constructs and older clinical population; external sample design differs from multi-item IRT Targeted bridge data can identify transformations unavailable from disconnected archives Text similarity can replace empirical bridge data without uncertainty 10.1016/j.annepidem.2025.01.002 https://doi.org/10.1016/j.annepidem.2025.01.002 verified_primary 2026-09-20
Binary file not shown.
File diff suppressed because one or more lines are too long
Binary file not shown.

After

Width:  |  Height:  |  Size: 328 KiB

@@ -0,0 +1,69 @@
# 阶段 0 文献定位与贡献边界
日期:2026-09-20
矩阵:`literature_matrix.csv`
范围:聚焦检索,不是 PRISMA 系统综述
## 检索问题
本轮只覆盖支撑 Gate 0 的五个问题:
1. 跨调查行为/心理问卷在统计建模前需要怎样的来源和题目审核?
2. NLP/LLM 题目表示是否已经用于问卷匹配或响应结构预测?
3. 文本特征是否已经用于预测 IRT 题目参数?
4. 多组 IRT、alignment 和 DIF 如何处理多国家或不同题集?
5. Bayesian 层级模型如何处理复杂和信息性抽样?
检索优先使用期刊页面、PubMed/PMC、ACL Anthology 和论文预印本原页。无法独立确认存在或元数据的来源不进入矩阵。同行评审与预印本分开标记。
## 证据综合
### 1. 数据协调必须先做题目和来源审计
Chen 等和 Kołczyńska 的工作支持先核对研究总体、题目正文、回答选项、计分方向、版本及转换记录。它们不支持仅凭相似变量名合并,也不支持把有文档的 crosswalk 当作测量等价证据。
### 2. 语义问卷协调不是空白领域
McElroy 等已用 Sentence-BERT 比较心理健康问卷题目,并在共同作答样本中检验语义相似与实证相关。Ravenda 等进一步展示题目语义与问卷响应结构及 unseen-item 响应预测的关系。因此,本项目不能宣称首次把语言模型用于心理问卷结构或协调。
### 3. 文本预测 IRT 参数已有直接先例
BERT-IRT 已把 BERT 嵌入和工程特征用于题目参数估计。Chen 与 Chen 的预印本及 Peters 等的预印本综述进一步表明 text-to-parameter / item-difficulty modeling 已形成独立研究线,并提示目标可靠性上限、重复交叉验证和 scale-free 指标的重要性。因此,“让嵌入预测题目难度”本身不是充分创新。
### 4. 多组、跨国和不同题集协调已有成熟比较对象
Muthén 与 Asparouhov 的 alignment、Mansolf 等对不同题集/回答格式的扩展,以及 Heinz 等对 46 国青少年量表的分析,构成项目必须比较或讨论的方法基础。共同问卷也可能不满足 scalar invariance;不同题集的 alignment 仍需要预设构念结构和经验连接。
### 5. 复杂抽样不能用简单加权似然一句带过
Savitsky 与 Williams 表明,在多层信息性抽样下,只给个体 likelihood 加权仍可能不足。Wu 与 Stephenson 的综述也显示 MRP、pseudo-likelihood 和 synthetic population 各自依赖不同估计目标和设计信息。本项目必须把设计型描述性估计、模型内权重处理和设计一致的区间/重抽样检查分开。
## 当前可辩护的贡献候选
现有矩阵没有发现一项已核验工作同时完成以下组合:
- 以官方题目、回答标签、时间窗口和真实复杂抽样回答建立跨 GSHS、YRBS、NSDUH 的可追溯题库;
- 在题目家族级留出下,让多语言语义表示条件化阈值、区分度和受约束 DIF;
- 同时评估未见题目、未见问卷、未见国家/地区和时间外迁移;
- 对未见题目/国家效应传播不确定性,而不是设为零;
- 与直接协调、传统/层级 IRT、人工标签和非语义模型做公平消融;
- 明确处理权重、PSU、stratum 及设计型不确定性。
这只是阶段 0 的候选贡献边界,不是首创证明。正式论文前仍需扩大检索、做引用追踪和更新检索。
## 对研究设计的约束
1. 语义只提供候选先验或预测信息,不能替代经验锚定与可识别性。
2. 训练/测试必须按题目家族和外层域分组;单次随机划分和 RMSE 不足以支持新题目校准。
3. 对经验题目参数进行监督学习时,需传播参数估计误差或使用后验样本。
4. 若跨调查图不连通,应提出桥接样本或限制主张,不用语义相似伪造共同标度。
5. alignment 和其他传统协调方法是必要基线,不是可省略的背景文献。
6. 复杂抽样区间需独立验证;Bayesian 标签本身不保证设计一致性。
## 矩阵构成
- 总条目:13
- 同行评审:11
- 预印本:2
- 每条均记录:证据类型、数据/样本、主要结果、项目用途、限制、允许支持与不允许支持的主张、DOI/永久链接和核验状态。
+138
View File
@@ -0,0 +1,138 @@
# Language-Conditioned Psychometric Harmonization Model:项目章程
版本:0.1.0
冻结日期:2026-09-20
状态:`passed_for_gate_0`
适用范围:阶段 0–2;阶段 2 完成后依据可识别性证据修订
## 1. 项目目的与主要研究问题
本项目不是普通个体风险预测竞赛,也不预设所有自杀相关变量属于一个单维量表。项目首先确认数据结构支持 IRT、层级测量、阶段/多结局模型还是只能直接协调。
**Primary RQ** 在 GSHS、YRBS 与 NSDUH 青少年相关数据中,题目正文、时间窗口和回答标签的语义表示,能否相对于最佳适用的无文本层级测量基线和人工构念标签基线,提高未见题目家族、未见问卷版本及未见国家/工具上的测量预测与校准,同时不造成有实质意义的区间质量退化?
支撑问题:
1. 现有调查的题目共现、锚定连接和题目家族多样性是否足以识别各构念?
2. 措辞、时间窗口、回答格式、语言、国家、年份和工具是否产生有实质意义的 DIF 或测量非等价?
3. 题目语义能否解释训练资料中的响应结构和题目参数,并跨题目家族泛化?
4. 对未见国家、工具和题目,模型的不确定性是否校准;zero-shot 与 few-shot 的表现如何?
5. 复杂抽样、缺失/跳题、锚题和总体定义的合理替代方案是否改变结论?
## 2. 可证伪假设
语言条件化测量模型在预先冻结的外层迁移任务中,相对于最佳适用的 M0/M1/M2 与人工标签基线,在指定共同主指标上达到预设最小有意义改善,且校准或覆盖不劣于预设界限。
改善阈值、非劣界限和主指标数值将在阶段 2 的信息量审计和阶段 3 的模拟/功效校准后冻结。若语义模型没有稳定增益,项目输出应是阴性 benchmark 或适用边界,不更换测试集或事后挑选指标。
## 3. 目标总体和时间地域范围
- 对象:青少年调查参与者。
- GSHS:跨国开发候选,主要是学校在读样本;常见核心年龄为 13–17 岁,但必须逐国家—年份组件核验。
- YRBS:问卷和时间迁移候选;美国高中生样本不得外推到全部美国青少年。
- NSDUH:外部候选;家庭抽样框与学校调查框不可静默视为同一总体。
- 跨工具比较优先使用年龄与适用总体共同支持域。精确年龄带在阶段 1 建立总体适用表后冻结;不能统一的总体分层报告。
- YRBS 原始年份候选为 1991–2023;当前处理入口只到 2019,2021/2023 在布局验证前隔离。
- NSDUH 当前处理候选为 2021–2024,只纳入确认属于青少年模块且有官方问题/编码证据的变量。
- 国家背景仅在测量层建立后加入,并按调查时点可得性匹配。
## 4. 构念与变量角色
测量层分开登记,不预设单维:
1. suicidal ideation
2. suicide plan
3. suicide attempt(二元和次数形式分开)
4. self-harm(是否共模需证据)
5. sadness/hopelessness(是否共模需证据)
欺凌、孤独、睡眠、物质使用、家庭/同伴支持、暴力暴露及国家背景属于解释或分层变量,不因相关或预测作用自动变成自杀风险量表题目。横断面关联不表述为因果或个体纵向预测。
阶段 2 为每个构念选择:IRT/多维测量、阶段/潜类别/多结局、直接二分类/频数协调,或分调查报告。
## 5. 数据角色与唯一候选入口
| 数据源 | 候选角色 | 唯一候选入口 | 状态 |
|---|---|---|---|
| GSHS | 跨国开发及未见国家验证 | `Dataset/GSHS-全球学生健康调查数据/GSHS/01_data/GSHS.csv` | `audit_required` |
| YRBS | 问卷版本及时间迁移 | 处理包 `YRBS/full_by_year`19912019);原始核验用 `YRBS_National_1991_2023/raw` | `conditional` |
| NSDUH | 外部工具/总体候选 | `可直接分析数据包_NSDUH_YRBS/NSDUH/nsduh_2021_2024_full.parquet` | `conditional` |
| World Bank/WHO context | 后期国家—年份解释层 | `WHO-country_context/world_bank_country_year_1990_2025.csv` 及来源长表 | `conditional` |
| 其他目录 | 外部候选、文档或 provenance archive | 无 | 不进入核心分析,除非书面改版 |
入口是审计入口,不表示其中变量已经通过题义、总体、设计或许可审计。
## 6. 冻结的迁移任务
1. **Unseen item family**:同一家族全部措辞、翻译、回答和监督信号留出。
2. **Unseen questionnaire version**:整个版本及泄漏等价版本留出。
3. **Unseen country**:该国全部年份和组件回答留出,未知国家效应从层级分布预测并传播不确定性。
4. **Unseen region**:整个预定地区留出。
5. **Temporal extrapolation**:早期训练、后期测试,背景按预测时点可得性处理。
6. **Cross-instrument**:总体和结局定义审计通过后再跨工具评估。
随机个体划分只用于调试或辅助比较,不作为核心迁移主张证据。
## 7. 估计目标和指标层级
主要目标是题目阈值/难度、区分度、DIF、共同标度成立时的潜在表型/群体分布,以及新题目/新域的预测分布与不确定性。
共同主指标候选:未见题目家族 held-out log score;未见国家/工具共同目标上的校准或分布误差;模拟已知真值下参数恢复与覆盖;群体估计的设计加权误差与区间表现。AUPRC、AUROC、Brier 等为辅助预测指标,不能单独证明测量协调成功。
## 8. 识别和调查设计原则
- 语义相似只产生候选连接,不证明锚题不变或共同标度成立。
- 国家均值、阈值平移和无约束 DIF 可能混淆;正式模型必须声明标度、参考组、锚题和 DIF 约束。
- 独立信息单位主要是题目家族/版本及国家/工具等外层单元;大量受访者不能补偿题目种类不足。
- 保留权重、PSU、stratum、国家—年份及组件标识。
- 权重进入似然不等同于设计一致区间;描述性设计方差、模型内权重和设计型重抽样分开验证。
- 未施测、不适用、合法跳题、拒答、不知道、普通缺失和真实阴性分开。
## 9. 范围外事项
- 临床诊断、个体自杀预测或资源分配工具。
- 用国家自杀死亡率作个体结局。
- 无总体限定地把学校样本外推到全部青少年。
- 在跨构念/数据集/任务证据不足时称为 foundation model。
- 在阶段 2/3 前冻结最终深度架构、最终嵌入模型或结果后阈值。
- 把 AI 预审登记为两名人类审核者。
- 分发未经确认允许的微观数据。
## 10. 伦理、许可和治理
- 数据按去标识化二次数据处理,但使用仍受来源现行条款和机构要求约束。
- 不重新识别参与者,不公开再分发许可未确认的微数据。
- 主要结局进入确认性分析前需两名实际审核者复核;AI 检查只标 provisional。
- 输出记录 AI 辅助、数据来源、版本、哈希、模型/包版本及 unknown 状态。
- 原始数据只读;衍生产物写入 `research/`
- 改变总体、构念、数据角色、测试任务或成功判据时,必须提高本章程版本并在 `Progression.md` 记录原因。
## 11. 阶段 0 待验证清单
1. 各调查组件的年龄、在校状态、地域、语言和问卷版本。
2. 主要结局官方正文、答案标签、时间窗口、跳题和派生规则。
3. GSHS 权重、PSU、stratum 的调查级有效性。
4. YRBS 2021/2023 官方布局、导入程序、记录数和本地文件一致性。
5. NSDUH 青少年/成人/COVID/派生变量边界及正确权重/方差设计。
6. 锚题图、题目共现、阳性事件、结构缺失及有效题目家族数。
7. 各来源现行许可、引用要求及衍生数据分发范围。
8. 最接近方法文献和真实贡献边界;不得预设首创。
9. 缺失的 R、复杂抽样、PyMC/ArviZ 和心理测量环境的锁定方案。
## 12. Gate 与停止条件
Gate 0 要求本章程、数据清单、文献矩阵和环境审计齐备。Gate 1 要求正式题目均有来源、编码、总体与缺失规则且主要结局完成人工双重核查。Gate 2 要求每个构念得到可进入潜变量模型/有条件/仅直接协调/不纳入的证据判定。
官方文本或编码无法确认、图不连通/标度不可识别、题目家族不足、设计字段不完整或参数恢复失败时按工作流降级;不得通过增加模型复杂度、替换最终测试或扩大主张掩盖失败。
## 13. 输入绑定
| 输入 | SHA-256 |
|---|---|
| `Document/research-workflow.md` | `394a191359298132fc2a56304dbc49b5934f47af22d2e0fff47d28a0d9edcaeb` |
| `Document/项目的正式定位.md` | `28a415ff5d42649df2a7c68079ae0abfd295cbffb5fead216934a895cca4cd13` |
| `Document/language-conditioned-psychometric-harmonization-model.md` | `39d2050735f4df28c6abe08f8d1de6094b7ac11d0b6f2ccf9fd66d4cc450a825` |
| `research/audit/data_manifest.csv` | `0427a19c3287a2691edbd9a4486917989fe369fcb0f729a01477ce27cd8f98f2` |
| `research/audit/data_source_registry.csv` | `ac68b9928f8855c5373034a30adda316da1da2f1a1e890b0c23f836f12d55b67` |
| `research/audit/environment_audit.md` | `c84a6b05718bbaf7762021aa45b2f983e7c9ef844d41592d2138439251e929eb` |