11 KiB
11 KiB
| 1 | source_id | citation | year | publication_status | evidence_type | domain_and_data | sample_or_scope | main_result | project_use | limitations | supports | does_not_support | doi_or_id | source_url | verification_status | checked_date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2 | L01 | Chen et al. Pre-statistical harmonization of behavioral instruments across eight surveys and trials | 2021 | peer_reviewed | empirical workflow/methods | Behavioral instruments across eight dementia surveys and trials | Eight studies; manual instrument review plus automated raw-data checks | Comparable-looking items often differed in wording, response options, scoring, or direction and required pre-statistical review | Defines the source-review and crosswalk work required before statistical linking | Different population and constructs; does not test semantic embeddings or cross-national DIF | Official wording, response options, scoring, and populations must be reviewed before pooling | Semantic similarity alone establishes psychometric equivalence | 10.1186/s12874-021-01431-6 | https://doi.org/10.1186/s12874-021-01431-6 | verified_primary | 2026-09-20 |
| 3 | L02 | Kołczyńska. Combining multiple survey sources: A reproducible workflow and toolbox for survey data harmonization | 2022 | peer_reviewed | methods/workflow | Four cross-national survey projects; trust items | ESS, EVS, EQLS and Eurobarometer example | Crosswalk-centered, human-auditable documentation improves reproducibility of ex-post harmonization | Supports machine-readable source crosswalks, recodes, provenance and status tracking | Focuses recoding/documentation rather than latent linking or item semantics | Harmonization decisions and transformations need reusable documentation | A documented crosswalk proves measurement invariance | 10.1177/20597991221077923 | https://doi.org/10.1177/20597991221077923 | verified_primary | 2026-09-20 |
| 4 | L03 | McElroy et al. Using natural language processing to facilitate the harmonisation of mental health questionnaires | 2024 | peer_reviewed | empirical validation | Five mental-health questionnaires in a UK adult sample | 2,058 participants; 741 item pairs | Sentence-BERT semantic similarity correlated moderately with empirical item correlations and predicted held-out pair correlations with small error | Closest evidence for semantic item matching and response-structure signal | Adult UK sample, overlapping questionnaires and shared respondents; manual rules still needed; no cross-country DIF or survey-design inference | Text embeddings can help propose harmonization candidates | Embedding similarity verifies psychometric equivalence or transportability | 10.1186/s12888-024-05954-2 | https://doi.org/10.1186/s12888-024-05954-2 | verified_primary | 2026-09-20 |
| 5 | L04 | Ravenda et al. Rethinking psychometrics through LLMs: how item semantics shape measurement and prediction in psychological questionnaires | 2025 | peer_reviewed | empirical proof-of-concept | Big Five, DASS-42, GAD-7 and PHQ-9 questionnaire data | Large public questionnaire datasets; proof-of-concept response prediction | Semantic structure predicted empirical correlation patterns and supported prediction of responses to unseen items | Shows semantic representations can encode response-structure information in psychological questionnaires | Cross-cultural and multilingual transport were not established; predictive proof-of-concept is not survey harmonization | Item semantics may explain part of response covariance | Universal psychometric equivalence, DIF recovery or calibrated cross-national latent scores | 10.1038/s41598-025-21289-8 | https://doi.org/10.1038/s41598-025-21289-8 | verified_primary | 2026-09-20 |
| 6 | L05 | Yancey et al. BERT-IRT: Accelerating Item Piloting with BERT Embeddings and Explainable IRT Models | 2024 | peer_reviewed_conference | method plus operational evaluation | Duolingo English Test items | High-stakes language assessment item bank; exact proprietary sample details require full-paper extraction | BERT embeddings and engineered features reduced pilot length while maintaining reported criterion validity and reliability | Direct precedent for text features predicting IRT item parameters | Educational test items differ from suicide-related survey items; does not address cross-country DIF, complex samples or latent phenotype harmonization | Text-derived item features can inform item-parameter estimation | This project's core method is unprecedented or immediately transferable to health surveys | ACL Anthology 2024.bea-1.35 | https://aclanthology.org/2024.bea-1.35/ | verified_primary | 2026-09-20 |
| 7 | L06 | Chen and Chen. From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings | 2026 | preprint | method/benchmark | Mathematics and medical-licensure item banks | Two item banks; repeated cross-validation and simulation-based ceilings | Difficulty was more predictable than other parameters; reliability/design ceilings and repeated splits changed interpretation | Requires uncertainty-aware targets, repeated grouped validation and ceiling analysis for semantic parameter prediction | Preprint; educational/assessment domains; not cross-national mental-health surveys | Parameter-prediction benchmarks need target reliability and design ceilings | Reported RMSE alone establishes useful semantic signal | arXiv:2607.07141 | https://arxiv.org/abs/2607.07141 | verified_preprint | 2026-09-20 |
| 8 | L07 | Peters et al. Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review | 2025 | preprint | systematic review | Automated item-difficulty prediction | 37 articles through May 2025 | Language models can predict item difficulty in some settings, but studies vary in datasets, splits, targets and metrics | Maps existing text-to-difficulty literature and prevents novelty overclaiming | Preprint; focuses large-scale assessment rather than health questionnaires, DIF or survey design | Text-based difficulty prediction is an established research area | Reported best-case metrics transfer to this project | arXiv:2509.23486 | https://arxiv.org/abs/2509.23486 | verified_preprint | 2026-09-20 |
| 9 | L08 | Muthén and Asparouhov. IRT studies of many groups: the alignment method | 2014 | peer_reviewed | method plus Monte Carlo | Binary knowledge items across many country groups | Two surveys plus simulation | Alignment estimates group factor means/variances without requiring exact invariance and reports parameter non-invariance | Core comparator for many-country measurement invariance and DIF | Requires a prespecified factor structure and adequate linkage; alignment is not proof that all groups share one construct | Approximate invariance can be studied across many groups | Alignment repairs absent empirical connections or identifies a scale from semantics alone | 10.3389/fpsyg.2014.00978 | https://doi.org/10.3389/fpsyg.2014.00978 | verified_primary | 2026-09-20 |
| 10 | L09 | Mansolf et al. Extensions of Multiple-Group Item Response Theory Alignment | 2020 | peer_reviewed | method, simulation and application | International psychiatric genomics consortium with disparate item sets and formats | Multiple sites/instruments plus real-data-based simulation | Extended alignment accommodated differing item sets and response categories and recovered parameters in simulation | Closest latent-harmonization comparator for psychiatric phenotypes with nonidentical instruments | Needs specified construct/factor model and empirical connections; population and sampling designs differ from this project | Disparate psychiatric item sets can sometimes be aligned with explicit assumptions | Semantic priors alone create a common scale or eliminate anchor requirements | 10.1177/0013164419897307 | https://doi.org/10.1177/0013164419897307 | verified_primary | 2026-09-20 |
| 11 | L10 | Heinz et al. Item response theory and differential test functioning analysis of the HBSC-Symptom-Checklist across 46 countries | 2022 | peer_reviewed | cross-national psychometric application | Eight-item adolescent HBSC symptom checklist | 229,906 adolescents across 46 countries | Configural/metric invariance was more defensible than scalar invariance; alignment identified item non-invariance | Demonstrates the scale of cross-country adolescent DIF and consequences for comparisons | Uses one established common instrument, not different survey tools or unseen-item semantic prediction | Cross-national adolescent comparisons require item-level invariance/DIF checks | A common questionnaire automatically yields scalar comparability | 10.1186/s12874-022-01698-3 | https://doi.org/10.1186/s12874-022-01698-3 | verified_primary | 2026-09-20 |
| 12 | L11 | Savitsky and Williams. Pseudo Bayesian Mixed Models under Informative Sampling | 2022 | peer_reviewed | method, simulation and application | Hierarchical models under informative multistage sampling | Simulation plus business-establishment survey example | Weighting only unit likelihood contributions can remain biased when random effects correlate with design; weighting random-effect distributions addresses this setting | Constrains how hierarchical country/survey effects and design weights can be combined | Not an IRT application; requires inclusion-probability information and design assumptions | Complex-sample Bayesian multilevel models need design-aware treatment beyond naive weighted likelihood | Multiplying every likelihood by a weight guarantees correct interval coverage | 10.2478/jos-2022-0039 | https://doi.org/10.2478/jos-2022-0039 | verified_primary | 2026-09-20 |
| 13 | L12 | Wu and Stephenson. Bayesian estimation methods for survey data with potential applications to health disparities research | 2024 | peer_reviewed_review | narrative methodological review | Bayesian analysis of complex survey data | Reviews MRP, weighted pseudo-likelihood and synthetic-population approaches | No single Bayesian survey method is universally sufficient; assumptions and target estimands determine the route | Provides the survey-design method map for later model specifications and sensitivity analyses | Review rather than project-specific validation; does not resolve IRT identification | Multiple defensible Bayesian survey strategies exist and must be chosen by estimand/design | Bayesian modeling automatically corrects informative sampling | 10.1002/wics.1633 | https://doi.org/10.1002/wics.1633 | verified_primary | 2026-09-20 |
| 14 | L13 | Wu et al. Statistical harmonization of versions of measures across studies using external data | 2025 | peer_reviewed | calibration-sample method | Self-rated health and memory measured with different response formats | External calibration sample of 300 participants | A bridge sample answering both versions enabled model-based statistical harmonization with moderate agreement | Shows why bridge data may be necessary when archival surveys lack empirical links | Different constructs and older clinical population; external sample design differs from multi-item IRT | Targeted bridge data can identify transformations unavailable from disconnected archives | Text similarity can replace empirical bridge data without uncertainty | 10.1016/j.annepidem.2025.01.002 | https://doi.org/10.1016/j.annepidem.2025.01.002 | verified_primary | 2026-09-20 |