Files

11 KiB

1source_idcitationyearpublication_statusevidence_typedomain_and_datasample_or_scopemain_resultproject_uselimitationssupportsdoes_not_supportdoi_or_idsource_urlverification_statuschecked_date
2L01Chen et al. Pre-statistical harmonization of behavioral instruments across eight surveys and trials2021peer_reviewedempirical workflow/methodsBehavioral instruments across eight dementia surveys and trialsEight studies; manual instrument review plus automated raw-data checksComparable-looking items often differed in wording, response options, scoring, or direction and required pre-statistical reviewDefines the source-review and crosswalk work required before statistical linkingDifferent population and constructs; does not test semantic embeddings or cross-national DIFOfficial wording, response options, scoring, and populations must be reviewed before poolingSemantic similarity alone establishes psychometric equivalence10.1186/s12874-021-01431-6https://doi.org/10.1186/s12874-021-01431-6verified_primary2026-09-20
3L02Kołczyńska. Combining multiple survey sources: A reproducible workflow and toolbox for survey data harmonization2022peer_reviewedmethods/workflowFour cross-national survey projects; trust itemsESS, EVS, EQLS and Eurobarometer exampleCrosswalk-centered, human-auditable documentation improves reproducibility of ex-post harmonizationSupports machine-readable source crosswalks, recodes, provenance and status trackingFocuses recoding/documentation rather than latent linking or item semanticsHarmonization decisions and transformations need reusable documentationA documented crosswalk proves measurement invariance10.1177/20597991221077923https://doi.org/10.1177/20597991221077923verified_primary2026-09-20
4L03McElroy et al. Using natural language processing to facilitate the harmonisation of mental health questionnaires2024peer_reviewedempirical validationFive mental-health questionnaires in a UK adult sample2,058 participants; 741 item pairsSentence-BERT semantic similarity correlated moderately with empirical item correlations and predicted held-out pair correlations with small errorClosest evidence for semantic item matching and response-structure signalAdult UK sample, overlapping questionnaires and shared respondents; manual rules still needed; no cross-country DIF or survey-design inferenceText embeddings can help propose harmonization candidatesEmbedding similarity verifies psychometric equivalence or transportability10.1186/s12888-024-05954-2https://doi.org/10.1186/s12888-024-05954-2verified_primary2026-09-20
5L04Ravenda et al. Rethinking psychometrics through LLMs: how item semantics shape measurement and prediction in psychological questionnaires2025peer_reviewedempirical proof-of-conceptBig Five, DASS-42, GAD-7 and PHQ-9 questionnaire dataLarge public questionnaire datasets; proof-of-concept response predictionSemantic structure predicted empirical correlation patterns and supported prediction of responses to unseen itemsShows semantic representations can encode response-structure information in psychological questionnairesCross-cultural and multilingual transport were not established; predictive proof-of-concept is not survey harmonizationItem semantics may explain part of response covarianceUniversal psychometric equivalence, DIF recovery or calibrated cross-national latent scores10.1038/s41598-025-21289-8https://doi.org/10.1038/s41598-025-21289-8verified_primary2026-09-20
6L05Yancey et al. BERT-IRT: Accelerating Item Piloting with BERT Embeddings and Explainable IRT Models2024peer_reviewed_conferencemethod plus operational evaluationDuolingo English Test itemsHigh-stakes language assessment item bank; exact proprietary sample details require full-paper extractionBERT embeddings and engineered features reduced pilot length while maintaining reported criterion validity and reliabilityDirect precedent for text features predicting IRT item parametersEducational test items differ from suicide-related survey items; does not address cross-country DIF, complex samples or latent phenotype harmonizationText-derived item features can inform item-parameter estimationThis project's core method is unprecedented or immediately transferable to health surveysACL Anthology 2024.bea-1.35https://aclanthology.org/2024.bea-1.35/verified_primary2026-09-20
7L06Chen and Chen. From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings2026preprintmethod/benchmarkMathematics and medical-licensure item banksTwo item banks; repeated cross-validation and simulation-based ceilingsDifficulty was more predictable than other parameters; reliability/design ceilings and repeated splits changed interpretationRequires uncertainty-aware targets, repeated grouped validation and ceiling analysis for semantic parameter predictionPreprint; educational/assessment domains; not cross-national mental-health surveysParameter-prediction benchmarks need target reliability and design ceilingsReported RMSE alone establishes useful semantic signalarXiv:2607.07141https://arxiv.org/abs/2607.07141verified_preprint2026-09-20
8L07Peters et al. Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review2025preprintsystematic reviewAutomated item-difficulty prediction37 articles through May 2025Language models can predict item difficulty in some settings, but studies vary in datasets, splits, targets and metricsMaps existing text-to-difficulty literature and prevents novelty overclaimingPreprint; focuses large-scale assessment rather than health questionnaires, DIF or survey designText-based difficulty prediction is an established research areaReported best-case metrics transfer to this projectarXiv:2509.23486https://arxiv.org/abs/2509.23486verified_preprint2026-09-20
9L08Muthén and Asparouhov. IRT studies of many groups: the alignment method2014peer_reviewedmethod plus Monte CarloBinary knowledge items across many country groupsTwo surveys plus simulationAlignment estimates group factor means/variances without requiring exact invariance and reports parameter non-invarianceCore comparator for many-country measurement invariance and DIFRequires a prespecified factor structure and adequate linkage; alignment is not proof that all groups share one constructApproximate invariance can be studied across many groupsAlignment repairs absent empirical connections or identifies a scale from semantics alone10.3389/fpsyg.2014.00978https://doi.org/10.3389/fpsyg.2014.00978verified_primary2026-09-20
10L09Mansolf et al. Extensions of Multiple-Group Item Response Theory Alignment2020peer_reviewedmethod, simulation and applicationInternational psychiatric genomics consortium with disparate item sets and formatsMultiple sites/instruments plus real-data-based simulationExtended alignment accommodated differing item sets and response categories and recovered parameters in simulationClosest latent-harmonization comparator for psychiatric phenotypes with nonidentical instrumentsNeeds specified construct/factor model and empirical connections; population and sampling designs differ from this projectDisparate psychiatric item sets can sometimes be aligned with explicit assumptionsSemantic priors alone create a common scale or eliminate anchor requirements10.1177/0013164419897307https://doi.org/10.1177/0013164419897307verified_primary2026-09-20
11L10Heinz et al. Item response theory and differential test functioning analysis of the HBSC-Symptom-Checklist across 46 countries2022peer_reviewedcross-national psychometric applicationEight-item adolescent HBSC symptom checklist229,906 adolescents across 46 countriesConfigural/metric invariance was more defensible than scalar invariance; alignment identified item non-invarianceDemonstrates the scale of cross-country adolescent DIF and consequences for comparisonsUses one established common instrument, not different survey tools or unseen-item semantic predictionCross-national adolescent comparisons require item-level invariance/DIF checksA common questionnaire automatically yields scalar comparability10.1186/s12874-022-01698-3https://doi.org/10.1186/s12874-022-01698-3verified_primary2026-09-20
12L11Savitsky and Williams. Pseudo Bayesian Mixed Models under Informative Sampling2022peer_reviewedmethod, simulation and applicationHierarchical models under informative multistage samplingSimulation plus business-establishment survey exampleWeighting only unit likelihood contributions can remain biased when random effects correlate with design; weighting random-effect distributions addresses this settingConstrains how hierarchical country/survey effects and design weights can be combinedNot an IRT application; requires inclusion-probability information and design assumptionsComplex-sample Bayesian multilevel models need design-aware treatment beyond naive weighted likelihoodMultiplying every likelihood by a weight guarantees correct interval coverage10.2478/jos-2022-0039https://doi.org/10.2478/jos-2022-0039verified_primary2026-09-20
13L12Wu and Stephenson. Bayesian estimation methods for survey data with potential applications to health disparities research2024peer_reviewed_reviewnarrative methodological reviewBayesian analysis of complex survey dataReviews MRP, weighted pseudo-likelihood and synthetic-population approachesNo single Bayesian survey method is universally sufficient; assumptions and target estimands determine the routeProvides the survey-design method map for later model specifications and sensitivity analysesReview rather than project-specific validation; does not resolve IRT identificationMultiple defensible Bayesian survey strategies exist and must be chosen by estimand/designBayesian modeling automatically corrects informative sampling10.1002/wics.1633https://doi.org/10.1002/wics.1633verified_primary2026-09-20
14L13Wu et al. Statistical harmonization of versions of measures across studies using external data2025peer_reviewedcalibration-sample methodSelf-rated health and memory measured with different response formatsExternal calibration sample of 300 participantsA bridge sample answering both versions enabled model-based statistical harmonization with moderate agreementShows why bridge data may be necessary when archival surveys lack empirical linksDifferent constructs and older clinical population; external sample design differs from multi-item IRTTargeted bridge data can identify transformations unavailable from disconnected archivesText similarity can replace empirical bridge data without uncertainty10.1016/j.annepidem.2025.01.002https://doi.org/10.1016/j.annepidem.2025.01.002verified_primary2026-09-20