Multimodal Behavioral Typicality as a Training-Free Screening Signal for Dementia

Leticia Pinto-Alva, Thomas Lord Department of Computer Science, University of Southern California, Los Angeles, California, USA and Institute for Creative Technologies, University of Southern California, Los Angeles, California, USA, pintoalv@usc.edu
Gale Lucas, USC Chan Division of Occupational Science and Occupational Therapy, University of Southern California, Los Angeles, California, USA; Institute for Creative Technologies, University of Southern California, Los Angeles, California, USA and Thomas Lord Department of Computer Science, University of Southern California, Los Angeles, California, USA, lucas@ict.usc.edu
Maja J Matarić, Thomas Lord Department of Computer Science, University of Southern California, Los Angeles, California, USA, mataric@usc.edu
Jesse Thomason, Thomas Lord Department of Computer Science, University of Southern California, Los Angeles, California, USA, jessetho@usc.edu

Dementia is commonly described as impairing what people attend to, say, and mean more than how they move their eyes or produce speech. We test this asymmetry with a cross-modal behavioral marker—negative log-likelihood (NLL) under frozen pretrained models across gaze, text, and audio—that separates semantic engagement from production mechanics without training on dementia data at any stage. On a new multimodal dataset of 39 participants (25 control, 14 dementia) performing the Cookie Theft task, semantic models separate control from dementia in gaze (Hedges’ g = 1.04) and text (g up to 1.31), with a smaller effect in linguistic speech (g = 0.66) that does not survive correction. Their mechanically-matched counterparts—bottom-up saliency and temporal gaze dynamics for gaze, an acoustic codec for audio—do not reach significance. For text, removing surface disfluencies strengthens rather than weakens the signal, indicating the effect does not reduce to fluency. The signal is model-independent: ten of eleven language models separate groups (g = 0.95 to 1.24). Three-way fusion reaches AUC = 0.94, and external validation on 549 DementiaBank Pitt Corpus transcripts replicates the direction across all eight autoregressive models tested. This is a screening signal rather than a diagnostic, and the implication for multimodal interaction systems is direct: adapt to semantic engagement, not production mechanics.

CCS Concepts:Human-centered computing → HCI design and evaluation methods; • Computing methodologies → Natural language processing;

Keywords: Dementia detection, eye gaze, multimodal interaction, training-free inference, negative log-likelihood, CLIP, Cookie Theft task

ACM Reference Format:
Leticia Pinto-Alva, Gale Lucas, Maja J Matarić, and Jesse Thomason. 2026. Multimodal Behavioral Typicality as a Training-Free Screening Signal for Dementia. In INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION (ICMI '26), October 05--09, 2026, Napoli, Italy. ACM, New York, NY, USA 10 Pages. https://doi.org/10.1145/3776574.3831111

Figure 1
Figure 1: Example from our multimodal dataset: temporal gaze scanpath for a participant with dementia shown with the simultaneously recorded description. Color encodes progression through the recording (blue = start, red = end) in both the scanpath and the transcript. The transcript shown is the manual annotation used for the text analyses (Section 3). This participant's three-way fusion score (z = +0.61) falls near the dementia group mean (+ 0.57). The Cookie Theft stimulus is from the Boston Diagnostic Aphasia Examination, 3rd ed. [14], © PRO-ED, Inc.

1 Introduction

The clinical diagnostic process for Alzheimer's disease (AD) remains slow and reactive, typically initiated only after a patient or family notices decline [1, 30]; there is a clear need for low-burden monitoring tools. Dementia degrades the voluntary control of looking more than its basic machinery: in mild-to-moderate AD, prosaccade latency and amplitude can be statistically indistinguishable from age-matched controls while uncorrected antisaccade errors rise roughly tenfold [8], though oculomotor findings vary with disease severity across studies [28]. The semantic processes that determine what to look at, what to say, and how to organize meaning are disrupted more consistently [9, 12, 28]. However, many approaches to automated cognitive screening train supervised classifiers on labeled dementia data [25, 26], which face three linked limitations: they require diagnosed patient cohorts for training that are expensive and demographically narrow; they risk learning surface artifacts (filler-word rates, recording conditions, demographics) rather than semantic, cognitive markers; and they generalize less readily across populations and spoken languages. We use no training on dementia data at any stage. We leverage pretrained models to score patient behavioral typicality via negative log-likelihood (NLL) across gaze, text, and audio collected during the Cookie Theft picture description task [14]. The central question becomes: does this person's behavior look typical to a model that has never seen any dementia data? Beyond detection, this framework surfaces a theoretically meaningful cross-modal pattern: models that measure the typicality of semantic content separate groups in gaze and text and directionally in audio, while mechanically-matched baselines do not. Our results suggest that dementia impairs what people communicate more than how they produce behavior. For conversational and assistive systems, this finding inverts what should be measured: coherence of engagement, not production. We collected a new multimodal dataset of 41 participants (27 control, 14 dementia; 39 analyzable) with simultaneously recorded gaze, speech, and video during the Cookie Theft task (Figure 1 shows a representative dementia case; group-level contrasts appear in Figure 3). Our study tests two questions: (RQ1) Do semantic models separate dementia from control more strongly than the mechanically-matched baselines we can construct within each modality, and is the text signal robust to removal of surface disfluencies? (RQ2) Does multimodal fusion improve detection beyond any single modality?

Our contributions are:

  1. A new behavioral marker: a cross-modal content-vs-mechanics pattern, shown by matched mechanical baselines in gaze (CLIP vs. GradCAM) and audio (wav2vec2 vs. EnCodec), with text showing surface-robustness (removing filler words strengthens the signal), corroborated across ten of eleven language models—evidence that dementia preferentially disrupts content over mechanics.
  2. A CLIP-based annotation-free gaze pipeline: a pretrained vision-language model as an automatic annotator identifies semantically relevant regions; a GMM fit over those regions defines the expected gaze distribution. This pipeline produces more effective control versus dementia separation than eight hand-crafted alternatives by effect size (g = 1.04, pFDR = .021).
  3. A unified NLL framework that treats gaze, text, and audio symmetrically under pretrained models. Three-way fusion reaches AUC = 0.94, with each modality contributing partially independent signal.
  4. External validation on 549 DementiaBank Pitt Corpus transcripts confirming text-NLL generalization across datasets collected decades apart.

2 Related Work

Our contribution touches two literatures: automated dementia detection from speech and text, and eye tracking for clinical assessment.

2.1 Dementia Detection from Speech and Text

The DementiaBank Pitt Corpus [4] is the most widely used dataset for speech-based dementia detection, providing Cookie Theft picture descriptions from elderly controls and individuals with dementia (primarily probable Alzheimer's disease). The ADReSS [26] and ADReSSo [25] challenges provided demographically balanced subsets, with top systems reaching 80–90% classification accuracy under supervised learning, chiefly by fine-tuning pretrained language models on ADReSS transcripts [3, 42]. A potential limitation of such supervised classifiers is that they may learn surface artifacts—filler-word rates, recording conditions, speaker demographics—rather than genuine cognitive markers, a concern we revisit empirically in Section 5 through a filler-removal analysis. Recent paired perplexity methods [13, 23, 40] contrast transcript likelihood under control-adapted and impaired language models—the latter obtained by training on dementia transcripts or by deliberately degrading the control model—and report further gains, but each relies on in-domain dementia or control transcripts. Beyond speech-only pipelines, Fraser et al. [12] combined linguistic and acoustic features to identify AD from Cookie Theft narratives, and König et al. [22] showed that automatic vocal markers can distinguish healthy controls, MCI, and early AD; both rely on supervised learning over hand-crafted feature sets, and neither includes a visual or gaze modality. Foundation models for medical domains [18, 29] have shown the promise of pretrained representations for clinical tasks, though medical imaging has dominated the focus, with little work on spontaneous behavior. Our work departs along a single axis: we use pretrained models directly, with no fine-tuning at any stage—distinguishing it from transfer learning [35], domain adaptation [5], and few-shot approaches [39], all of which rely on target-domain data. To our knowledge no prior dementia-detection work adopts this fully pretrained-only stance, a property valuable for rare conditions, underrepresented populations, and resource-limited settings.

2.2 Eye Tracking and Vision-Language Models for Clinical Assessment

Eye tracking has long been used to assess cognitive function in dementia through fixation analysis, saccade dynamics, and visual-exploration statistics [8, 28, 31]. Existing pipelines rely on either low-level oculomotor features or hand-labeled regions of interest defined by clinicians or annotators. Findings on scene exploration in AD are mixed, but several report a suggestive dissociation: patients overlook informative or incongruous scene regions while basic saccade measures remain comparable to controls [28]. Our gaze results give this clinical observation a computational form. Gaussian Mixture Models [27] are a standard tool for density estimation over spatial distributions, and EyeFormer [21] predicts personalized scanpaths with transformer-guided reinforcement learning. Vision-language models such as CLIP [32] learn joint image-text representations that enable zero-shot transfer, and have been applied successfully to medical imaging tasks [18, 44]. To our knowledge, no prior work uses a VLM to approximate a semantically meaningful expected-gaze distribution for AD assessment. We formulate an automated semantic annotation pipeline using CLIP scores that replaces hand-labeled regions-of-interest for gaze analysis.

3 Data

Participants viewed the Cookie Theft (Figure 1) picture from the Boston Diagnostic Aphasia Examination [14] and described everything they observed for up to two minutes, with the option to end the description earlier when they felt they had finished. This standardized clinical task elicits spontaneous narrative speech and has been used extensively in dementia research [9]. We collected data from 41 participants at a clinical research site (27 control undiagnosed, 14 diagnosed with dementia); modality-specific technical issues affecting three control sessions (detailed at the end of this section) reduce this to 39 analyzable participants (25 control, 14 dementia) in each single-modality analysis, and 38 for three-way fusion. Participants self-reported their age, gender, race, and education. Eligibility required participants to be aged 50 or older, sighted, and English-speaking; most participants were 70 or older. Across the 41 collected participants, the sample was 59% female and predominantly White (73%), with Black, Hispanic, Asian, and other participants comprising the remainder. Age was recorded in five-year bands rather than exact years. The two groups did not differ in age (Mann-Whitney U = 176.0, p = .53; ages available for 26 control and 12 dementia participants) and the dementia group was not older than the control group: the median band was 75–79 for controls and 70–74 for the dementia group. Gender composition was similar across groups (of the 41: control 16 female, 11 male; dementia 8 female, 6 male). Education was recorded for only seven participants and is therefore not analyzed. All data were collected under IRB approval with written informed consent.

What we know and do not know about our sample. Participants diagnosed with dementia were existing patients of the research site. Recruitment was opportunistic over several months, reflecting constraints on multimodal collection from a vulnerable population: dementia patients require additional IRB protections around consent and decision capacity, and clinical sites are conservative about research access. Site clinicians verbally confirmed each participant's control/dementia status to the research team; no written diagnostic records (severity staging such as CDR/MMSE, or etiology labels such as AD, vascular, Lewy body, or mixed) were shared. Because the site is a dementia research center, the dementia participant group in our study is likely enriched for Alzheimer's disease relative to the general dementia population, but we do not have verified, per-participant etiology and make no claims about severity or subtype composition. Controls were family members and study partners accompanying patients, who volunteered. Controls self-reported an absence of dementia diagnosis and were not formally screened—a common convenience-sample approach in clinical research. Recruitment was not extended to the surrounding community because diagnostic status could not be confirmed outside the clinic, which constrains sample size and diversity but keeps group assignment tied to clinician confirmation. We apply these characterization limits symmetrically when interpreting cross-dataset comparisons (Section 5.7).

Gaze was recorded using a Tobii Pro X3-120 eye tracker mounted on a consumer laptop. The device samples at 120 Hz; effective per-participant rates varied with tracking quality, since blinks, eye closure, and momentary loss of the corneal reflection are common in an elderly cohort. Effective rate did not differ between groups (Mann-Whitney U = 214.5, p = .25), and the number of usable gaze points differed only marginally (control M = 8, 514; dementia M = 6, 318; U = 241.0, p = .055). Because NLL is computed over each participant's pooled gaze coordinates rather than over detected fixations, sample density affects the precision of a participant's score but not its expected value. Audio was captured via the laptop microphone. Transcripts were produced independently by two sources: Whisper [33] for automatic speech recognition, and manual human annotation; manual transcripts were used for the text NLL analyses reported here, with Whisper transcripts retained as an ASR reference. Video was recorded via laptop webcam at 60 FPS but is not analyzed here. All modalities were synchronized to the Cookie Theft stimulus presentation window, with recording durations ranging from 35 seconds to 2 minutes.

Three control participants were affected by modality-specific technical issues: one session produced no usable recordings; one had an eye-tracker calibration failure (speech and transcripts preserved); one had a microphone-channel issue (gaze preserved). Because transcripts derive from the audio, text and audio share an exclusion set. Accounting for these issues yields n = 39 for each single-modality analysis (25 control, 14 dementia), with gaze excluding the fully-missing and calibration-failure sessions, and text/audio excluding the fully-missing and microphone-issue sessions. Three-way fusion requires usable data across all three modalities and is restricted to n = 38 (24 control, 14 dementia).

4 Methods

Three properties carry through every subsection: every score is an NLL under a pretrained model, every modality is z-normalized before fusion, and no parameters are estimated from dementia-labeled data.

4.1 NLL as a Unified Typicality Metric

Across modalities, we compute NLL under pretrained models as a measure of behavioral typicality. Pretrained models encode distributional expectations about “normal” behavior—normal gaze patterns, normal language, normal speech. Dementia-related behavioral changes should produce higher NLL (more surprising behavior) without any model fine-tuning. Once z-normalized, NLL becomes a common currency across modalities. Beyond scoring, we also compare two classes of pretrained model within each modality. Semantic models (CLIP for gaze, language models for text, self-supervised speech models for audio) are trained to represent content rather than low-level signal structure. Mechanical models (bottom-up saliency, free-viewing temporal dynamics, an acoustic codec) are trained to represent low-level sensory or temporal structure independent of content. Holding this axis constant where possible (gaze and audio) lets the cross-modal pattern (Section 5) surface cleanly: a mechanical baseline that fails where its semantic counterpart succeeds is hard to attribute to a pipeline artifact.

4.2 Gaze: Systematic Comparison of Golden Target Strategies

Each participant's real recorded gaze is scored against a golden target: a distribution over expected fixation locations derived from the stimulus image alone, identical for every participant and estimated without reference to any gaze data. Group separation therefore arises from how real gaze scores against a shared reference—no gaze is simulated. A central challenge in gaze-based cognitive screening is the absence of ground truth: there is no “correct” way to look at the Cookie Theft picture. Constructing the golden target is therefore the central design question, and we systematically tested ten strategies organized into two categories. Our experimental design is hypothesis-driven: the ten gaze strategies test a specific theoretical prediction (semantic vs. mechanical contrast), not a search for optimal configurations.

Hand-annotated semantic targets (8 strategies): We annotated the Cookie Theft image with pixel-level labels for 11 semantic objects (characters, key action zones, scene context). Eight clinically motivated configurations were constructed encoding different hypotheses about which scene elements matter. Automated targets (2 strategies): A bottom-up saliency strategy (GradCAM) uses ResNet50 [15] and GradCAM [36] activation maps to identify regions of high visual contrast and edge density—where low-level features draw the eye regardless of meaning. A vision-language grounding strategy (CLIP) uses CLIP (ViT-B/16) [32] as an automatic annotator rather than a classifier. The Cookie Theft image (1302 × 943 pixels) is scanned with 64 × 64 patches at stride 32, yielding 1,092 overlapping patches. We define seven descriptive prompts capturing the Cookie Theft narrative (e.g., “a boy standing on a stool stealing cookies from a jar,” “water overflowing from the sink onto the floor”); these prompts were selected a priori based on the narrative elements of the picture and were not modified after initial evaluation. Each image patch and each prompt are encoded by CLIP's vision and text encoders into a shared 512-dimensional space. For each patch pi, we average its cosine similarity across the seven descriptive prompts:

\begin{equation} s_i = \frac{1}{7}\sum _{j=1}^{7} \frac{\text{CLIP}_{\text{image}}(p_i) \cdot \text{CLIP}_{\text{text}}(t_j)}{||\text{CLIP}_{\text{image}}(p_i)|| \, ||\text{CLIP}_{\text{text}}(t_j)||} \end{equation}
(1)
Patches scoring above the 60th percentile are selected as the expected-gaze region set (GradCAM uses the 70th; both use K = 8). The downstream analysis (GMM fitting, NLL scoring) is identical across strategies; only the target construction differs. Automated targets sample 3,000 coordinates from selected regions (K = 8); hand-annotated targets use every 15th labeled pixel, with one component per annotated region initialized at its centroid:
\begin{equation} \text{NLL}_{\text{gaze}} = -\frac{1}{N} \sum _{n=1}^{N} \ln \sum _{k=1}^{K} \pi _k \, \mathcal {N}(\mathbf {x}_n \mid \boldsymbol {\mu }_k, \boldsymbol {\Sigma }_k) \end{equation}
(2)
where xn are the participant's gaze coordinates, and GMM parameters are estimated from image annotations only—never from participant data, preserving the training-free property. Raw gaze coordinates outside the valid normalized range [0, 1] were filtered prior to scoring to remove off-screen and tracker-artifact points.

Prompt-specificity ablation for CLIP. Replacing the Cookie Theft-specific prompts with task-unrelated random prompts (e.g., “a mountain landscape”) tests whether the effect depends on task semantics rather than on CLIP's vision encoder alone; results are reported in Section 5.2. We additionally evaluated EyeFormer [21], a transformer-based scanpath model, as a temporal-dynamics baseline; per-participant temporal NLL was computed over non-overlapping 7-second windows (mean 13.4 per participant), matching the free-viewing clip duration EyeFormer was trained on.

4.3 Text: Pretrained Language Models

For causal models, NLL is the mean per-token cross-entropy loss. For masked language models (BERT [11], RoBERTa [24]), we compute pseudo-log-likelihood [37]. We evaluate eleven models spanning three generations: GPT-2 (small: 117M, medium: 345M, large: 774M) [34], XLNet [41], BERT, and RoBERTa (2019); GPT-Neo 1.3B [7], GPT-J 6B [38], and OPT-1.3B [43] (2021–2022); Phi-2 (2.7B) [19] and Mistral-7B [20] (2023). Three transcript conditions test whether the NLL signal resides in surface disfluency (filler words such as “uh,” “um,” and pauses) or deeper semantic content: (1) paragraph (full transcript), (2) paragraph with filler words removed, (3) utterance-by-utterance (transcripts split into sentence-level units scored independently, removing cross-utterance discourse context).

4.4 Audio: Linguistic vs. Acoustic Models

We distinguish two categories of audio models. Linguistic speech models: HuBERT [17] and wav2vec2 [2] are self-supervised models pretrained on raw speech; they encode phonemic and lexical structure. We note these representations are not purely linguistic: they also carry prosodic, speaker-dependent, and recording-condition information. The audio arm is therefore a weaker test of the content-vs-mechanics contrast than gaze, where CLIP and GradCAM operate on identical input, and we interpret it accordingly (Section 5.7). Whisper [33], weakly rather than self-supervised, is reported alongside them for comparison. Acoustic codec (control condition): EnCodec [10] provides a content-agnostic baseline, trained to compress and reconstruct any audio signal without modeling linguistic content. Reconstruction MSE—the Gaussian NLL up to an affine constant, so g and AUC are unchanged—therefore measures purely acoustic typicality: voice quality and spectral characteristics independent of what is being said. This contrast between linguistic and acoustic models parallels CLIP-vs-GradCAM in gaze: if EnCodec separates groups, the signal lies at least partly in how people sound.

4.5 Late Multimodal Fusion

Per-participant NLL scores are standardized to z-scores within each modality. Late fusion is computed as the arithmetic mean of z-scores:

\begin{equation} z_{\text{fusion}} = \frac{1}{M} \sum _{m=1}^{M} \frac{\text{NLL}_m - \bar{\mu }_m}{\sigma _m}, \end{equation}
(3)
where $\bar{\mu }_m$ and σm are the mean and standard deviation of NLL for modality m across the sample. We deliberately chose arithmetic averaging over learned fusion to preserve the training-free property. The normalization statistics $\bar{\mu }_m$ and σm are computed over the pooled sample without reference to group labels, so no diagnostic-label information enters the scoring at any point; single-modality results are unchanged by this step, which serves only to place modalities on a common scale. In deployment, these statistics would be fixed from a reference cohort rather than estimated in-sample—a practical constraint on the framework, not a property of the NLL scores.

4.6 Grid-Search Over Model Configurations

As a robustness audit for our technique, we conducted an exhaustive grid search over all 1,320 combinations of gaze strategies (n = 10)  ×  language models and transcript conditions (n = 33)  ×  audio models (n = 4), using three-way fusion effect size as the evaluation criterion. Rather than selecting a “winner,” this audit asks whether the content-mechanics pattern holds across the entire configuration space—i.e., whether high-performing combinations consistently pair semantic models while mechanical baselines consistently fail. Results are reported in Section 5.6.

4.7 Evaluation

Because the dementia group is small (n = 14), we triangulate across three statistical test families to evaluate the separation of control and dementia participant scores: Welch's t-test (robust to unequal variance), Mann-Whitney U (robust to non-normality), and permutation tests (10,000 two-sided iterations, assumption-free). Effect sizes are reported as Hedges’ g [16] with small-sample bias correction and 95% bootstrap confidence intervals; positive g indicates higher NLL in dementia than control. Benjamini-Hochberg FDR correction [6] is applied within each modality, since our hypotheses concern within-modality semantic-vs-mechanical contrasts. Hedges’ g (sample-size invariant) is our primary effect-size metric; AUC (threshold-independent) is the primary classification metric. Per-modality analyses use n = 39 (see Section 3 for exclusions); three-way fusion uses n = 38.

5 Results

Three findings follow. Semantic models separate groups in gaze and text and directionally in audio, while mechanical baselines do not—a pattern corroborated by our 1,320-configuration audit. Three-way fusion reaches g = 1.79 (AUC = 0.94). External validation replicates the text signal but reverses for audio, marking the boundary of training-free generalization.

5.1 Content vs. Mechanics: A Cross-Modal Pattern

The comparison takes two forms (Table 1, Figure 2). In gaze and audio, we pair each semantic model with a mechanically-matched baseline that scores a different kind of structure within the same modality: CLIP vs. GradCAM and EyeFormer in gaze, wav2vec2 vs. EnCodec in audio. In gaze, the semantic model reaches FDR significance while both mechanical baselines do not; in audio, the semantic model shows a positive but nominal effect (Table 4) while the acoustic baseline is flat. In text, we cannot construct a fully content-free baseline of the same model family, so we instead test whether the signal depends on the most obvious surface production feature: filler words. Removing them strengthens rather than weakens the effect (Mistral-7B: g = 1.24 → 1.31), indicating the signal does not reduce to surface-level disfluency.

Table 1: Cross-modal pattern. In gaze and audio the semantic representation shows a positive effect while the mechanically-matched baseline does not; text is tested instead by disfluency removal. Gaze and text effects survive Benjamini-Hochberg correction within modality; the audio effect is nominal only († : pFDR = .119 across four audio models). With n = 39 the design is powered to detect g ≈ 0.96, so the baselines are non-significant rather than demonstrated nulls (EnCodec 95% CI [ − 0.56, +0.81]). All single-modality analyses on n = 39 (25 control, 14 dementia), with modality-specific exclusions; pperm from 10,000 two-sided permutations.
Modality Level g pperm Sig
Gaze CLIP Semantic +1.04 .002
GradCAM Bottom-up − 0.28 .407
EyeFormer Temporal − 0.02 .944
Text Mistral-7B (para) Content(+ disf.) +1.24 < .001
Mistral-7B (− disf.) Content only +1.31 < .001
Audio wav2vec2 Linguistic +0.66 .046
EnCodec Acoustic − 0.03 .945
Figure 2
Figure 2: Cross-modal pattern: effect sizes (Hedges’ g) by modality (n = 39). In each modality, at least one semantic representation shows a positive effect (CLIP and Mistral-7B reach FDR significance; wav2vec2 is nominal), while matched mechanical baselines (GradCAM, EyeFormer, EnCodec) do not. Hatching marks effects that do not survive FDR correction. The text signal strengthens when disfluencies are removed. Stars denote permutation-test significance (*p < .05, **p < .01, ***p < .001); † marks a nominal effect that does not survive FDR correction (pFDR = .119).

The pattern supports the interpretation: dementia impairs the content of behavioral engagement—what people attend to and say—while the production-level signals we can measure (low-level gaze statistics, acoustic reconstruction, transcript length) do not reach significance. We detail the per-modality results below.

5.2 Gaze: Spatial Strategies (n = 39)

Table 2: Gaze NLL across ten golden target strategies (n = 39: 25 control, 14 dementia). Two strategies reach FDR significance at α = 0.05 (marked ✓ in the Sig column). Bolded row marks the strongest-performing strategy by effect size (CLIP semantic), used as the gaze component in all subsequent fusion analyses. † marks the two automated strategies; the remaining eight are hand-annotated.
Strategy g pperm pFDR AUC Sig
CLIP semantic +1.04 .002 .021 .734
Full narrative / All annotated +0.93 .006 .031 .743
Characters + actions +0.65 .056 .187 .637
Unusual events +0.44 .182 .456 .583
Mother's activity +0.38 .253 .505 .560
Action/danger +0.03 .921 .991 .566
Kids’ mischief +0.02 .951 .991 .529
Scene context +0.02 .951 .991 .457
Characters only − 0.00 .991 .991 .471
GradCAM saliency† − 0.28 .407 .678 .380

CLIP achieved the strongest effect (g = 1.04, pperm = .002, pFDR = .021, AUC = 0.73), outperforming all eight hand-annotated alternatives by effect size. The direct contrast between CLIP (semantic) and GradCAM (bottom-up saliency, g = −0.28) constitutes a controlled comparison: both strategies produce probability distributions over the same image and feed an identical GMM-NLL pipeline, differing in whether the target encodes semantic content or low-level saliency. The thresholds also differ (CLIP 60th, GradCAM 70th percentile), but in the direction that disfavors CLIP: its broader region set flattens the target density and reduces discriminative power. The prompt ablation isolates that semantic difference further. Replacing task-relevant prompts with unrelated ones (g ≈ 0.67) preserves CLIP's architecture and the entire downstream pipeline while removing task semantics; pure saliency removes semantics altogether (g = −0.28). The ordering task-relevant > random ≫ saliency is a dose-response on semantic specificity. Among hand-crafted strategies, narrower configurations limited to single aspects of the scene—characters alone (g ≈ 0.00), scene context (g = 0.02), specific action zones (g ≈ 0 for action/danger and kids’ mischief)—fail to separate groups. The Full narrative / All annotated configuration, which covers the complete set of narrative-relevant objects, reaches g = 0.93 and FDR significance. CLIP's advantage over careful full-narrative labeling is modest (Δg = 0.11), but CLIP-derived annotation recovers that effect automatically, without the task-specific labor.

5.3 Gaze: Temporal Dynamics (n = 39)

EyeFormer temporal NLL showed no group separation (g = −0.02, pperm = .944). EyeFormer was trained and evaluated on free-viewing clips of roughly seven seconds, so we segmented each recording into non-overlapping 7-second windows (mean 13.4 per participant); the null persists with duration matched. The remaining mismatch is task rather than duration—free-viewing versus narrative description—and our framework predicts that a description-trained temporal model would recover effects EyeFormer does not, a falsifiable prediction testable once such a model exists. The null, combined with GradCAM's null and CLIP's positive effect, is consistent with localizing the difference to top-down, meaning-driven attention allocation, though the present design cannot establish this mechanism directly. A null in low-level gaze metrics alongside impaired content selection has precedent in the clinical scene-exploration literature [28].

5.4 Text: Pretrained Language Models (n = 39)

Table 3: Text NLL—paragraph condition (n = 39: 25 control, 14 dementia). Ten of eleven models reach significance and survive Benjamini-Hochberg correction across the eleven models (pFDR ≤ .010); XLNet does not (pFDR = .095). ‡ indicates masked LMs (pseudo-NLL).
Model Year g pperm AUC Sig
Mistral-7B 2023 +1.24 < .001 .789
BERT‡ 2019 +1.16 .001 .766
Phi-2 2023 +1.08 .002 .754
OPT-1.3B 2022 +1.05 .003 .746
GPT-2 large 2019 +1.02 .005 .737
GPT-Neo 1.3B 2021 +1.01 .004 .726
GPT-J 6B 2021 +1.00 .005 .734
GPT-2 medium 2019 +0.98 .006 .706
RoBERTa‡ 2019 +0.95 .007 .729
GPT-2 2019 +0.95 .009 .711
XLNet 2019 +0.54 .095 .657

Ten of eleven models achieved significance (Table 3), with effect sizes ranging from g = 0.95 to g = 1.24 for the significant models. This consistency across architectures, sizes (110M to 7B parameters), and release years (2019–2023) indicates that the signal is model-independent rather than an artifact of any single architecture.

Content vs. disfluency. For every autoregressive model, removing filler words strengthened the NLL-based group separation. Specifically, Mistral-7B's effect size rose from g = 1.24 (with fillers) to g = 1.31 (fillers removed)—a pattern also dominant in our grid-search audit (Section 5.6). While filler-word rate does differ between groups (dementia participants produced more; d = 0.94, p < .001), other surface features do not: word count, vocabulary size, and type-token ratio all fail to reach significance.

Context-dependence of the signal. Scoring transcripts utterance-by-utterance—which strips cross-utterance discourse context—yields a weaker but still significant effect for Mistral-7B (g = 0.91, pperm = .011). The attenuation from g = 1.24 (paragraph) to g = 0.91 (utterance) is consistent with the content-integration interpretation: some of the signal lies in how utterances cohere across the full description, not only within individual sentences. The method is not uniformly significant. Broad agreement across language models reflects their shared large-scale English pretraining rather than an indiscriminate positive: most hand-crafted gaze targets fail to separate groups (Table 2), the mechanical baselines do not reach significance in either gaze or audio (Table 1), XLNet does not reach significance (Table 3), and BERT and RoBERTa reverse direction on the external corpus (Section 5.7).

5.5 Audio (n = 39)

Table 4: Audio results (n = 39: 25 control, 14 dementia). wav2vec2 and Whisper show positive effects of comparable magnitude; neither survives Benjamini-Hochberg correction across the four audio models (pFDR = .119), so the audio effect is nominal (†). HuBERT is directionally consistent but weaker, and the acoustic codec shows no separation. Scores are per-chunk mean NLL (HuBERT, wav2vec2, Whisper) or reconstruction MSE (EnCodec). Bolded row marks wav2vec2, used as the audio component in fusion because it is a self-supervised linguistic representation (Section 4.4), not because it outperforms Whisper.
Model Type g pperm pFDR AUC Sig
Whisper Speech +0.69 .060 .119 .697
wav2vec2 Linguistic +0.66 .046 .119 .649
HuBERT Linguistic +0.44 .188 .251 .637
EnCodec Acoustic − 0.03 .945 .945 .540

The linguistic speech model wav2vec2 shows a positive effect (Table 4; g = 0.66, pperm = .046), as does Whisper (g = 0.69, pperm = .060); neither survives FDR correction across the four audio models (pFDR = .119), and the three test families disagree at the margin (wav2vec2: Welch p = .073, Mann-Whitney p = .132; Whisper: Welch p = .043, Mann-Whitney p = .045). HuBERT shows a positive but weaker trend (g = 0.44). EnCodec, which measures acoustic reconstruction independent of linguistic content, shows no group difference (g = −0.03, pperm = .945), directly paralleling the CLIP-vs-GradCAM contrast in gaze. The audio arm therefore supports the content-vs-mechanics pattern in direction, but not with the statistical strength of gaze and text (Section 4.4).

5.6 Multimodal Fusion and Grid-Search Robustness

Gaze+text reaches g = 1.69 (AUC = 0.88); adding audio yields g = 1.79 (AUC = 0.94). The remaining pairs are gaze+audio (g = 1.21, AUC = 0.84) and text+audio (g = 1.37, AUC = 0.82); all fusion analyses use n = 38. Fusion balances complementary error profiles (Figure 3). We report the paragraph condition (not the marginally stronger no-fillers condition, g = 1.31) since it preserves the transcript as naturally produced; as the grid-search audit shows, the result holds either way. An exhaustive audit of all 1,320 configurations (10 gaze strategies × 33 text model-condition pairs × 4 audio models) establishes that the content-mechanics pattern is not configuration-dependent. The grid covers spatial strategies only, so EyeFormer is excluded; GradCAM and EnCodec are its mechanical baselines. All top-10 configurations pair semantic gaze targets (CLIP or full-narrative annotation), large language models (Mistral-7B, BERT, or OPT-1.3B), and linguistic audio models (wav2vec2, Whisper, or HuBERT). No mechanical baseline (GradCAM, EnCodec) appears anywhere in the top 10, and neither does XLNet—the one language model that failed to reach individual significance. Seven of the top-10 use filler-removed transcripts, corroborating the “content, not disfluency” claim. The effect-size spread among the top-10 is narrow (Δg = 0.14 from rank #1 to #10), and the configuration we report throughout this paper (CLIP + Mistral-7B paragraph + wav2vec2) ranks #3 of 1,320—broad robustness rather than a single lucky combination.

Figure 3
Figure 3: Multimodal separation. Left: gaze × text, each dot one participant; the dashed diagonal marks the fusion threshold (zgaze + ztext = 0). Control participants cluster in the lower-left, dementia participants drift toward the upper-right. Right: three-way fusion scores per participant, sorted (g = 1.79, AUC = 0.94); controls (blue) fall negative, dementia participants (red) positive.

5.7 External Validation on Pitt Corpus

We applied the text pipeline to 549 DementiaBank Pitt Corpus transcripts (243 control, 306 dementia; ∼ 14 × as many as our sample) and the audio pipeline to the 292 with usable recordings. We report Cohen's d for consistency with prior DementiaBank literature, and note that at this sample size Hedges’ g and Cohen's d are numerically equivalent up to two decimals. Text generalizes robustly. Of the eleven text models from our primary analysis, ten are reported here (XLNet omitted due to individual non-significance). All 8 autoregressive models replicated the direction (dementia > control); 7/8 reached significance (mean d = 0.35, range 0.16–0.45), with the large-model cluster at d = 0.41–0.45. BERT and RoBERTa replicated in the opposite direction, which we attribute to pseudo-log-likelihood scoring interacting with Pitt's transcript format—consistent with our emphasis on autoregressive NLL as the more robust metric. The reduced magnitude (d ≈ 0.35 vs. our g ≈ 1.0+) is a boundary condition rather than a contradiction: with severity and etiology unavailable for both samples (Section 3) we cannot attribute the attenuation to composition, but Pitt's decades-long, multi-equipment collection is a heterogeneity source known to attenuate rather than invert effects. The central claim is direction rather than magnitude. Pitt is longitudinal: these 549 transcripts derive from 291 unique individuals, so repeat sessions are not independent and significance levels should be read as approximate. Audio marks the boundary of training-free generalization. Audio models applied to Pitt showed effects in the opposite direction from our primary sample (wav2vec2 d = −0.38, HuBERT d = −0.29; both pFDR < .05), while EnCodec remained null as designed (d ≈ 0). Decades of equipment heterogeneity disproportionately affects acoustic features exposed to speech models trained on modern recording standards. The text pipeline, invariant to recording conditions, generalizes; the audio pipeline requires standardized protocols.

6 Discussion

6.1 A Unified Account: Why the Content-vs-Mechanics Pattern Holds Up

Our training-free framework recovers, without any supervision, the clinically-observed gradient in which semantic and top-down control processes are disrupted more consistently than the basic machinery of production [8, 9, 28]. CLIP, Mistral-7B, and wav2vec2 have no knowledge of cortical anatomy, yet collectively they mark the same boundary between affected semantic and comparatively spared production processes in AD.

6.2 The Filler-Removal Result: What Supervised Approaches Risk

The filler-removal result bears directly on the surface-artifact concern of Section 2, which motivated the demographically balanced ADReSS benchmark [26]: classifiers may learn filler rate rather than cognitive markers. That explicitly encoding pause structure measurably improves supervised accuracy [42] makes the concern concrete—disfluency is exploitable signal. Filler rate differs between our groups (d = 0.94), so any classifier with a flexible objective could exploit it as a shortcut; a model trained on in-domain transcripts would have no reason not to. Our result runs the other way: removing fillers strengthens the effect. A training-free NLL approach instead scores transcripts under a pretrained language model whose surprise is shaped by general English—a distribution in which filler words are common—so the signal reflects content coherence, not fluency.

6.3 Fusion as a Probe of Separable Cognitive Sub-Processes

The fusion gains (Section 5.6) suggest that gaze, language, and linguistic speech probe at least partially separable cognitive sub-processes: narrative visual integration, linguistic semantic coherence, and lexical-phonemic retrieval. The errors are partially non-overlapping: cases that three-way fusion catches but single modalities miss account for the AUC gain from 0.780 (text alone, n = 38) to 0.938 (three-way).

6.4 Limitations of the Framework

The primary sample is single-site and small (n = 39). The gaze strategy carried into fusion was the strongest of ten in this same sample, so fusion effect sizes should be read as optimistic; the grid-search audit (Section 5.6) bounds how much this matters. The Pitt validation extends this for text but not audio, which reversed direction under recording heterogeneity (Section 5.7); no external gaze corpus exists. Control participants self-reported the absence of a dementia diagnosis and were not formally screened (Section 3), so undetected preclinical impairment in the control group cannot be excluded and would attenuate rather than inflate the reported separation. Our sample is demographically homogeneous (59% female, 73% White) and the pipeline is English-only, so cross-linguistic deployment requires per-population validation. Finally, per-participant severity and etiology labels were unavailable for both samples (Section 3), precluding analyses relating the NLL signal to disease stage or subtype. Gaze sampling density was lower in the dementia group at the margin of significance (p = .055); because fewer points yield a noisier rather than a shifted NLL estimate, this would attenuate rather than inflate the reported separation.

6.5 Implications for Multimodal Interaction

The content-vs-mechanics pattern reframes what multimodal interaction systems should measure. Much of current HCI assumes behavioral production—fluency of speech, smoothness of gaze, voice quality—carries the signal of cognitive state. Our results suggest this assumption is misplaced for dementia. Instead, the signal is in what users engage with, not only how they produce behavior. The concrete move for conversational and assistive systems is a shift from disfluency-based to coherence-based adaptation. Coherence-based flagging—whether verbal descriptions integrate across a scene, whether gaze samples narratively relevant regions—captures the actual locus of cognitive change.

7 Conclusion

We presented a training-free multimodal framework that scores behavioral typicality under frozen pretrained models across gaze, text, and audio. Models that never saw a dementia label separate the groups in gaze and text and reach AUC = 0.94 in three-way fusion, while mechanically-matched baselines do not—a pattern holding across ten of eleven language models, 549 external transcripts, and a 1,320-configuration audit. Because nothing in the pipeline was fitted to these groups, the separation is a property of the behavior rather than of a trained classifier, which is what makes the approach portable to settings where labeled cohorts do not exist. This is a screening signal, not a diagnostic, and the implication for multimodal interaction is concrete: adapt to what users engage with, not only how they produce behavior.

8 Safe and Responsible Innovation

This is a screening framework, not a diagnostic tool: false positives risk distress, false negatives delayed evaluation. Responsible deployment requires clinician oversight, multi-site validation, and demographic fairness audits. Our English-only pipeline inherits the priors of its pretrained components: speakers of under-represented dialects (AAVE, non-native or accented English) may register as more “atypical,” a pattern the framework reflects rather than creates. Addressing this limitation fundamentally requires broader linguistic coverage in pretraining, not just downstream correction.

9 Reproducibility

All models are frozen and publicly available; none is trained or fine-tuned on dementia labels. Scoring code and the analysis scripts producing every table and figure are released at https://github.com/leticiapinto/icmi2026-multimodal-typicality-dementia. The DementiaBank Pitt Corpus is available under DementiaBank's data use agreement; our recordings contain identifiable speech and gaze and cannot be redistributed, so we release derived per-participant NLL scores.

Acknowledgments

This work was supported by UPenn's Alzheimer's Disease Research Center, under the NIH “Penn Artificial Intelligence and Technology Collaboratory for Healthy Aging” subaward for An Accessible Machine Learning-Based ADRD Screening Tool for Families and Caregivers (Award 5-P30-AG-073105-02). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. This work was also supported by the Army Research Office under Award W911NF-25-2-0040. We thank the participants and their families for their time. We also thank the USC Viterbi research programs CURVE, SURE, SHINE, and VSI, and the students who contributed to data collection, analysis, and coding: Leslie Moreno, Gwen Bradforth, Cecily Chung, Riley Ashford, Terry Tao, and Julie Kim.

References

  • Alzheimer's Association. 2025. 2025 Alzheimer's Disease Facts and Figures. https://www.alz.org/alzheimers-dementia/facts-figures. Accessed: 2026-04-19.
  • Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations, In NeurIPS 2020. Advances in neural information processing systems 33, 12449–12460.
  • Aparna Balagopalan, Benjamin Eyre, Frank Rudzicz, and Jekaterina Novikova. 2020. To BERT or not to BERT: comparing speech and language-based approaches for Alzheimer's disease detection. In Interspeech 2020.
  • James T Becker, François Boller, Oscar L Lopez, Judith Saxton, and Karen L McGonigle. 1994. The natural history of Alzheimer's disease: description of study cohort and accuracy of diagnosis. Archives of neurology 51, 6 (1994), 585–594.
  • Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan. 2010. A theory of learning from different domains. Machine learning 79, 1 (2010), 151–175.
  • Yoav Benjamini and Yosef Hochberg. 1995. Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal statistical society: series B (Methodological) 57, 1 (1995), 289–300.
  • Sid Black, Gao Leo, Phil Wang, Connor Leahy, and Stella Biderman. 2021. Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow.
  • Trevor J Crawford, Steve Higham, Jenny Mayes, Mark Dale, Sandip Shaunak, and Godwin Lekwuwa. 2013. The role of working memory and attentional disengagement on inhibitory control: effects of aging and Alzheimer's disease. Age 35, 5 (2013), 1637–1650.
  • Bernard Croisile, Bernadette Ska, Marie-Josee Brabant, Annick Duchene, Yves Lepage, Gerard Aimard, and Marc Trillet. 1996. Comparative Study of Oral and Written Picture Description in Patients with Alzheimer's Disease. Brain and Language 53, 1 (1996), 1–19. https://doi.org/10.1006/brln.1996.0033
  • Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2022. High fidelity neural audio compression. arXiv preprint arXiv:2210.13438 (2022).
  • Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 4171–4186.
  • Kathleen C Fraser, Jed A Meltzer, and Frank Rudzicz. 2016. Linguistic features identify Alzheimer's disease in narrative speech. Journal of Alzheimer's Disease 49, 2 (2016), 407–422.
  • Julian Fritsch, Sebastian Wankerl, and Elmar Nöth. 2019. Automatic diagnosis of Alzheimer's disease using neural network language models. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 5841–5845.
  • Harold Goodglass, Edith Kaplan, and Barbara Barresi. 2001. BDAE: The Boston Diagnostic Aphasia Examination (3 ed.). Lippincott Williams & Wilkins, Philadelphia, PA.
  • Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition. 770–778.
  • Larry V Hedges. 1981. Distribution theory for Glass's estimator of effect size and related estimators. journal of Educational Statistics 6, 2 (1981), 107–128.
  • Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM transactions on audio, speech, and language processing 29 (2021), 3451–3460.
  • Shih-Cheng Huang, Liyue Shen, Matthew P. Lungren, and Serena Yeung. 2021. GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-Efficient Medical Image Recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 3942–3951.
  • Mojan Javaheripi, Sébastien Bubeck, Marah Abdin, Jyoti Aneja, Caio César Teodoro Mendes, Weizhu Chen, Allie Del Giorno, Ronen Eldan, Sivakanth Gopi, Suriya Gunasekar, Piero Kauffmann, Yin Tat Lee, Yuanzhi Li, Anh Nguyen, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Michael Santacroce, Harkirat Singh Behl, Adam Taumann Kalai, Xin Wang, Rachel Ward, Philipp Witte, Cyril Zhang, and Yi Zhang. 2023. Phi-2: The surprising power of small language models. Microsoft Research Blog. https://www.microsoft.com/en-us/research/blog/phi-2-the-surprising-power-of-small-language-models/
  • Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023).
  • Yue Jiang, Zixin Guo, Hamed Rezazadegan Tavakoli, Luis A Leiva, and Antti Oulasvirta. 2024. EyeFormer: predicting personalized scanpaths with transformer-guided reinforcement learning. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–15.
  • Alexandra König, Aharon Satt, Alexander Sorin, Ron Hoory, Orith Toledo-Ronen, Alexandre Derreumaux, Valeria Manera, Frans Verhey, Pauline Aalten, Phillipe H Robert, et al. 2015. Automatic speech analysis for the assessment of patients with predementia and Alzheimer's disease. Alzheimer's & Dementia: Diagnosis, Assessment & Disease Monitoring 1, 1 (2015), 112–124.
  • Changye Li, David Knopman, Weizhe Xu, Trevor Cohen, and Serguei Pakhomov. 2022. GPT-D: Inducing dementia-related linguistic anomalies by deliberate degradation of artificial neural language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1866–1877.
  • Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019).
  • Saturnino Luz, Fasih Haider, Sofia de la Fuente, Davida Fromm, and Brian MacWhinney. 2021. Detecting cognitive decline using speech only: The ADReSSo challenge. arxiv:2104.09356 [cs.CL]
  • Saturnino Luz, Fasih Haider, Sofia de la Fuente Garcia, Davida Fromm, and Brian MacWhinney. 2021. Alzheimer's Dementia Recognition through Spontaneous Speech: The ADReSS Challenge. Frontiers in computer science 3 (2021), 780169.
  • Geoffrey J McLachlan, Sharon X Lee, and Suren I Rathnayake. 2019. Finite mixture models. Annual Review of Statistics and Its Application 6, 1 (2019), 355–378.
  • Robert J Molitor, Philip C Ko, and Brandon A Ally. 2015. Eye movements in Alzheimer's disease. Journal of Alzheimer's disease 44, 1 (2015), 1–12.
  • Michael Moor, Oishi Banerjee, Zahra Shakeri Hossein Abad, Harlan M. Krumholz, Jure Leskovec, Eric J. Topol, and Pranav Rajpurkar. 2023. Foundation Models for Generalist Medical Artificial Intelligence. Nature 616, 7956 (2023), 259–265.
  • National Institute on Aging. 2022. How Is Alzheimer's Disease Diagnosed?https://www.nia.nih.gov/health/alzheimers-symptoms-and-diagnosis/how-alzheimers-disease-diagnosed. Content reviewed December 8, 2022; accessed 2026-04-19.
  • Ivanna M Pavisic, Nicholas C Firth, Samuel Parsons, David Martinez Rego, Timothy J Shakespeare, Keir XX Yong, Catherine F Slattery, Ross W Paterson, Alexander JM Foulkes, Kirsty Macpherson, et al. 2017. Eyetracking metrics in young onset Alzheimer's disease: a window into cognitive visual functions. Frontiers in neurology 8 (2017), 377.
  • Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PmLR, 8748–8763.
  • Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning. PMLR, 28492–28518.
  • Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9.
  • Maithra Raghu, Chiyuan Zhang, Jon Kleinberg, and Samy Bengio. 2019. Transfusion: Understanding Transfer Learning for Medical Imaging. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 32. 3347–3357. https://proceedings.neurips.cc/paper_files/paper/2019/file/eb1e78328c46506b46a4ac4a1e378b91-Paper.pdf
  • Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In Proc. ICCV. 618–626. https://doi.org/10.1109/ICCV.2017.74
  • Alex Wang and Kyunghyun Cho. 2019. BERT has a mouth, and it must speak: BERT as a Markov random field language model. In Proceedings of the workshop on methods for optimizing and evaluating neural language generation. 30–36.
  • Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  • Yaqing Wang, Quanming Yao, James T Kwok, and Lionel M Ni. 2020. Generalizing from a few examples: A survey on few-shot learning. ACM computing surveys (csur) 53, 3 (2020), 1–34.
  • Yao Xiao, Heidi Christensen, and Stefan Goetze. 2025. Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models. arXiv preprint arXiv:2506.09315.
  • Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems 32.
  • Jiahong Yuan, Yuchen Bian, Xingyu Cai, Jiaji Huang, Zheng Ye, and Kenneth Church. 2020. Disfluencies and Fine-Tuning Pre-Trained Language Models for Detection of Alzheimer's Disease.. In Interspeech, Vol. 2020. 2162–6.
  • Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068 (2022).
  • Sheng Zhang, Yanbo Xu, Naoto Usuyama, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, Cliff Wong, Matthew P. Lungren, Tristan Naumann, and Hoifung Poon. 2023. BiomedCLIP: a Multimodal Biomedical Foundation Model Pretrained from Fifteen Million Scientific Image-Text Pairs. arXiv preprint arXiv:2303.00915 (2023).

Footnote

Code and derived per-participant scores are available at https://github.com/leticiapinto/icmi2026-multimodal-typicality-dementia.

CC-BY non-commercial, no derivatives license image
This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

ICMI '26, Napoli, Italy

© 2026 Copyright held by the owner/author(s).
ACM ISBN 979-8-4007-2318-6/26/10.
DOI: https://doi.org/10.1145/3776574.3831111