Validation of a Multi-level Self-Report Rapport Scale and its Impact on Multimodal Modeling of Small Group Interaction

Justine Reverdy, ALMAnaCH, Inria, Paris, France, justine.reverdy@inria.fr
Oussama Silem, ALMAnaCH, Inria, Paris, France, oussama.silem@inria.fr
Justine Cassell, ALMAnaCH, Inria, Paris, France, justine.cassell@inria.fr

To investigate how personality traits and multimodal behaviors shape small-group rapport, we introduce the French Self-reported Group and Dyadic Rapport (SGDR) scale. Study 1 validates the SGDR on 30 face-to-face tetrads (groups of four) engaged in collaborative tasks. The results demonstrate high internal consistency for both its dyadic and group subscales, alongside convergent validity with the Inclusion-of-Other-in-Self scale and task satisfaction.

Study 2 investigates the impact of the OCEAN personality traits on rapport, revealing that dyadic and group rapport rely on distinct mechanisms. Dyadic rapport is driven by the target's (interlocutor's) personality: agreeableness predicts it positively, conscientiousness negatively. Conversely, group-level rapport is not significantly explained by either a rater's own traits or the mean levels of personality traits across a tetrad. It is instead anchored in the rater's most extreme dyadic rapport ratings, with both the best and the worst of these ratings independently predicting the rating of group rapport. Notably, conscientiousness operates in opposite directions across levels: it negatively predicts rapport when it is the target's trait at the dyadic level, and positively predicts rapport at the group level (though only as a trend) when it is the rater's own trait. Study 3, on a subset of 16 tetrads, integrates multimodal behaviors. For dyadic rapport, pitch synchrony emerges as the strongest negative predictor, and this association is moderated by rater's openness. Group rapport again shows no significant effects of personality, while tetrad-level synchrony accounts for more of its variance (even if the association remains a trend). These findings demonstrate that rapport is a complex multimodal multi-level phenomenon and that group rapport is not reducible to the sum of dyadic interactions, validating the need for a multi-level scale, such as the SGDR.

CCS Concepts:Human-centered computing → Collaborative interaction; • General and reference → Measurement;

Keywords: psychometric scale, group collaboration, rapport, nonverbal behavioral analysis

ACM Reference Format:
Justine Reverdy, Oussama Silem, and Justine Cassell. 2026. Validation of a Multi-level Self-Report Rapport Scale and its Impact on Multimodal Modeling of Small Group Interaction. In INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION (ICMI '26), October 05--09, 2026, Napoli, Italy. ACM, New York, NY, USA 15 Pages. https://doi.org/10.1145/3776574.3831156

1 Introduction

Rapport, defined as a sense of mutual connection and harmony between individuals, has been considered a key component of both social and task interactions [3, 57, 61]. It emerges from the fine-grained temporal coordination of behavioral signals [46]. While its positive impact on dyadic communication is well-established [6, 12, 42], for humans and conversational agents [33], its role in small groups remains underexplored, overlooking the multi-leveled structure of group dynamics where individuals may address another member or the whole group [25]. As conversational agents transition from personal assistants into multi-party environments, developing robust, interpretable measures of group rapport becomes as important as it is for single-user dialogue systems [13]. Group settings introduce additional complexities as they may emerge through intra-group dyadic and whole-group relationships, while the group composition itself (i.e., the different levels of personality traits within a tetrad) may play a role. This multi-leveled structure raises fundamental questions about how to measure rapport.

We therefore investigate self-reported rapport in groups of four participants (tetrads) engaged in collaborative conversations, based on tasks used in [31, 48]. To this end, building on existing rapport questionnaires [6, 22], we design a post-task questionnaire that captures both dyadic and group-level perceptions. To round out the picture, we collect responses to the Inclusion-of-Other-in-Self scale [2], and ask about task satisfaction. Following work by [31], we also collect self-reported personality traits before the participants gather, via the IPIP-NEO-60 OCEAN inventory online [26, 44, 45] allowing us to see how these individual differences relate to rapport both at dyadic and group levels.

Our analysis proceeds to: (1) a psychometric validation using factor analysis and multi-level modeling; (2) an exploration of the links between personality tests and rapport; and (3) a multimodal exploration using nested model comparisons (likelihood-ratio tests) to assess the respective predictive power of personality and audio-visual features.

2 Related Work

The study of rapport and, more broadly, the pragmatics of social interaction, spans psychology, linguistics, and computational science, yet there exists a gap in how these constructs translate from dyads to groups.

Rapport and Synchrony in Multimodal Group Modeling. Interpersonal rapport is indexed by ways of speaking, nonverbal devices, and behavioral synchrony [6, 14, 57, 61]. The OCEAN personality traits of interlocutors (namely, Openness to Experience, Conscientiousness, Extraversion, Agreeableness, and Neuroticism) have also been shown to play a role [30], with extraversion and agreeableness emerging as the strongest predictors in human-agent conversation [13, 16]. However, the vast majority of this research has focused on dyadic settings [6, 66, 71], and has often focused on either nonverbal behavior or spoken language (both content and acoustic features), but not both. Additionally, prior work demonstrates that processes observed between two people do not necessarily generalize to multi-party settings [25, 38]. On the other hand, numerous studies have explored various other aspects of small group collaboration, such as creative problem solving [53], influential statements [48], curiosity in learning contexts [51], group engagement [49, 55], or group cohesion [34]. However, other phenomena, such as group cohesion, often defined in terms of alignment in attitudes or opinions, do not fully capture rapport, which instead reflects the moment-to-moment signaling of interpersonal connection and harmony during interaction. It therefore remains important to model group rapport.

In this context, [22] links rapport to interactional synchrony and group performance, highlighting its functional role in collaborative settings, and proposes a group rapport scale based on earlier dyadic measures [8]; however, she does not explore the relationship between dyadic and group rapport in groups, nor the impact of personality traits. Building on [22], who found that non-verbal behavioral synchrony was predictive of group rapport, our work adapts a compact subset of her scale's group rapport items to suit the analysis of a dataset that also examines the impact of personality (see section 4).

From another angle, [47] has demonstrated that multimodal behavioral features can detect low dyadic rapport in small groups. Collecting self-reported rapport at the dyadic level, they found the best predictor to be facial expressions. However, their work does not examine the relationship between dyadic and group rapport, nor does it investigate how psychological traits such as personality influence rapport signaling.

Evaluation of Rapport. How rapport is measured differs widely. Some studies rely on self-report questionnaires pertaining to an entire interaction [8, 13]. Others have relied on “thin-slice" observer judgments [1, 29, 41], where video is divided into short segments, and each is annotated by a naive observer.

The 18-item scale of Bernieri [6], originally developed for inter-human interaction, has served as a common reference point for computational models that have relied on self-report [28, 30, 47, 68]. Within these different perspectives, scales tailored to specific interaction contexts have also been proposed, including for teacher-student [69], investigative interviewing [18], and therapy [43]. In human-machine interaction research, dedicated rapport scales have further been introduced [13, 39, 40, 67].

These self-report scales are motivated by work that suggests that external observers may fail to capture the subjective experience of interactants [8, 13, 50], and highlights a mismatch between ‘how participants describe their feelings’ and ‘how observers perceive them’ [17, 21, 50]. For example, [8, 13] reported low consistency between self-reported and observer assessment of rapport, as well as different behavioral correlates across the two perspectives. However, an alternative perspective argues that people may access their internal states largely through their behavior, and that such access can fade by the end of an interaction [5], with end-of-task judgments disproportionately shaped by initial and final impressions [20, 58].

Rapport and Personality in Groups. Much research has examined the role of the five OCEAN personality traits in group outcomes. In fact, a meta-analysis has shown that team personality composition influences performance — often assessed through aggregated supervisor ratings of team effectiveness across task-related dimensions — with higher and more homogeneous levels of agreeableness and Conscientiousness predicting better team outcomes [52]. [9] also highlights the importance of group personality composition for both individual- and group-level outcomes, as coded by observers who rated group behaviors such as conflict, cooperation, communication, and coordination. However, personality remains insufficiently integrated into frameworks that jointly consider multimodal social signals and multi-level perceptions of rapport. In addition, existing approaches often treat personality as a holistic predictor, rather than modeling the different traits and how they interact with behavioral signals to shape emergent group processes such as rapport.

A Multi-Level Approach. Our work addresses the above limitations by examining self-reported dyadic and group-level rapport within the same small-group setting, and by focusing on the impact of observable behaviors as well as the underlying psychological factors of personality traits.

To the best of our knowledge, this is the first scale that proposes to integrate this multi-level approach of rapport into a relatively short questionnaire. By then integrating personality traits and multimodal behavioral synchrony into our analysis, we can investigate how mechanisms that have been shown to underlie dyadic rapport may extend to the group-level —an important extension to current literature. This approach allows us to validate self-report scales while bridging the gap between fine-grained temporal behavioral signals and the underlying psychological traits of personality and rapport.

3 Dataset

The dataset1 consists of groups of four native or near-native French-speaking participants (tetrads) engaging in three verbal tasks designed to elicit collaborative conversations. It is designed to be comparable to the MATRICS corpus recorded in Japanese [31, 48], in order to allow future cross-cultural/cross-linguistic comparisons (small changes were made in order to capture the upper bodies on video). Tetrads participated in three 20-minute tasks: (1) Guest Selection: choosing and scheduling five celebrity guests for a festival and outlining their stage performances; (2) Booth Planning: designing a plan for six food booths distributed across key festival areas; and (3) Travel Planning: organizing a two-day sightseeing itinerary in Paris for a friend, including visits and rest periods.

Recording procedure and setup. At the beginning of each task, written instructions were distributed, then also given verbally by the experimenters, allowing participants to ask questions. Participants were asked to reach a consensus on each task. Before and after each task, participants engaged in 7 minutes of free talk. Video was recorded on four GoPro cameras, and audio was collected with four wireless microphones mounted on headsets. Participants sat around a small table in a quiet room with cameras facing each participant, placed at chest level to ensure they could freely move their upper bodies. Participants received €25 as compensation.

Demographics and subset. 30 tetrads constitute the data for Studies 1 and 2, with a subset of 16 tetrads for Study 3. All groups were same-gender, with an equal number of male and female tetrads. Participants were recruited through a local online research site. Mean participant age was 28 (σ = 9.94, min = 18). Participants were not acquainted beforehand.

4 Scale development

The goal here is to construct a scale that measures both group and dyadic rapport between every pair of participants within a group. The dyadic rapport scale was developed based on [8]. However, the original 18-item scale was not suitable here, where each participant evaluated their level of rapport with every other group member as well as for the group. To limit fatigue while preserving reliability [19], we therefore initially reduced the dyadic rapport scale to 6 items, selecting two for each core component of rapport as conceptualized in [61]—(positivity: Q1, Q2; coordination: Q3–Q4; attentiveness: Q5–Q6). Preliminary analyses revealed insufficient internal consistency (see Study 1 in section 5), and thus the scale was extended with one additional item per component from the original scale, resulting in a 9-item dyadic rapport scale. For group rapport, we relied on prior work by Fultz [22], who proposed a 19-item scale for groups, grounded in rapport theory and the [8] dyadic scale. The validation of the scale in [22] identified two main components through PCA: rapport and compliance with group wishes. We retained a subset of 6 items, with strong loadings on the rapport component (Positivity, Coordination, and Attention). More questions were retained for attentiveness than for coordination to avoid inadvertently capturing group-task functioning instead of rapport, as per [22]. Additionally, we incorporated 3 items to assess task satisfaction, as strong loadings between task satisfaction items and group rapport were reported in [22]. This resulted in a final 9-item group-level scale (6 group rapport and 3 task satisfaction items). To account for potential acquiescence effects [32] in responses, as well as for subsequent validity checks [37, 65], we reverse-coded Q2 and Q4 in the dyadic rapport scale, and Q3 and Q5 in the group scale. All items were rated on a 7-point Likert scale.

To assess convergent validity, we also collected measures of interpersonal closeness using the Inclusion-of-Other-in-Self (IOS) scale [2], which has been shown to correlate with constructs such as liking and intimacy [22, 56], as well as rapport [23]. Participants completed the scale for each partner and then an adapted group-level version, the Inclusion of Ingroup in the Self (IIS) scale for group evaluations [22, 63].

The items from the dyadic and group rapport scales were translated into French by the authors and subsequently reviewed by two native French speakers for clarity and linguistic accuracy (see Appendix A for full questionnaires). Participants recorded responses to all scales on computers right after the interaction.

5 Study 1: Multi-level Psychometric Evaluation of Dyadic and Group Rapport Scales

Following established procedures in psychometric scale validation, we assess the internal consistency of the dyadic and group rapport scales and their convergent validity with task satisfaction, the Inclusion of Other in the Self (IOS) and Inclusion of Ingroup in the Self (IIS) measures2. Dyadic and group scales were analyzed separately within a multi-level framework. Sample sizes are consistent with respondent-to-item heuristics for both scales at their respective units of analysis [10] (see footnote 3). An exploratory factor analysis (EFA) was conducted separately for the two scales to examine their underlying structure. Internal consistency was assessed using Cronbach's α and McDonald's ω [54, 59]. In order to ensure comparability, internal consistency was subsequently re-evaluated.3 Convergent validity was assessed by estimating linear mixed-effects models predicting Dyadic IOS, Group IIS, and Task Satisfaction from rapport scores. Spearman correlation coefficients were computed to examine associations among study variables, after averaging item-level scores into construct-level means for each tetrad and individual at both the group and dyadic levels, merging these aggregates, and using pairwise complete observations. To further investigate interpersonal dynamics, dyadic reciprocity was calculated using Spearman correlations between mutual rapport ratings (rater-to-target and target-to-rater) across the 180 dyads within the 30 tetrads. Finally, to account for the nested structure of the data, a linear mixed-effects model with tetrad as a random intercept was used to examine the effect of gender on dyadic rapport.

5.1 Results

Table 1: Descriptive statistics, reliability, and correlations between group and dyadic measures.
Var. M SD α ω 1 2 3 4 5 6 7
1. G-IIS 5.19 1.33 1.00
2. G-Rap6 6.05 .82 .80 .82 .42** 1.00
3. G-Rap5 6.03 .88 .84 .85 .45** .94** 1.00
4. Task Sat. 6.25 .92 .85 .86 .27** .49** .45** 1.00
5. D-IOS 4.62 1.36 .66** .31** .34** .28** 1.00
6. D-Rap6 5.89 .65 .78 .79 .53** .64** .61** .51* .50* 1.00
7. D-Rap9 5.71 .81 .91 .91 .48** .58** .66** .31** .44** 1.00
Note. G-IIS = Inclusion of Ingroup in the Self; G-Rap6 = 6-item Group Rapport; G-Rap5 = 5-item Group Rapport; D-IOS = Dyadic IOS; D-Rap6 = 6-item Dyadic Rapport; D-Rap9 = 9-item Dyadic Rapport. Scale range: 1–7. G-Rap5 excludes chemistry. Spearman correlations reported. Correlation between D-Rap6 and D-Rap9 unavailable because scales were collected in different experimental phases. *p < .05, **p < .01.

Exploratory factor analyses supported a unidimensional structure for both the dyadic and group rapport scales. Detailed loadings are given in Appendix  C. For the dyadic scale, all items loaded positively on a single factor, with loadings ranging from .40 to .86 for the 6-item version and from .38 to .85 for the 9-item version. The 9-item version exhibited a stronger and more consistent structure, accounting for 52% of variance compared to the 6-item version (40%). For the group rapport scale, chemistry showed weak loadings and low communality, meaning that this dimension of group rapport was not perceived similarly by respondents. Removal resulted in a clearer factor structure, and the 5-item version was retained for subsequent analyses.

As shown in Table 1, internal consistency was acceptable for the 6-item dyadic scale (α = .78, ω = .79) and excellent for the 9-item version (α = .91, ω = .91). The refined 5-item group scale also demonstrated good internal consistency (α = .84, ω = .85). We note the high means and low standard deviations for both dyadic and group rapport ratings, compared to slightly lower means and higher standard deviations for dyadic and group IOS. Spearman correlational analyses revealed a moderately strong statistically significant (rs = .66) association between both the refined rapport measures, indicating that rapport at different interaction levels is related but not redundant. Results from linear mixed-effects models (Table 2) showed that dyadic rapport was significantly associated with IOS (β = .75, p < .001), while group rapport was significantly associated with IIS (β = .32, p = .02). Group rapport was also a significant positive predictor of task satisfaction (β = .73, std. β = .69, p < .001). Group rapport and task satisfaction were nonetheless only moderately correlated (rs = .45; Table 1), sharing roughly a fifth of their variance and thus remaining empirically distinct rather than redundant. Dyadic reciprocity analyses indicated a weak and non-significant association between mutual ratings within dyads (Spearman's ρ = .11, p = .15). Furthermore, gender did not significantly predict dyadic rapport (β = −.05, p = .773). A null mixed-effects model revealed an intra-class correlation coefficient (ICC) of .253, indicating that 25.3% of the variance in dyadic rapport ratings occurred between tetrads rather than within tetrads. This suggests that dyadic rapport ratings were not independent of the specific tetrad in which participants interacted.

Table 2: Convergent validity results from linear mixed-effects models.
Outcome Pred. β Std. β SE 95% CI p
Group IIS G. Rapport .32 .21 .14 [.05, .59] .020
Dyadic IOS D. Rapport .75 .43 .14 [.46, 1.03] <.001
Task Sat. G. Rapport .73 .69 .07 [.59, .86] <.001
G. Rapport D. Rapport .59 .52 .09 [.41, .76] <.001
D. Rapport Gender (M) -.05 -.07 .16 [-.37, .27] .773

5.2 Discussion

These findings provide initial support for the psychometric properties of the proposed dyadic and group rapport scales. First, the factor analyses support a one factor conceptualization of rapport, echoing previous findings [7]. The addition of items in the dyadic scale and the removal of a weak item in the group scale resulted in clearer factor structures and improved internal consistency.

Second, the convergent validity results provide evidence for a hierarchical structure across dyadic and group levels of interaction. Dyadic rapport was strongly associated with IOS and group rapport was moderately associated with IIS and predicted task satisfaction well, while remaining only moderately correlated with it. This pattern suggests that group rapport may capture aspects of collective interaction quality and coordination in addition to interpersonal closeness [63].

Additional analyses found no gender effect. Crucially, rapport in this setting is not reciprocal at the dyadic level. Again, the substantial variance attributable to tetrad-level clustering indicates that rapport is partly shaped by the specific group dynamic present in each tetrad.

This highlights the importance of accounting for group-level dynamics when studying rapport in multi-party interactions.

6 Study 2: Rapport and Personality Traits

Building on the validated dyadic and group rapport measures, this study investigated whether individual differences in personality traits from the Big Five (OCEAN) predict self-reported rapport. Personality traits were assessed using a 60-item adaptation drawn from the Canadian-French IPIP-NEO-300 [27]. See Appendix B for details and the full questionnaire. Internal consistency estimates for the five personality scales are reported in Appendix C.2.

6.1 Methodology

Linear mixed-effects models were used to examine the relationship between OCEAN personality traits and rapport at both the dyadic and group levels. In all models, the distinction between rater (the participant rating another participant) and target (the participant being rated) was maintained.

For dyadic rapport, fixed effects included both rater and target scores on all five OCEAN traits, with crossed random intercepts for session (tetrad), rater, and target. Crossed random rater and target intercepts allow the model to separate a participant's general tendency to report high or low rapport from their general tendency to be rated highly or poorly by others.

For group rapport, each rater's own report of overall group rapport was retained as an individual observation, so that a rater's traits and the compositional average of their tetrad's personality traits could be estimated as separate, simultaneous predictors. Two complementary specifications were tested against a common null model (random intercept for session only): a composition model, in which each rater's own OCEAN scores and their tetrad's mean OCEAN scores jointly predict group rapport; and a link-anchor model, in which each rater's own OCEAN scores are combined with the maximum and minimum dyadic rapport score they reported across their partners in the same tetrad, motivated by the peak-end rule governing the retrospective evaluation of multi-episode experiences [20].

Nested model comparisons were evaluated via likelihood-ratio tests, with models refit via maximum likelihood for this purpose and via restricted maximum likelihood for parameter estimation. The sample was of N = 360 directed dyads within 30 tetrads for the dyadic model, and N = 120 raters within the same 30 tetrads for the two group-level models.

6.2 Results

6.2.1 Dyadic Rapport. Adding the OCEAN fixed effects significantly improved model fit for dyadic rapport relative to the random-effects-only baseline (χ2(10) = 22.27, p = .014), capturing a marginal variance of $R^2_m = .080$ (Table 3). The targets’ personality traits were the primary drivers of dyadic rapport: targets scoring higher on Agreeableness were rated with significantly higher rapport (β = .126, p = .013), while targets scoring higher on Conscientiousness were rated with significantly lower rapport (β = −.137, p = .010; Table 4). Rater-level traits showed a weaker pattern: rater Extraversion showed a positive trend (β = .160, p = .076), while no other rater-level trait reached significance.

6.2.2 Group Rapport: Composition and Anchoring Effects. Neither a rater's own OCEAN traits nor the tetrad-level mean personality traits (composition) significantly predicted group rapport (χ2(10) = 8.43, p = .587; $R^2_m = .073$); indeed, no individual predictor in this model approached significance.

By contrast, the link-anchor model was highly significant (χ2(7) = 39.43, p < .001; $R^2_m = .283$). Both the maximum and minimum dyadic rapport a rater experienced within their tetrad significantly and positively predicted their group rapport rating (max: β = .388, p = .022; min: β = .248, p = .011; Table 4). The rater's own Conscientiousness showed a positive, non-significant trend in this model (β = .115, p = .185).

Table 3: Model comparisons for Study 2: personality predictors of dyadic and group rapport.
Context / Model AIC BIC $R^2_m$ p
Dyadic Rapport
   OCEAN Full 922.37 980.66 .080 .014*
Group Rapport
   Composition 355.62 391.86 .073 .587
   Link-Anchor 318.62 346.50 .283 < .001***
Note. p-values from likelihood-ratio tests against a random-intercepts-only null model. *p < .05, **p < .01, ***p < .001.
Table 4: Notable OCEAN predictors of dyadic rapport (Model A) and the group link-anchor model (Model C). Full parameter tables for all three models, including the group-composition model (B), are given in Appendix  C.
Model Predictor β SE p
A. Dyadic Rater Extraversion .160 .090 .076
Rapport Target Agreeableness .126 .050 .013*
Target Conscientiousness -.137 .052 .010*
C. Group Rater Conscientiousness .115 .086 .185
Link-Anchor Max Dyadic Rapport .388 .167 .022*
Min Dyadic Rapport .248 .096 .011*
*p < .05, **p < .01, ***p < .001.

6.3 Discussion

These results indicate that dyadic and group rapport are governed by distinct mechanisms. At the dyadic level, rapport is shaped primarily by the target's personality, as one might expect: participants high in Agreeableness may have facilitated rapport, consistent with Agreeableness's established role in fostering positive impressions in initial dyadic interactions [16]. Target Conscientiousness, on the other hand, limited rapport. Participants that scored higher in Conscientiousness thus may appear more rigid and therefore be evaluated less positively in this collaborative setting.

At the group level, personality, whether captured by a rater's traits or by the tetrad-level mean personality traits does not appear to explain group rapport, suggesting that group-level rapport cannot be straightforwardly mapped onto the aggregation of mean individual levels of OCEAN personality traits. Instead, group rapport is strongly shaped by the extremes of a rater's dyadic experiences within the tetrad: both their most positive and their most negative dyadic interaction independently and positively predict their overall sense of group rapport. This pattern is consistent with a bounded mental shortcut for evaluation [20]: group rapport appears to be bounded above by the best interaction experienced, and bounded below by whether any interaction fell beneath an acceptable relational threshold, rather than built up as an average of dyadic impressions.

The positive, though non-significant, association between rater Conscientiousness and group rapport may reflect rating the collective achievement of consensus goals over individual relational bonding, echoing evidence that Conscientiousness predicts positive team-level outcomes [52, 64].

This stands in contrast to Conscientiousness's negative association with dyadic rapport, where a conscientious partner may be perceived as task-focused rather than easy to bond with. This structural difference may imply that group-level evaluations are driven more by behavioral coordination than by targets’ personality traits.

7 Study 3: Personality and Non-verbal Behavioral Synchrony

Building on the validated rapport measures established in Study 1 and the personality effects identified in Study 2, we examine how non-verbal behavioral synchrony and individual personality differences might jointly contribute to rapport, and whether these associations generalize to unseen groups. Specifically, we asked three questions: (Q1) whether behavioral synchrony explains variance in rapport beyond personality traits, (Q2) whether personality traits retain predictive value once synchrony is considered, and (Q3) whether personality moderates the relationship between behavioral synchrony and rapport. The sample comprises 16 tetrads (64 participants, constituting 192 dyads), the subset of the full sample for which synchronized audio-visual recordings were processed.

7.1 Methodology

We modeled rapport as a function of both individual- and group-level predictors while accounting for the nested data structures. Visual and acoustic representations were extracted from the tetrads using both self-supervised embeddings and handcrafted features. VideoMAE [62] encoded spatiotemporal behavioral dynamics, while WavLM-Large [15] captured prosodic profiles. Facial Action Units and head pose were extracted using OpenFace [4], focusing on AU12 (smiling intensity) and nodding-related motion. Acoustic correlates of vocal arousal were obtained via fundamental frequency (F0) and intensity using Praat [11] through Parselmouth [35]. Synchrony was quantified via Pearson correlation on the behavioral time series (AU12, head pose, pitch, and intensity) and via cosine similarity for the embeddings, capturing actual coordination (see full Table 5). All six synchrony features were z-scored prior to modeling so that coefficients were directly comparable in scale to the OCEAN predictors.

For each outcome (dyadic and group rapport), four nested linear mixed-effects models were compared to address Q1 and Q2: a baseline model with random effects only (M0); a model adding the five OCEAN personality traits (M1); a model adding the six synchrony features (M2); and a full model combining both (M3). The dyadic models retained the crossed random-intercept structure for session, rater, and target. For the group models, each rater's own report of group rapport was retained as an individual observation, with a random intercept for tetrad, and each rater's own OCEAN scores entered as individual-level predictors. Synchrony, being a property of a dyad rather than of an individual, was instead aggregated at the tetrad level: each feature was averaged across all directed dyads within a session, giving a single tetrad-wide value per feature.

To address Q3, a fifth dyadic model (M4) extended the full model (M3) with interaction terms between pitch synchrony – the strongest and most consistent synchrony predictor across M2 and M3 – and the three OCEAN traits that emerged as significant predictors on their own in M1: rater Openness, target Agreeableness, and target Conscientiousness. This analysis was restricted to dyadic rapport; the group-level models, with only 16 tetrads, would not support interaction terms reliably.

Nested model comparisons were evaluated via likelihood-ratio tests, with all models fit via maximum likelihood for this purpose and refit via restricted maximum likelihood for the fixed-effect estimates reported below. To assess whether these associations generalize beyond the sample on which they were estimated, we additionally conducted leave-one-tetrad-out cross-validation (LOGO-CV): for each of the 16 tetrads in turn, all models were refit on the remaining 15 tetrads and used to predict rapport scores for the held-out tetrad. Prediction accuracy is summarized as the root mean squared error (RMSE) and the squared correlation between predicted and observed scores across all out-of-fold predictions.

Table 5: Multimodal Feature Set for Predicting Perceived Dyadic and Group Rapport
Feature Description Modality
Deep Video Sync Cosine similarity of VideoMAE embeddings; captures movement coordination. Visual
Deep Audio Sync Cosine similarity of WavLM-Large embeddings; captures prosodic entrainment. Audio
Facial Mimicry Correlation of AU12 intensity; measures shared smiles. Visual
Head Pose Sync Correlation of head pitch (Rx); captures nodding alignment. Visual
Pitch Sync Correlation of F0 (Hz); indicates pitch alignment. Audio
Intensity Sync Correlation of energy (dB); indicates shared arousal. Audio

7.2 Results

7.2.1 Dyadic Rapport (Q1, Q2). Personality and synchrony features both improved model fit for dyadic rapport, though to different degrees (Table 6). The OCEAN-only model was a marginal improvement (χ2(10) = 16.70, p = .081; $R^2_m=.111$); the synchrony-only (χ2(6) = 22.71, p < .001; $R^2_m=.082$) and full (χ2(16) = 34.35, p = .005; $R^2_m=.162$) models were both significant improvements over baseline.

Adding synchrony to the personality-only model significantly improved fit (χ2(6) = 17.65, p = .007); adding personality to the synchrony-only model did not (χ2(10) = 11.64, p = .310) – synchrony captures variance in dyadic rapport that personality does not.

At the coefficient level (Table 7), the OCEAN-only model reproduced Study 2’s target-level pattern: target Agreeableness positively predicted rapport (β = .150, p = .030) and target Conscientiousness negatively predicted it (β = −.156, p = .036). In addition, within this 16-tetrad multimodal subset, rater Openness emerged as a significant negative predictor (β = −.272, p = .029), although this effect was not observed in the full 30-tetrad personality analysis. Pitch synchrony was the strongest and only significant synchrony predictor, negative both alone (β = −.383, p < .001) and in the full model (β = −.328, p = .002). Intensity synchrony showed a positive trend alone (p = .103). Adding synchrony weakened all three personality effects to trend level.

7.2.2 Personality-Synchrony Interactions (Q3). Adding the three personality-by-pitch-synchrony interactions significantly improved fit over the additive model (χ2(3) = 10.00, p = .019; $R^2_m$: .162 → .183; Table 8). Only rater Openness significantly moderated the pitch-synchrony effect (β = −.126, p = .046); target Agreeableness (p = .527) and target Conscientiousness (p = .149) did not. At mean Openness, pitch synchrony was only a trend (β = −.205, p = .062); the interaction shows this effect is roughly four times stronger for high-Openness raters (predicted slope ≈ −.33 at + 1 SD) than for low-Openness raters (≈ −.08 at − 1 SD).

7.2.3 Group Rapport (Q1, Q2). Personality traits did not improve model fit for group rapport (χ2(5) = 3.25, p = .661; $R^2_m=.050$; Table 6; no individual trait reached significance, Appendix  C). Tetrad-level synchrony fared better, though still short of significance: it explained three times more variance than personality ($R^2_m=.154$ vs. .050) and approached significance against baseline (χ2(6) = 10.54, p = .104). The full model was not significant either (χ2(11) = 13.45, p = .265; $R^2_m=.192$), and neither incremental test reached significance (full vs. OCEAN: p = .117; full vs. synchrony: p = .713).

No individual predictor reached significance in either group model (Table 7), but two synchrony features were consistently the largest coefficients: deep video synchrony was negative (β = −.832, p = .153 alone; − .798, p = .202 in the full model) and head-pose synchrony was positive (β = .506, p = .232; .438, p = .329). Rater Conscientiousness showed a similarly consistent, non-significant positive trend (β = .196 in both models, p = .214 and .220).

7.2.4 Cross-Validation. Leave-one-tetrad-out cross-validation tested whether these associations generalize to unseen groups (Table 9). For dyadic rapport, no model beat the baseline (RMSE = .994, R2 = .154); the moderation model (M4) outperformed the additive full model, suggesting part of the interaction generalizes beyond the estimation sample. For group rapport, OCEAN and the full model again underperformed the baseline (RMSE = 1.149 and 1.175 vs. 1.105), while the synchrony-only model was the only model to beat the baseline on RMSE (1.075 vs. 1.105), though not on cross-validated R2 (.063 vs. .242). 4

Table 6: Model comparisons for Study 3: personality and synchrony as predictors of dyadic and group rapport.
Context / Model AIC BIC $R^2_m$ χ2 (df) p
Dyadic Rapport
   M0 Base 495.28 511.57 .000
   M1 OCEAN 498.58 547.45 .111 16.70 (10) .081
   M2 Synchrony 484.57 520.40 .082 22.71 (6) < .001***
   M3 Full 492.93 561.34 .162 34.35 (16) .005**
Group Rapport
   M0 Base 198.36 204.83 .000
   M1 OCEAN 205.10 222.37 .050 3.25 (5) .661
   M2 Synchrony 199.82 219.25 .154 10.54 (6) .104
   M3 Full 206.90 237.13 .192 13.45 (11) .265
Note. χ2 and p from likelihood-ratio tests against the corresponding M0 baseline. Group-level synchrony predictors are tetrad-wide averages. *p < .05, **p < .01, ***p < .001.
Table 7: Notable OCEAN and synchrony predictors of dyadic and group rapport (Study 3). For the group models, where no predictor reached significance, the two largest-magnitude coefficients are shown for transparency. Full parameter tables for all models are given in Appendix  C.
Model Predictor β SE p
Dyadic M1 Rater Openness -.272 .121 .029*
OCEAN Target Agreeableness .150 .067 .030*
Target Conscientiousness -.156 .073 .036*
Dyadic M2 Pitch Synchrony -.383 .099 < .001***
Synchrony Intensity Synchrony .187 .114 .103
Dyadic M3 Rater Openness -.236 .124 .061
Full Target Agreeableness .108 .066 .108
Target Conscientiousness -.106 .071 .141
Pitch Synchrony -.328 .103 .002**
Group M2 Deep Video Sync -.832 .574 .153
Synchrony Head Pose Sync .506 .419 .232
Group M3 Rater Conscientiousness .196 .158 .220
Full Deep Video Sync -.798 .618 .202
Head Pose Sync .438 .445 .329
*p < .05, **p < .01, ***p < .001.
Table 8: Personality × pitch-synchrony interactions predicting dyadic rapport (Model M4).
Predictor β SE p
Pitch Synchrony (main effect) -.205 .109 .062
Rater Openness × Pitch Synchrony -.126 .063 .046*
Target Agreeableness × Pitch Synchrony .036 .057 .527
Target Conscientiousness × Pitch Synchrony -.082 .056 .149
Note. Model comparison: M4 (with interactions) vs. M3 (additive only), χ2(3) = 10.00, p = .019. $R^2_m$: M3 = .162, M4 = .183. Full parameter table (all main effects) in Appendix  C. *p < .05.
Table 9: Leave-one-tetrad-out cross-validation (LOGO-CV) results.
Outcome Model RMSE R2
Dyadic M0 Base .994 .154
M1 OCEAN 1.026 .005
M2 Synchrony .999 .008
M3 Full 1.044 .006
M4 Moderation 1.033 .013
Group M0 Base 1.105 .242
M1 OCEAN 1.149 .011
M2 Synchrony 1.075 .063
M3 Full 1.175 .024
Note. RMSE and R2 are computed across all out-of-fold predictions; R2 is the squared correlation between predicted and observed values. All models converged on all 16 folds for both outcomes.

7.3 Discussion

Behavioral synchrony, particularly pitch synchrony, is thus associated with dyadic rapport beyond what personality alone captures (Q1), and this is not symmetrical: personality adds little once synchrony is included (Q2), though the target-level effects replicated in M1 and lost significance rather than magnitude in M3. The prominence of vocal pitch synchrony is consistent with prior work on acoustic-prosodic entrainment [70], but its direction is not: higher pitch synchrony predicted lower dyadic rapport. Yet, this direction is not unique to our data: pitch entrainment lowered trust in conversational avatars [24], and a meta-analysis also found vocal pitch synchrony negatively associated with therapeutic alliance [36]. A possible explanation is that it reflects interactional mechanisms specific to multi-party collaboration. Additionally, this relationship is not uniform across individuals (Q3): within the multimodal subset, rater Openness significantly moderated the association between pitch synchrony and rapport. Because rater Openness was not a significant main effect in the full 30-tetrad personality analysis, this moderation should be interpreted as preliminary until replicated in larger multimodal datasets.

At the group level, personality showed no association with rapport, alone or combined with synchrony, though the rater Conscientiousness trend from Study 2 persisted unchanged. Tetrad-level synchrony fared better across two independent tests: it explained more variance than personality in-sample ($R^2_m=.154$ vs. .050, p = .104) and was the only model, dyadic or group, to beat its own baseline on cross-validated RMSE, though not on cross-validated R2. Deep video synchrony (negative) and head-pose synchrony (positive) were the largest individual contributors, though neither reached significance alone, suggesting that further research is needed.

Given the small number of tetrads (16), the group-level patterns are suggestive rather than conclusive. The dyadic moderation finding rests on a larger sample (192 dyads) and is backed by the likelihood-ratio test, providing stronger evidence than the group-level trends. Overall, dyadic rapport appears shaped jointly by personality and behavioral coordination, while group rapport looks more tied to tetrad-wide coordination than to raters’ own personality traits.

8 Conclusion

We introduced and validated the Self-reported Group and Dyadic Rapport (SGDR) scale, a compact instrument measuring rapport at both the dyadic and the group level in small-group interaction. Across three studies (30 tetrads, with 16 in study 3’s multimodal subset), the scale showed solid reliability and meaningful links to interpersonal closeness and task satisfaction at both levels, supporting its use in multimodal research. Findings suggest that dyadic and group rapport may function differently. Dyadic rapport was associated with personality traits and behavioral synchrony: higher target Agreeableness and lower target Conscientiousness predicted higher rapport; when behavioral synchrony was added, pitch synchrony became the strongest predictor. Moreover, these two classes of predictors interact rather than being merely additive: the association between pitch synchrony and rapport was roughly four times stronger for raters high in Openness, a moderation supported by the likelihood-ratio test. Group rapport, by contrast, was explained neither by a rater's own traits nor by tetrad-level mean personality traits. It was instead explained by the extremes of the rater's dyadic experiences (the best and the worst dyadic interactions independently predicted the group rating), while tetrad-wide behavioral coordination showed a suggestive, not-yet-conclusive association that nonetheless outperformed personality both in explained variance and in generalization to unseen groups. Notably, Conscientiousness was significantly negative for dyadic rapport, and trended positive at the group level — a further indication that the two levels may be influenced by different mechanisms. Taken together, these findings caution against treating group rapport as a scaled-up version of dyadic rapport and support modeling the two levels separately. The SGDR provides a practical tool for doing so, and a step toward multimodal systems able to track rapport in task-oriented multi-party interaction.

9 Limitations & Future Work

The present findings should be interpreted in light of several limitations. While studies 1 and 2 included 30 tetrads, Study 3 relied on only 16, limiting the ability to detect more than moderate effects. Group-level findings that include nonverbal and paraverbal behaviors should thus be considered exploratory. The Openness moderation identified in Study 3 should also be interpreted with caution, as it was observed only in the 16-tetrad multimodal subset and not in the full 30-tetrad personality analysis. This discrepancy, together with the comparatively modest reliability of the Openness scale (see Appendix C), suggests that this finding is preliminary and requires replication in larger multimodal datasets.

Results presented within these three studies, moreover, rely on self-report completed at the end of the sessions, while externally assessed rapport annotations, currently underway, will allow self-perceptions to be triangulated with observer judgments. Finally, rapport ratings were generally high with restricted variance, and the dataset covers a single linguistic and cultural context; the planned comparison with the MATRICS corpus [31, 48] will help establish cross-cultural generalizability, as will a planned comparison with our Korean colleagues, also collecting an identical dataset.

10 Generative AI Use Disclosure

Generative AI tools (ChatGPT, Claude) were used for editing and polishing the manuscript. All scientific content, experimental design, and results were produced by the authors.

11 Safe and Responsible Innovation Statement

This work follows ethical and responsible research practices for multimodal interaction. In compliance with GDPR, data are securely stored, pseudonymized, and anonymized where possible, with risks such as personal disclosure mitigated through screening and controlled access. The study protocol, recruitment, and materials were approved by a local Institutional Review Board, and all participants provided informed consent after full information. We support transparent, privacy-preserving, and socially beneficial human–agent interaction while minimizing misuse risks. We also consider bias, inclusivity, and cross-cultural validity in the dataset design.

Acknowledgments

This work was supported by the Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korean government (MSIT) (RS-2022-II220043, Adaptive Personality for Intelligent Agents). Warm thanks to our research collaborators at KETI (Korean Electronics Technology Institute), Jeehyeong Kim, Mira Lee, and Hyoseon Kye, who have inspired many aspects of this research and have been of invaluable help in designing the dataset and working with us to prepare for a cross-cultural comparison. We also thank Professor Yukiko Nakano for her inspirational work on the Japanese MATRICS dataset and for her invaluable advice. Finally, we are grateful for the support of the other members of the ArticuLab group in the ALMAnaCH project-team for their assistance, and particularly master's student Adriano Rivierez for his help with data collection and processing, and lab manager Sophie Etling, for her support with the logistics of data collection.

References

  • Nalini Ambady and Robert Rosenthal. 1992. Thin slices of expressive behavior as predictors of interpersonal consequences: A meta-analysis.Psychological bulletin 111, 2 (1992), 256.
  • Arthur Aron, Elaine N Aron, and Danny Smollan. 1992. Inclusion of other in the self scale and the structure of interpersonal closeness.Journal of personality and social psychology 63, 4 (1992), 596.
  • Zachary G Baker, Emily M Watlington, and C Raymond Knee. 2020. The role of rapport in satisfying one's basic psychological needs. Motivation and emotion 44, 2 (2020), 329–343.
  • Tadas Baltrušaitis, Peter Robinson, and Louis-Philippe Morency. 2016. Openface: an open source facial behavior analysis toolkit. In 2016 IEEE winter conference on applications of computer vision (WACV). IEEE, 1–10.
  • Daryl J Bem. 1972. Self-perception theory. In Advances in experimental social psychology. Vol. 6. Elsevier, 1–62.
  • Frank J Bernieri. 1988. Coordinated movement and rapport in teacher-student interactions. Journal of Nonverbal behavior 12, 2 (1988), 120–138.
  • Frank J Bernieri, Janet M Davis, Robert Rosenthal, and C Raymond Knee. 1994. Interactional synchrony and rapport: Measuring synchrony in displays devoid of sound and facial affect. Personality and social psychology bulletin 20, 3 (1994), 303–311.
  • Frank J Bernieri, John S Gillis, Janet M Davis, and Jon E Grahe. 1996. Dyad rapport and the accuracy of its judgment across situations: a lens model analysis.Journal of Personality and Social Psychology 71, 1 (1996), 110.
  • Eva Bleckmann, Richard Rau, Oliver Lüdtke, Sascha Krause, and Jenny Wagner. 2026. How group personality composition affects person and group outcomes: An integrative analysis using the group actor–partner interdependence model.Journal of Personality and Social Psychology (2026).
  • Godfred O Boateng, Torsten B Neilands, Edward A Frongillo, Hugo R Melgar-Quiñonez, and Sera L Young. 2018. Best practices for developing and validating scales for health, social, and behavioral research: a primer. Frontiers in public health 6 (2018), 149.
  • Paul Boersma and David Weenink. 2020. Praat: Doing Phonetics by Computer [Computer program]. Version 6.1.38, retrieved from http://www.praat.org/.
  • Justine Cassell, Alastair Gill, and Paul Tepper. 2007. Coordination in conversation and rapport. In Proceedings of the workshop on Embodied Language Processing. 41–50.
  • Aleksandra Cerekovic, Oya Aran, and Daniel Gatica-Perez. 2016. Rapport with virtual agents: What do human social cues and personality explain?IEEE Transactions on Affective Computing 8, 3 (2016), 382–395.
  • Tanya L Chartrand and John A Bargh. 1999. The chameleon effect: The perception-behavior link and social interaction.Journal of Personality and Social Psychology 76, 6 (1999), 893–910.
  • Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al. 2022. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing 16, 6 (2022), 1505–1518.
  • Ronen Cuperman and William Ickes. 2009. Big Five predictors of behavior and perceptions in initial dyadic interactions: Personality similarity helps extraverts and introverts, but hurts “disagreeables”.Journal of personality and social psychology 97, 4 (2009), 667.
  • Sidney D'Mello, Ed Dieterle, and Angela Duckworth. 2017. Advanced, analytic, automated (AAA) measurement of engagement during learning. Educational psychologist 52, 2 (2017), 104–123.
  • Misty C Duke, James M Wood, Brock Bollin, Matthew Scullin, and Julia LaBianca. 2018. Development of the Rapport Scales for Investigative Interviews and Interrogations (RS3i), Interviewee Version.Psychology, Public Policy, and Law 24, 1 (2018), 64.
  • Gudrun Eisele, Hugo Vachon, Ginette Lafit, Peter Kuppens, Marlies Houben, Inez Myin-Germeys, and Wolfgang Viechtbauer. 2022. The effects of sampling frequency and questionnaire length on perceived burden, compliance, and careless responding in experience sampling data in a student population. Assessment 29, 2 (2022), 136–151.
  • Barbara L Fredrickson and Daniel Kahneman. 1993. Duration neglect in retrospective evaluations of affective episodes.Journal of personality and social psychology 65, 1 (1993), 45.
  • Kathryn A Fuller, Nilushi S Karunaratne, Som Naidu, Betty Exintaris, Jennifer L Short, Michael D Wolcott, Scott Singleton, and Paul J White. 2018. Development of a self-report instrument for measuring in-class student engagement reveals that pretending to engage is a significant unrecognized problem. PloS one 13, 10 (2018), e0205828.
  • Amber A. Fultz. 2023. The Relationships Between Synchrony, Rapport, and Small Group Performance. Doctoral dissertation. Oregon State University. https://ir.library.oregonstate.edu/concern/graduate_thesis_or_dissertations/47429j58kAdvisor: Frank J. Bernieri.
  • Simon Gächter, Chris Starmer, and Fabio Tufano. 2015. Measuring the closeness of relationships: a comprehensive evaluation of the'inclusion of the other in the self'scale. PloS one 10, 6 (2015), e0129478.
  • Ramiro H Gálvez, Agustín Gravano, Štefan Beňuš, Rivka Levitan, Marian Trnka, and Julia Hirschberg. 2020. An empirical study of the effect of acoustic-prosodic entrainment on the perceived trustworthiness of conversational avatars. Speech Communication 124 (2020), 46–67.
  • Erving Goffman. 1981. Forms of talk. University of Pennsylvania Press.
  • Lewis R Goldberg et al. 1999. A broad-bandwidth, public domain, personality inventory measuring the lower-level facets of several five-factor models. Personality psychology in Europe 7, 1 (1999), 7–28.
  • L. Gravel. 2001. French (Canadian) Translation of the IPIP Version of the NEO PI-R. https://ipip.ori.org/FrenchCanadian300-Item-IPIP-NEO.htm. Accessed: 2026-04-12.
  • Juan Lorenzo Hagad, Roberto Legaspi, Masayuki Numao, and Merlin Suarez. 2011. Predicting Levels of Rapport in Dyadic Interactions through Automatic Detection of Posture and Posture Congruence. In 2011 IEEE Third International Conference on Privacy, Security, Risk and Trust and 2011 IEEE Third International Conference on Social Computing. 613–616. https://doi.org/10.1109/PASSAT/SocialCom.2011.143
  • Judith A. Hall, Debra L. Roter, Danielle C. Blanch, and Richard M. Frankel. 2009. Observer-Rated Rapport in Interactions between Medical Students and Standardized Patients. Patient Education and Counseling 76, 3 (Sept. 2009), 323–327. https://doi.org/10.1016/j.pec.2009.05.009
  • Takato Hayashi, Ryusei Kimura, Ryo Ishii, and Shogo Okada. 2025. Investigating Role of Big Five Personality Traits in Audio-Visual Rapport Estimation. In 2025 IEEE 19th International Conference on Automatic Face and Gesture Recognition (FG). IEEE, 1–10.
  • Yuki Hayashi, Fumio Nihei, Yukiko Nakano, Hung-Hsuan Huang, and Shogo Okada. 2015. Construction of a Group Discussion Corpus and Analysis of its Relationship with Personality Traits. IPSJ Journal 56, 4 (2015), 1217–1227.
  • Andreas Hinz, Dominik Michalski, Reinhold Schwarz, and Philipp Yorck Herzberg. 2007. The acquiescence effect in responding to a questionnaire. GMS Psycho-Social Medicine 4 (2007), Doc07.
  • Lixing Huang, Louis-Philippe Morency, and Jonathan Gratch. 2011. Virtual Rapport 2.0. In Intelligent Virtual Agents, Hannes Högni Vilhjálmsson, Stefan Kopp, Stacy Marsella, and Kristinn R. Thórisson (Eds.). Springer, Berlin, Heidelberg, 68–79. https://doi.org/10.1007/978-3-642-23974-8_8
  • Hayley Hung and Daniel Gatica-Perez. 2010. Estimating cohesion in small groups using audio-visual nonverbal behavior. IEEE Transactions on Multimedia 12, 6 (2010), 563–575.
  • Yannick Jadoul, Bill Thompson, and Bart de Boer. 2018. Introducing Parselmouth: A Python interface to Praat. Journal of Phonetics (2018). praat-parselmouth version 0.4.7.
  • Simone Jennissen, Julia Huber, Beate Ditzen, and Ulrike Dinger. 2025. Association between nonverbal synchrony, alliance, and outcome in psychotherapy: systematic review and meta-analysis. Psychotherapy Research 35, 7 (2025), 1213–1228.
  • Krisztián Józsa and George A Morgan. 2017. Reversed items in Likert scales: Filtering out invalid responders. Journal of Psychological and Educational Research 25, 1 (2017), 7–25.
  • Nale Lehmann-Willenbrock, Hayley Hung, and Joann Keyton. 2017. New frontiers in analyzing dynamic group interactions: Bridging social and computer science. Small group research 48, 5 (2017), 519–531.
  • Ting-Han Lin, Guan Chen, Bilge Mutlu, J Gregory Trafton, and Sarah Sebo. 2026. The Reduced-Length Connection-Coordination Rapport (CCR) Scale. ACM Transactions on Human-Robot Interaction (2026).
  • Ting-Han Lin, Hannah Dinner, Tsz Long Leung, Bilge Mutlu, J Gregory Trafton, and Sarah Sebo. 2025. Connection-Coordination rapport (CCR) scale: a dual-factor scale to measure human-robot rapport. In 2025 20th ACM/IEEE International Conference on Human-Robot Interaction (HRI). IEEE, 869–879.
  • Nichola Lubold and Heather Pon-Barry. 2014. Acoustic-Prosodic Entrainment and Rapport in Collaborative Learning Dialogues. In Proceedings of the 2014 ACM Workshop on Multimodal Learning Analytics Workshop and Grand Challenge. ACM, Istanbul Turkey, 5–12. https://doi.org/10.1145/2666633.2666635
  • Michael Madaio, Kun Peng, Amy Ogan, and Justine Cassell. 2018. A Climate of Support: A Process-Oriented Analysis of the Impact of Rapport on Peer Tutoring.Grantee Submission (2018).
  • Darlene Magito McLaughlin and Edward G Carr. 2005. Quality of rapport as a setting event for problem behavior: Assessment and intervention. Journal of Positive Behavior Interventions 7, 2 (2005), 68–91.
  • Jessica L Maples-Keller, Rachel L Williamson, Chelsea E Sleep, Nathan T Carter, W Keith Campbell, and Joshua D Miller. 2019. Using item response theory to develop a 60-item representation of the NEO PI–R using the International Personality Item Pool: Development of the IPIP–NEO–60. Journal of personality assessment 101, 1 (2019), 4–15.
  • Robert R McCrae and Oliver P John. 1992. An introduction to the five-factor model and its applications. Journal of personality 60, 2 (1992), 175–215.
  • Lynden K Miles, Louise K Nind, and C Neil Macrae. 2009. The rhythm of rapport: Interpersonal synchrony and social perception. Journal of experimental social psychology 45, 3 (2009), 585–589.
  • Philipp Müller, Michael Xuelin Huang, and Andreas Bulling. 2018. Detecting low rapport during natural interactions in small groups from non-verbal behaviour. In Proceedings of the 23rd International Conference on Intelligent User Interfaces. 153–164.
  • Fumio Nihei, Yukiko I. Nakano, Yuki Hayashi, Hung-Hsuan Huang, and Shogo Okada. 2014. Predicting Influential Statements in Group Discussions Using Speech and Head Motion Information. In Proceedings of the 16th International Conference on Multimodal Interaction (Istanbul, Turkey) (ICMI ’14). Association for Computing Machinery, New York, NY, USA, 136–143. https://doi.org/10.1145/2663204.2663248
  • Catharine Oertel and Giampiero Salvi. 2013. A gaze-based method for relating group involvement to individual engagement in multimodal multiparty dialogue. In Proceedings of the 15th ACM on International conference on multimodal interaction. 99–106.
  • Sam O'Connor Russell, Justine Reverdy, Benjamin Cowan, and Naomi Harte. 2025. Prediction of self-reported and external observations of conversational engagement in online group discussions. Journal on Multimodal User Interfaces (2025), 1–18.
  • Bhargavi Paranjape, Zhen Bai, and Justine Cassell. 2018. Predicting the temporal and social dynamics of curiosity in small group learning. In International conference on artificial intelligence in education. Springer, 420–435.
  • Miranda AG Peeters, Harrie FJM Van Tuijl, Christel G Rutte, and Isabelle MMJ Reymen. 2006. Personality and team performance: a meta-analysis. European journal of personality 20, 5 (2006), 377–396.
  • Gerard J Puccio, Cyndi Burnett, Selcuk Acar, Jo A Yudess, Molly Holinger, and John F Cabra. 2020. Creative problem solving in small groups: The effects of creativity training on idea generation, solution creativity, and leadership effectiveness. The Journal of Creative Behavior 54, 2 (2020), 453–471.
  • Er B Ravinder and AB Saraswathi. 2020. Literature review of Cronbach alpha coefficient (A) and Mcdonald's omega coefficient (Ω). European Journal of Molecular & Clinical Medicine 7, 6 (2020), 2943–2949.
  • Justine Reverdy, Sam O'Connor Russell, Louise Duquenne, Diego Garaialde, Benjamin R Cowan, and Naomi Harte. 2022. RoomReader: A multimodal corpus of online multiparty conversational interactions. In Proceedings of the Thirteenth Language Resources and Evaluation Conference. 2517–2527.
  • Tanja Schneeberger, Anna Lea Reinwarth, Robin Wensky, Manuel Silvio Anglet, Patrick Gebhard, and Janet Wessler. 2023. Fast friends: generating interpersonal closeness between humans and socially interactive agents. In Proceedings of the 23rd ACM international conference on intelligent virtual agents. 1–8.
  • Helen Spencer-Oatey. 2000. Rapport management: A framework for analysis. Culturally speaking: Managing rapport through talk across cultures 1146 (2000).
  • Jessica Sullivan. 2019. The primacy effect in impression formation: Some replications and extensions. Social Psychological and Personality Science 10, 4 (2019), 432–439.
  • Mohsen Tavakol and Reg Dennick. 2011. Making sense of Cronbach's alpha. International journal of medical education 2 (2011), 53.
  • Benjamin Thiry and Maëva Piolti. 2023. IPIP NEO 300, adaptation française européenne. https://benjaminthiry.netlify.app/posts/2023-02-12-ipipneo300fr/. Accessed: 2026-04-12.
  • Linda Tickle-Degnen and Robert Rosenthal. 1990. The nature of rapport and its nonverbal correlates. Psychological inquiry 1, 4 (1990), 285–293.
  • Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems 35 (2022), 10078–10093.
  • Linda R Tropp and Stephen C Wright. 2001. Ingroup identification as the inclusion of ingroup in the self. Personality and Social Psychology Bulletin 27, 5 (2001), 585–600.
  • Annelies EM Van Vianen and Carsten KW De Dreu. 2001. Personality in teams: Its relationship to social cohesion, task cohesion, and team performance. European journal of work and organizational psychology 10, 2 (2001), 97–120.
  • Andreu Vigil-Colet, David Navarro-González, and Fabia Morales-Vives. 2020. To reverse or to not reverse Likert-type items: That is the question. Psicothema 32, 1 (2020), 108–114.
  • Alessandro Vinciarelli, Maja Pantic, and Hervé Bourlard. 2009. Social signal processing: Survey of an emerging domain. Image and vision computing 27, 12 (2009), 1743–1759.
  • Ning Wang and Jonathan Gratch. 2009. Rapport and facial expression. In 2009 3rd International Conference on Affective Computing and Intelligent Interaction and Workshops. IEEE, 1–6.
  • Wenqing Wei, Sixia Li, Candy Olivia Mawalim, Xiguang Li, Kazunori Komatani, and Shogo Okada. 2025. Influence of Personality Traits and Demographics on Rapport Recognition Using Adversarial Learning. Multimodal Technologies and Interaction 9, 3 (March 2025), 18. https://doi.org/10.3390/mti9030018
  • Janie H Wilson and Rebecca G Ryan. 2013. Professor–student rapport scale: Six items predict student outcomes. Teaching of Psychology 40, 2 (2013), 130–133.
  • Camille J Wynn and Stephanie A Borrie. 2022. Classifying conversational entrainment of speech behavior: An expanded framework and review. Journal of Phonetics 94 (2022), 101173.
  • Ran Zhao, Tanmay Sinha, Alan W Black, and Justine Cassell. 2016. Socially-aware virtual agents: Automatically assessing dyadic rapport from temporal patterns of behavior. In International conference on intelligent virtual agents. Springer, 218–233.

A Self-Reported Rapport Scales

This appendix presents the French version of the questionnaires given at the end of the group collaborative discussion. We also give the English translated version. We also note here that given the length of the questionnaire, we accounted for potential participant fatigue and attentional lapses by incorporating several reverse-coded items. To assess the impact of this, we initially conducted a sensitivity analysis to determine if these items introduced systematic response bias or inconsistent patterns. Following the application of an attentional filter—which excluded participants who failed to respond consistently to validated reverse-coded checks—we re-evaluated the factor stability. As the exclusion of these responses did not yield statistically significant differences in the final scores, we retained the full dataset to maintain statistical power, concluding that attention bias did not meaningfully distort the results.

Table 10: Items du questionnaire (version française). Échelle de Likert en 7 points (1 = Pas du tout d'accord, 7 = Tout à fait d'accord). Les items inversés sont marqués (R).
Code Item Échelle
Maintenant que la session est finie, évalue ton groupe
G1 1. Entoure l'image qui décrit le mieux la relation que tu as eue avec le groupe 1–7
GQ1 Q1: L'atmosphère au sein de mon groupe était positive 1–7
GQ2 Q2: Il y avait une forte connexion entre les membres de mon groupe 1–7
GQ3 Q3: Les membres de mon groupe n’étaient pas attentifs les uns aux autres 1–7 (R)
GQ4 Q4: Les échanges au sein de mon groupe étaient fluides 1–7
GQ5 Q5: Il n'y avait pas une bonne alchimie au sein du groupe 1–7 (R)
GQ6 Q6: Pendant l'activité, mon groupe était très investi dans la tâche 1–7
Task satisfaction
GQ7 Q7: Étais-tu d'accord avec l'ordre de passage établi? 1–7
GQ8 Q8: Étais-tu d'accord avec les aliments et emplacements choisis? 1–7
GQ9 Q9: Étais-tu d'accord avec l'itinéraire choisi? 1–7
Et maintenant, évalue ton interaction avec chaque personne avec qui tu as discuté
d1 2.1 Entoure l'image qui décrit le mieux la relation que tu as eue avec 1–7
dq1 Q1: Notre interaction était positive : 1–7
dq2 Q2: Notre interaction était peu amicale 1–7 (R)
dq3 Q3: Notre interaction était fluide 1–7
dq4 Q4: Notre interaction était peu coordonnée 1–7 (R)
dq5 Q5: Nous étions attentifs l'un à l'autre 1–7
dq6 Q6: Notre interaction était captivante 1–7
dq7 Q7: Notre interaction était chaleureuse 1–7
dq8 Q8: Nous étions sur la même longueur d'onde 1–7
dq9 Q9: Nous étions très impliqués l'un vers l'autre 1–7
Table 11: English questionnaire items. All items use a 7-point Likert scale (1 = Strongly disagree, 7 = Strongly agree). Reverse-coded items are marked (R).
Code Item Scale
Now that the session is over, please evaluate your group
G1 1. Circle the image that best describes the relationship you had with the group 1–7
GQ1 Q1: The atmosphere within my group was positive 1–7
GQ2 Q2: There was a strong connection between the members of my group 1–7
GQ3 Q3: The members of my group were not attentive to one another 1–7 (R)
GQ4 Q4: Exchanges within my group were seamless 1–7
GQ5 Q5: There was not a good chemistry within the group 1–7 (R)
GQ6 Q6: During the activity, my group was very invested in the task 1–7
Task satisfaction
GQ7 Q7: Did you agree with the established order of appearance? 1–7
GQ8 Q8: Did you agree with the food items and locations chosen? 1–7
GQ9 Q9: Did you agree with the chosen itinerary? 1–7
And now, evaluate your interaction with each person you spoke with
d1 2.1 Circle the image that best describes the relationship you had with 1–7
dq1 Q1: Our interaction was positive 1–7
dq2 Q2: Our interaction was unfriendly 1–7 (R)
dq3 Q3: Our interaction was fluid 1–7
dq4 Q4: Our interaction was poorly coordinated 1–7 (R)
dq5 Q5: We were attentive to one another 1–7
dq6 Q6: Our interaction was engaging 1–7
dq7 Q7: Our interaction was warm 1–7
dq8 Q8: We were on the same wavelength 1–7
dq9 Q9: We were very involved toward one another 1–7

B Personality Questionnaire Item Selection and Linguistic Adaptation

Concerning the personality questionnaire used for trait estimation, the specific test was selected following a linguistic and psychometric sensitivity analysis conducted for each item by a bilingual researcher (native French speaker with extensive academic experience in English-speaking environments). This review process demonstrated that the Canadian-French version proposed by [27] maintained a higher semantic fidelity to the original English-language personality constructs and also retained neutrality as opposed to the ’European adaptation’ proposed by [60], which was excluded due to the introduction of regional Belgian idiomatic phrasings.

To ensure cross-cultural validity, the political orientation item (‘I tend to vote for liberal political candidates’) was inverted to a conservative-oriented construct, as the term ‘liberal’ possesses significantly different socio-political connotations in France than in the US context, where the test was originally designed. Furthermore, the item ‘I believe in the importance of art’ was translated using a standard, neutral construction (‘Je crois en l'importance de l'art’) to maintain internal consistency within the IPIP-NEO-60 framework. Finally, to prevent order effects and avoid block-related bias, all items were randomized across the survey instrument rather than grouped by factor. See Table below for the French personality questionnaire: items, OCEAN trait, item identifier, and scoring direction.

Table 12: Questionnaire de personnalité (version française, 60 items) : énoncés, trait OCEAN, identifiant et sens de cotation.
Item Énoncé Trait Id. Cot.
Q1 Je suis facilement stressé(e) N N2 +
Q2 J'aime aider les autres A A5 +
Q3 Je prends les choses en main E E5 +
Q4 Je ne m'aime pas N N6 +
Q5 J'aime rêvasser O O2 +
Q6 Je sais comment accomplir les choses C C2 +
Q7 Je me mets facilement en colère N N3 +
Q8 J'ai confiance dans les autres A A1 +
Q9 Je laisse du désordre dans ma chambre C C4
Q10 J'aime la vie E E12 +
Q11 Je profite des autres A A4
Q12 Je reste calme même dans les situations tendues N N11
Q13 J’évite la foule E E4
Q14 Je fixe des standards élevés pour moi et les autres C C8 +
Q15 Je ressens mes émotions très intensément O O5 +
Q16 Je dis la vérité C C5 +
Q17 Je préfère m'en tenir à ce que je connais O O7
Q18 Je me fais facilement des amis E E1 +
Q19 Je prends des décisions sans réfléchir C C12
Q20 J'ai une imagination vive O O1 +
Q21 Je ne tiens pas mes promesses C C6
Q22 J'aime les grandes fêtes E E3 +
Q23 Je parviens à contrôler mes envies N N10
Q24 Je me sens souvent déprimé(e) N N5 +
Q25 J'aime ranger C C3 +
Q26 Je crois en l'importance de l'art O O3 +
Q27 Je m'inquiète pour beaucoup de choses N N1 +
Q28 Je cherche l'aventure E E10 +
Q29 Je triche pour avancer A A3
Q30 Je suis à l'aise avec les autres E E2 +
Q31 J'ai du mal à commencer les tâches C C10
Q32 Je crois que les autres ont de bonnes intentions A A2 +
Q33 Je suis toujours en mouvement E E8 +
Q34 J'insulte les gens A A7
Q35 Je garde mon calme sous pression N N12
Q36 Je n'aime pas l'idée de changement O O8
Q37 J'ai une haute opinion de moi-même A A10
Q38 Je travaille dur C C7 +
Q39 Je m'amuse beaucoup E E11 +
Q40 Je me venge des autres A A8
Q41 J'ai de la sympathie pour ceux qui sont moins bien lotis que moi A A12 +
Q42 Je suis toujours occupé(e) E E7 +
Q43 Je perds facilement mon sang-froid N N4 +
Q44 J’évite les discussions philosophiques O O9
Q45 Je prends des décisions hâtives C C11
Q46 J'ai de la compassion pour les sans-abri A A11 +
Q47 Je gère les tâches avec fluidité C C1 +
Q48 Je n'aime pas l'art O O4
Q49 Je me sens facilement intimidé(e) N N8 +
Q50 J'essaie de diriger les autres E E6 +
Q51 Je m'inquiète pour les autres A A6 +
Q52 Je me crois supérieur(e) aux autres A A9
Q53 Je commence mes tâches immédiatement C C9 +
Q54 J'adore l'excitation E E9 +
Q55 Les discussions théoriques ne m'intéressent pas O O10
Q56 Je ne suis pas facilement influencé(e) par mes émotions O O6
Q57 Je me laisse rarement aller à l'excès N N9
Q58 Je crois en une seule vraie religion O O12
Q59 J'ai du mal à aller vers les autres N N7 +
Q60 J'ai tendance à voter pour des candidats conservateurs O O11

C Complementary Tables

C.1 Study 1

Table 13: EFA factor loadings (λ) and communalities (h2) for the 6-item dyadic rapport scale (early-phase sessions, directed dyadic ratings, single-factor solution). R: reverse-coded.
Code Item (component) λ h2
dq1 Positive (Positivity) .52 .27
dq2 UnfriendlyR (Positivity) .40 .16
dq3 Fluid (Coordination) .86 .74
dq4 Poorly coordinatedR (Coordination) .57 .32
dq5 Attentive (Attentiveness) .66 .44
dq6 Engaging (Attentiveness) .68 .47
Proportion variance .40
N (dyadic ratings) 72
Table 14: EFA factor loadings (λ) and communalities (h2) for the 9-item dyadic rapport scale (later-phase sessions, directed dyadic ratings, single-factor solution). R: reverse-coded.
Code Item (component) λ h2
dq1 Positive (Positivity) .79 .62
dq2 UnfriendlyR (Positivity) .38 .14
dq3 Fluid (Coordination) .85 .72
dq4 Poorly coordinatedR (Coordination) .50 .25
dq5 Attentive (Attentiveness) .78 .60
dq6 Engaging (Attentiveness) .77 .59
dq7 Warm (Positivity) .81 .66
dq8 Same wavelength (Coordination) .73 .54
dq9 Involved (Attentiveness) .76 .58
Proportion variance .52
N (dyadic ratings) 288
Table 15: EFA factor loadings (λ) and communalities (h2) for the original 6-item group rapport scale. The chemistry item (gq5) shows the weakest loading and communality and was removed from subsequent analyses.
Code Item (component) λ h2
gq1 Positive atmosphere (Positivity) .78 .61
gq2 Strong connection (Positivity) .76 .57
gq3 Not attentiveR (Attention) .54 .29
gq4 Seamless exchanges (Coordination) .80 .65
gq5 Poor chemistryR (Chemistry) .26 .07
gq6 Invested in task (Attention) .71 .51
Proportion variance .45
N (raters) 120
Table 16: EFA factor loadings (λ) and communalities (h2) for the retained 5-item group rapport scale (chemistry item removed).
Code Item (component) λ h2
gq1 Positive atmosphere (Positivity) .79 .62
gq2 Strong connection (Positivity) .75 .57
gq3 Not attentiveR (Attention) .53 .28
gq4 Seamless exchanges (Coordination) .80 .64
gq6 Invested in task (Attention) .72 .51
Proportion variance .53
N (raters) 120

C.2 Study 2

Table 17: Internal consistency of OCEAN traits
Trait Cronbach's α McDonald's ω
Openness .59 .71
Conscientiousness .80 .81
Extraversion .76 .77
Agreeableness .77 .78
Neuroticism .76 .76
Table 18: Full fixed-effects estimates for Study 2 models.
Model Predictor β SE df p
A. Dyadic Rapport Intercept .001 .088 27.7 .990
Rater Agreeableness .046 .077 105.5 .553
Rater Conscientiousness .015 .080 109.2 .852
Rater Extraversion .160 .090 108.2 .076
Rater Neuroticism -.069 .086 111.5 .427
Rater Openness -.094 .083 113.8 .257
Target Agreeableness .126 .050 94.3 .013*
Target Conscientiousness -.137 .052 95.6 .010*
Target Extraversion .088 .059 96.8 .139
Target Neuroticism .040 .057 98.3 .482
Target Openness -.066 .056 100.1 .235
B. Group Composition Intercept .000 .108 24.0 1.000
Rater Agreeableness .016 .102 85.0 .871
Rater Conscientiousness .164 .108 85.0 .131
Rater Extraversion .122 .121 85.0 .313
Rater Neuroticism .092 .118 85.0 .441
Rater Openness -.133 .117 85.0 .258
Tetrad-mean Agreeableness .141 .312 29.9 .655
Tetrad-mean Conscientiousness -.107 .280 32.8 .703
Tetrad-mean Extraversion -.244 .367 30.0 .512
Tetrad-mean Neuroticism -.432 .281 35.1 .133
Tetrad-mean Openness .304 .259 37.1 .248
C. Group Link-Anchor Intercept -.038 .146 83.7 .798
Rater Agreeableness .020 .084 105.3 .808
Rater Conscientiousness .115 .086 110.0 .185
Rater Extraversion -.011 .100 109.4 .911
Rater Neuroticism .067 .093 111.6 .474
Rater Openness -.017 .089 111.7 .851
Max Dyadic Rapport .388 .167 111.5 .022*
Min Dyadic Rapport .248 .096 111.9 .011*
*p < .05, **p < .01, ***p < .001. df estimated via Satterthwaite approximation.

C.3 Study 3

Table 19: Full fixed-effects estimates for Study 3 models.
Model Predictor β SE df p
Dyadic – M1 OCEAN Intercept .090 .108 64.3 .408
Rater Agreeableness -.015 .103 57.4 .888
Rater Conscientiousness .070 .112 57.1 .534
Rater Extraversion .087 .123 59.1 .486
Rater Neuroticism -.027 .112 57.6 .812
Rater Openness -.272 .121 60.6 .029*
Target Agreeableness .150 .067 50.5 .030*
Target Conscientiousness -.156 .073 49.3 .036*
Target Extraversion .114 .083 52.7 .174
Target Neuroticism .079 .073 51.0 .284
Target Openness -.039 .084 53.9 .642
Dyadic – M2 Synchrony Intercept .111 .108 69.0 .311
Deep Video Sync -.008 .052 135.3 .873
Deep Audio Sync -.065 .068 136.2 .344
Facial Mimicry (AU12) .038 .082 148.5 .642
Head Pose Sync .005 .061 140.5 .933
Pitch Sync -.383 .099 150.4 < .001***
Intensity Sync .187 .114 139.6 .103
Dyadic – M3 Full Intercept .097 .108 62.0 .371
Rater Agreeableness -.040 .105 57.2 .704
Rater Conscientiousness .100 .113 56.4 .380
Rater Extraversion .046 .125 58.9 .714
Rater Neuroticism -.034 .115 59.3 .769
Rater Openness -.236 .124 60.6 .061
Target Agreeableness .108 .066 50.4 .108
Target Conscientiousness -.106 .071 49.2 .141
Target Extraversion .076 .082 52.1 .357
Target Neuroticism .082 .074 51.2 .270
Target Openness -.017 .083 54.0 .835
Deep Video Sync -.001 .053 128.8 .991
Deep Audio Sync -.045 .071 126.5 .534
Facial Mimicry (AU12) .044 .083 135.6 .596
Head Pose Sync -.011 .062 137.3 .858
Pitch Sync -.328 .103 150.2 .002**
Intensity Sync .146 .119 131.0 .222
Dyadic – M4 Moderation Intercept .101 .108 62.2 .356
Rater Agreeableness -.047 .106 57.6 .659
Rater Conscientiousness .102 .114 56.7 .375
Rater Extraversion .093 .127 60.4 .468
Rater Neuroticism -.027 .116 59.6 .814
Rater Openness -.238 .125 60.9 .061
Target Agreeableness .093 .065 53.1 .160
Target Conscientiousness -.102 .068 48.2 .141
Target Extraversion .100 .080 52.0 .216
Target Neuroticism .078 .071 49.7 .273
Target Openness -.016 .080 52.4 .839
Deep Video Sync .032 .053 125.1 .541
Deep Audio Sync -.041 .070 119.3 .555
Facial Mimicry (AU12) .108 .084 120.4 .202
Head Pose Sync .035 .063 126.6 .578
Pitch Sync -.205 .109 134.1 .062
Intensity Sync .133 .117 127.3 .257
Rater Openness × Pitch Sync -.126 .063 132.5 .046*
Target Agreeableness × Pitch Sync .036 .057 134.6 .527
Target Conscientiousness × Pitch Sync -.082 .056 141.8 .149
Group – M1 OCEAN Intercept -.035 .140 58.0 .805
Rater Agreeableness .052 .142 58.0 .716
Rater Conscientiousness .196 .155 58.0 .214
Rater Extraversion .034 .169 58.0 .842
Rater Neuroticism .033 .155 58.0 .830
Rater Openness -.120 .164 58.0 .468
Group – M2 Synchrony (tetrad-mean) Intercept -.029 .133 57.0 .829
Deep Video Sync -.832 .574 57.0 .153
Deep Audio Sync .204 .242 57.0 .405
Facial Mimicry (AU12) .129 .233 57.0 .584
Head Pose Sync .506 .419 57.0 .232
Pitch Sync -.248 .424 57.0 .561
Intensity Sync -.050 .281 57.0 .860
Group – M3 Full (tetrad-mean sync) Intercept -.032 .137 52.0 .817
Rater Agreeableness .010 .145 52.0 .944
Rater Conscientiousness .196 .158 52.0 .220
Rater Extraversion .000 .177 52.0 .999
Rater Neuroticism .061 .156 52.0 .696
Rater Openness -.138 .168 52.0 .415
Deep Video Sync -.798 .618 52.0 .202
Deep Audio Sync .263 .261 52.0 .318
Facial Mimicry (AU12) .110 .244 52.0 .653
Head Pose Sync .438 .445 52.0 .329
Pitch Sync -.104 .456 52.0 .820
Intensity Sync -.086 .305 52.0 .779
*p < .05, **p < .01, ***p < .001. df estimated via Satterthwaite approximation. Conditional $R^2_c$ by model: Dyadic M1 = .655, M2 = .706, M3 = .697; Group M1 = .050, M2 = .154, M3 = .192 (equal to $R^2_m$ in the group models, reflecting near-zero session-level residual variance beyond the fixed effects).

Footnote

1The dataset is planned for release once the full 30 tetrads are processed with transcriptions and externally assessed rapport annotations.

2These correspond to items G1 (IIS) and d1 (IOS) in the questionnaire in Appendix A.

3As described in section 4, three additional items were added partway through data collection (after tetrad 6) to improve reliability. The dyadic EFAs therefore use directed dyadic ratings from disjoint subsets (6 and 24 tetrads; N = 72 and 288), while the group EFAs use the 120 participant-level ratings. The 6-item dyadic solution rests on fewer independent raters and is treated as preliminary.

4The intercept-only baselines retain the highest cross-validated R2 for both outcomes. This reflects the nested design: out-of-fold, the random intercepts carrying most of the systematic variance revert to the grand mean, making RMSE the more informative criterion here.

CC-BY non-commercial, no derivatives license image
This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

ICMI '26, Napoli, Italy

© 2026 Copyright held by the owner/author(s).
ACM ISBN 979-8-4007-2318-6/26/10.
DOI: https://doi.org/10.1145/3776574.3831156