Validation of a Multi-level Self-Report Rapport Scale and its Impact on Multimodal Modeling of Small Group Interaction
DOI: https://doi.org/10.1145/3776574.3831156
ICMI '26: INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION, Napoli, Italy, October 2026
To investigate how personality traits and multimodal behaviors shape small-group rapport, we introduce the French Self-reported Group and Dyadic Rapport (SGDR) scale. Study 1 validates the SGDR on 30 face-to-face tetrads (groups of four) engaged in collaborative tasks. The results demonstrate high internal consistency for both its dyadic and group subscales, alongside convergent validity with the Inclusion-of-Other-in-Self scale and task satisfaction.
Study 2 investigates the impact of the OCEAN personality traits on rapport, revealing that dyadic and group rapport rely on distinct mechanisms. Dyadic rapport is driven by the target's (interlocutor's) personality: agreeableness predicts it positively, conscientiousness negatively. Conversely, group-level rapport is not significantly explained by either a rater's own traits or the mean levels of personality traits across a tetrad. It is instead anchored in the rater's most extreme dyadic rapport ratings, with both the best and the worst of these ratings independently predicting the rating of group rapport. Notably, conscientiousness operates in opposite directions across levels: it negatively predicts rapport when it is the target's trait at the dyadic level, and positively predicts rapport at the group level (though only as a trend) when it is the rater's own trait. Study 3, on a subset of 16 tetrads, integrates multimodal behaviors. For dyadic rapport, pitch synchrony emerges as the strongest negative predictor, and this association is moderated by rater's openness. Group rapport again shows no significant effects of personality, while tetrad-level synchrony accounts for more of its variance (even if the association remains a trend). These findings demonstrate that rapport is a complex multimodal multi-level phenomenon and that group rapport is not reducible to the sum of dyadic interactions, validating the need for a multi-level scale, such as the SGDR.
ACM Reference Format:
Justine Reverdy, Oussama Silem, and Justine Cassell. 2026. Validation of a Multi-level Self-Report Rapport Scale and its Impact on Multimodal Modeling of Small Group Interaction. In INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION (ICMI '26), October 05--09, 2026, Napoli, Italy. ACM, New York, NY, USA 15 Pages. https://doi.org/10.1145/3776574.3831156
1 Introduction
Rapport, defined as a sense of mutual connection and harmony between individuals, has been considered a key component of both social and task interactions [3, 57, 61]. It emerges from the fine-grained temporal coordination of behavioral signals [46]. While its positive impact on dyadic communication is well-established [6, 12, 42], for humans and conversational agents [33], its role in small groups remains underexplored, overlooking the multi-leveled structure of group dynamics where individuals may address another member or the whole group [25]. As conversational agents transition from personal assistants into multi-party environments, developing robust, interpretable measures of group rapport becomes as important as it is for single-user dialogue systems [13]. Group settings introduce additional complexities as they may emerge through intra-group dyadic and whole-group relationships, while the group composition itself (i.e., the different levels of personality traits within a tetrad) may play a role. This multi-leveled structure raises fundamental questions about how to measure rapport.
We therefore investigate self-reported rapport in groups of four participants (tetrads) engaged in collaborative conversations, based on tasks used in [31, 48]. To this end, building on existing rapport questionnaires [6, 22], we design a post-task questionnaire that captures both dyadic and group-level perceptions. To round out the picture, we collect responses to the Inclusion-of-Other-in-Self scale [2], and ask about task satisfaction. Following work by [31], we also collect self-reported personality traits before the participants gather, via the IPIP-NEO-60 OCEAN inventory online [26, 44, 45] allowing us to see how these individual differences relate to rapport both at dyadic and group levels.
Our analysis proceeds to: (1) a psychometric validation using factor analysis and multi-level modeling; (2) an exploration of the links between personality tests and rapport; and (3) a multimodal exploration using nested model comparisons (likelihood-ratio tests) to assess the respective predictive power of personality and audio-visual features.
2 Related Work
The study of rapport and, more broadly, the pragmatics of social interaction, spans psychology, linguistics, and computational science, yet there exists a gap in how these constructs translate from dyads to groups.
Rapport and Synchrony in Multimodal Group Modeling. Interpersonal rapport is indexed by ways of speaking, nonverbal devices, and behavioral synchrony [6, 14, 57, 61]. The OCEAN personality traits of interlocutors (namely, Openness to Experience, Conscientiousness, Extraversion, Agreeableness, and Neuroticism) have also been shown to play a role [30], with extraversion and agreeableness emerging as the strongest predictors in human-agent conversation [13, 16]. However, the vast majority of this research has focused on dyadic settings [6, 66, 71], and has often focused on either nonverbal behavior or spoken language (both content and acoustic features), but not both. Additionally, prior work demonstrates that processes observed between two people do not necessarily generalize to multi-party settings [25, 38]. On the other hand, numerous studies have explored various other aspects of small group collaboration, such as creative problem solving [53], influential statements [48], curiosity in learning contexts [51], group engagement [49, 55], or group cohesion [34]. However, other phenomena, such as group cohesion, often defined in terms of alignment in attitudes or opinions, do not fully capture rapport, which instead reflects the moment-to-moment signaling of interpersonal connection and harmony during interaction. It therefore remains important to model group rapport.
In this context, [22] links rapport to interactional synchrony and group performance, highlighting its functional role in collaborative settings, and proposes a group rapport scale based on earlier dyadic measures [8]; however, she does not explore the relationship between dyadic and group rapport in groups, nor the impact of personality traits. Building on [22], who found that non-verbal behavioral synchrony was predictive of group rapport, our work adapts a compact subset of her scale's group rapport items to suit the analysis of a dataset that also examines the impact of personality (see section 4).
From another angle, [47] has demonstrated that multimodal behavioral features can detect low dyadic rapport in small groups. Collecting self-reported rapport at the dyadic level, they found the best predictor to be facial expressions. However, their work does not examine the relationship between dyadic and group rapport, nor does it investigate how psychological traits such as personality influence rapport signaling.
Evaluation of Rapport. How rapport is measured differs widely. Some studies rely on self-report questionnaires pertaining to an entire interaction [8, 13]. Others have relied on “thin-slice" observer judgments [1, 29, 41], where video is divided into short segments, and each is annotated by a naive observer.
The 18-item scale of Bernieri [6], originally developed for inter-human interaction, has served as a common reference point for computational models that have relied on self-report [28, 30, 47, 68]. Within these different perspectives, scales tailored to specific interaction contexts have also been proposed, including for teacher-student [69], investigative interviewing [18], and therapy [43]. In human-machine interaction research, dedicated rapport scales have further been introduced [13, 39, 40, 67].
These self-report scales are motivated by work that suggests that external observers may fail to capture the subjective experience of interactants [8, 13, 50], and highlights a mismatch between ‘how participants describe their feelings’ and ‘how observers perceive them’ [17, 21, 50]. For example, [8, 13] reported low consistency between self-reported and observer assessment of rapport, as well as different behavioral correlates across the two perspectives. However, an alternative perspective argues that people may access their internal states largely through their behavior, and that such access can fade by the end of an interaction [5], with end-of-task judgments disproportionately shaped by initial and final impressions [20, 58].
Rapport and Personality in Groups. Much research has examined the role of the five OCEAN personality traits in group outcomes. In fact, a meta-analysis has shown that team personality composition influences performance — often assessed through aggregated supervisor ratings of team effectiveness across task-related dimensions — with higher and more homogeneous levels of agreeableness and Conscientiousness predicting better team outcomes [52]. [9] also highlights the importance of group personality composition for both individual- and group-level outcomes, as coded by observers who rated group behaviors such as conflict, cooperation, communication, and coordination. However, personality remains insufficiently integrated into frameworks that jointly consider multimodal social signals and multi-level perceptions of rapport. In addition, existing approaches often treat personality as a holistic predictor, rather than modeling the different traits and how they interact with behavioral signals to shape emergent group processes such as rapport.
A Multi-Level Approach. Our work addresses the above limitations by examining self-reported dyadic and group-level rapport within the same small-group setting, and by focusing on the impact of observable behaviors as well as the underlying psychological factors of personality traits.
To the best of our knowledge, this is the first scale that proposes to integrate this multi-level approach of rapport into a relatively short questionnaire. By then integrating personality traits and multimodal behavioral synchrony into our analysis, we can investigate how mechanisms that have been shown to underlie dyadic rapport may extend to the group-level —an important extension to current literature. This approach allows us to validate self-report scales while bridging the gap between fine-grained temporal behavioral signals and the underlying psychological traits of personality and rapport.
3 Dataset
The dataset1 consists of groups of four native or near-native French-speaking participants (tetrads) engaging in three verbal tasks designed to elicit collaborative conversations. It is designed to be comparable to the MATRICS corpus recorded in Japanese [31, 48], in order to allow future cross-cultural/cross-linguistic comparisons (small changes were made in order to capture the upper bodies on video). Tetrads participated in three 20-minute tasks: (1) Guest Selection: choosing and scheduling five celebrity guests for a festival and outlining their stage performances; (2) Booth Planning: designing a plan for six food booths distributed across key festival areas; and (3) Travel Planning: organizing a two-day sightseeing itinerary in Paris for a friend, including visits and rest periods.
Recording procedure and setup. At the beginning of each task, written instructions were distributed, then also given verbally by the experimenters, allowing participants to ask questions. Participants were asked to reach a consensus on each task. Before and after each task, participants engaged in 7 minutes of free talk. Video was recorded on four GoPro cameras, and audio was collected with four wireless microphones mounted on headsets. Participants sat around a small table in a quiet room with cameras facing each participant, placed at chest level to ensure they could freely move their upper bodies. Participants received €25 as compensation.
Demographics and subset. 30 tetrads constitute the data for Studies 1 and 2, with a subset of 16 tetrads for Study 3. All groups were same-gender, with an equal number of male and female tetrads. Participants were recruited through a local online research site. Mean participant age was 28 (σ = 9.94, min = 18). Participants were not acquainted beforehand.
4 Scale development
The goal here is to construct a scale that measures both group and dyadic rapport between every pair of participants within a group. The dyadic rapport scale was developed based on [8]. However, the original 18-item scale was not suitable here, where each participant evaluated their level of rapport with every other group member as well as for the group. To limit fatigue while preserving reliability [19], we therefore initially reduced the dyadic rapport scale to 6 items, selecting two for each core component of rapport as conceptualized in [61]—(positivity: Q1, Q2; coordination: Q3–Q4; attentiveness: Q5–Q6). Preliminary analyses revealed insufficient internal consistency (see Study 1 in section 5), and thus the scale was extended with one additional item per component from the original scale, resulting in a 9-item dyadic rapport scale. For group rapport, we relied on prior work by Fultz [22], who proposed a 19-item scale for groups, grounded in rapport theory and the [8] dyadic scale. The validation of the scale in [22] identified two main components through PCA: rapport and compliance with group wishes. We retained a subset of 6 items, with strong loadings on the rapport component (Positivity, Coordination, and Attention). More questions were retained for attentiveness than for coordination to avoid inadvertently capturing group-task functioning instead of rapport, as per [22]. Additionally, we incorporated 3 items to assess task satisfaction, as strong loadings between task satisfaction items and group rapport were reported in [22]. This resulted in a final 9-item group-level scale (6 group rapport and 3 task satisfaction items). To account for potential acquiescence effects [32] in responses, as well as for subsequent validity checks [37, 65], we reverse-coded Q2 and Q4 in the dyadic rapport scale, and Q3 and Q5 in the group scale. All items were rated on a 7-point Likert scale.
To assess convergent validity, we also collected measures of interpersonal closeness using the Inclusion-of-Other-in-Self (IOS) scale [2], which has been shown to correlate with constructs such as liking and intimacy [22, 56], as well as rapport [23]. Participants completed the scale for each partner and then an adapted group-level version, the Inclusion of Ingroup in the Self (IIS) scale for group evaluations [22, 63].
The items from the dyadic and group rapport scales were translated into French by the authors and subsequently reviewed by two native French speakers for clarity and linguistic accuracy (see Appendix A for full questionnaires). Participants recorded responses to all scales on computers right after the interaction.
5 Study 1: Multi-level Psychometric Evaluation of Dyadic and Group Rapport Scales
Following established procedures in psychometric scale validation, we assess the internal consistency of the dyadic and group rapport scales and their convergent validity with task satisfaction, the Inclusion of Other in the Self (IOS) and Inclusion of Ingroup in the Self (IIS) measures2. Dyadic and group scales were analyzed separately within a multi-level framework. Sample sizes are consistent with respondent-to-item heuristics for both scales at their respective units of analysis [10] (see footnote 3). An exploratory factor analysis (EFA) was conducted separately for the two scales to examine their underlying structure. Internal consistency was assessed using Cronbach's α and McDonald's ω [54, 59]. In order to ensure comparability, internal consistency was subsequently re-evaluated.3 Convergent validity was assessed by estimating linear mixed-effects models predicting Dyadic IOS, Group IIS, and Task Satisfaction from rapport scores. Spearman correlation coefficients were computed to examine associations among study variables, after averaging item-level scores into construct-level means for each tetrad and individual at both the group and dyadic levels, merging these aggregates, and using pairwise complete observations. To further investigate interpersonal dynamics, dyadic reciprocity was calculated using Spearman correlations between mutual rapport ratings (rater-to-target and target-to-rater) across the 180 dyads within the 30 tetrads. Finally, to account for the nested structure of the data, a linear mixed-effects model with tetrad as a random intercept was used to examine the effect of gender on dyadic rapport.
5.1 Results
| Var. | M | SD | α | ω | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1. G-IIS | 5.19 | 1.33 | — | — | 1.00 | ||||||
| 2. G-Rap6 | 6.05 | .82 | .80 | .82 | .42** | 1.00 | |||||
| 3. G-Rap5 | 6.03 | .88 | .84 | .85 | .45** | .94** | 1.00 | ||||
| 4. Task Sat. | 6.25 | .92 | .85 | .86 | .27** | .49** | .45** | 1.00 | |||
| 5. D-IOS | 4.62 | 1.36 | — | — | .66** | .31** | .34** | .28** | 1.00 | ||
| 6. D-Rap6 | 5.89 | .65 | .78 | .79 | .53** | .64** | .61** | .51* | .50* | 1.00 | |
| 7. D-Rap9 | 5.71 | .81 | .91 | .91 | .48** | .58** | .66** | .31** | .44** | — | 1.00 |
| Note. G-IIS = Inclusion of Ingroup in the Self; G-Rap6 = 6-item Group Rapport; G-Rap5 = 5-item Group Rapport; D-IOS = Dyadic IOS; D-Rap6 = 6-item Dyadic Rapport; D-Rap9 = 9-item Dyadic Rapport. Scale range: 1–7. G-Rap5 excludes chemistry. Spearman correlations reported. Correlation between D-Rap6 and D-Rap9 unavailable because scales were collected in different experimental phases. *p < .05, **p < .01. |
|||||||||||
Exploratory factor analyses supported a unidimensional structure for both the dyadic and group rapport scales. Detailed loadings are given in Appendix C. For the dyadic scale, all items loaded positively on a single factor, with loadings ranging from .40 to .86 for the 6-item version and from .38 to .85 for the 9-item version. The 9-item version exhibited a stronger and more consistent structure, accounting for 52% of variance compared to the 6-item version (40%). For the group rapport scale, chemistry showed weak loadings and low communality, meaning that this dimension of group rapport was not perceived similarly by respondents. Removal resulted in a clearer factor structure, and the 5-item version was retained for subsequent analyses.
As shown in Table 1, internal consistency was acceptable for the 6-item dyadic scale (α = .78, ω = .79) and excellent for the 9-item version (α = .91, ω = .91). The refined 5-item group scale also demonstrated good internal consistency (α = .84, ω = .85). We note the high means and low standard deviations for both dyadic and group rapport ratings, compared to slightly lower means and higher standard deviations for dyadic and group IOS. Spearman correlational analyses revealed a moderately strong statistically significant (rs = .66) association between both the refined rapport measures, indicating that rapport at different interaction levels is related but not redundant. Results from linear mixed-effects models (Table 2) showed that dyadic rapport was significantly associated with IOS (β = .75, p < .001), while group rapport was significantly associated with IIS (β = .32, p = .02). Group rapport was also a significant positive predictor of task satisfaction (β = .73, std. β = .69, p < .001). Group rapport and task satisfaction were nonetheless only moderately correlated (rs = .45; Table 1), sharing roughly a fifth of their variance and thus remaining empirically distinct rather than redundant. Dyadic reciprocity analyses indicated a weak and non-significant association between mutual ratings within dyads (Spearman's ρ = .11, p = .15). Furthermore, gender did not significantly predict dyadic rapport (β = −.05, p = .773). A null mixed-effects model revealed an intra-class correlation coefficient (ICC) of .253, indicating that 25.3% of the variance in dyadic rapport ratings occurred between tetrads rather than within tetrads. This suggests that dyadic rapport ratings were not independent of the specific tetrad in which participants interacted.
| Outcome | Pred. | β | Std. β | SE | 95% CI | p |
|---|---|---|---|---|---|---|
| Group IIS | G. Rapport | .32 | .21 | .14 | [.05, .59] | .020 |
| Dyadic IOS | D. Rapport | .75 | .43 | .14 | [.46, 1.03] | <.001 |
| Task Sat. | G. Rapport | .73 | .69 | .07 | [.59, .86] | <.001 |
| G. Rapport | D. Rapport | .59 | .52 | .09 | [.41, .76] | <.001 |
| D. Rapport | Gender (M) | -.05 | -.07 | .16 | [-.37, .27] | .773 |
5.2 Discussion
These findings provide initial support for the psychometric properties of the proposed dyadic and group rapport scales. First, the factor analyses support a one factor conceptualization of rapport, echoing previous findings [7]. The addition of items in the dyadic scale and the removal of a weak item in the group scale resulted in clearer factor structures and improved internal consistency.
Second, the convergent validity results provide evidence for a hierarchical structure across dyadic and group levels of interaction. Dyadic rapport was strongly associated with IOS and group rapport was moderately associated with IIS and predicted task satisfaction well, while remaining only moderately correlated with it. This pattern suggests that group rapport may capture aspects of collective interaction quality and coordination in addition to interpersonal closeness [63].
Additional analyses found no gender effect. Crucially, rapport in this setting is not reciprocal at the dyadic level. Again, the substantial variance attributable to tetrad-level clustering indicates that rapport is partly shaped by the specific group dynamic present in each tetrad.
This highlights the importance of accounting for group-level dynamics when studying rapport in multi-party interactions.
6 Study 2: Rapport and Personality Traits
Building on the validated dyadic and group rapport measures, this study investigated whether individual differences in personality traits from the Big Five (OCEAN) predict self-reported rapport. Personality traits were assessed using a 60-item adaptation drawn from the Canadian-French IPIP-NEO-300 [27]. See Appendix B for details and the full questionnaire. Internal consistency estimates for the five personality scales are reported in Appendix C.2.
6.1 Methodology
Linear mixed-effects models were used to examine the relationship between OCEAN personality traits and rapport at both the dyadic and group levels. In all models, the distinction between rater (the participant rating another participant) and target (the participant being rated) was maintained.
For dyadic rapport, fixed effects included both rater and target scores on all five OCEAN traits, with crossed random intercepts for session (tetrad), rater, and target. Crossed random rater and target intercepts allow the model to separate a participant's general tendency to report high or low rapport from their general tendency to be rated highly or poorly by others.
For group rapport, each rater's own report of overall group rapport was retained as an individual observation, so that a rater's traits and the compositional average of their tetrad's personality traits could be estimated as separate, simultaneous predictors. Two complementary specifications were tested against a common null model (random intercept for session only): a composition model, in which each rater's own OCEAN scores and their tetrad's mean OCEAN scores jointly predict group rapport; and a link-anchor model, in which each rater's own OCEAN scores are combined with the maximum and minimum dyadic rapport score they reported across their partners in the same tetrad, motivated by the peak-end rule governing the retrospective evaluation of multi-episode experiences [20].
Nested model comparisons were evaluated via likelihood-ratio tests, with models refit via maximum likelihood for this purpose and via restricted maximum likelihood for parameter estimation. The sample was of N = 360 directed dyads within 30 tetrads for the dyadic model, and N = 120 raters within the same 30 tetrads for the two group-level models.
6.2 Results
6.2.1 Dyadic Rapport. Adding the OCEAN fixed effects significantly improved model fit for dyadic rapport relative to the random-effects-only baseline (χ2(10) = 22.27, p = .014), capturing a marginal variance of $R^2_m = .080$ (Table 3). The targets’ personality traits were the primary drivers of dyadic rapport: targets scoring higher on Agreeableness were rated with significantly higher rapport (β = .126, p = .013), while targets scoring higher on Conscientiousness were rated with significantly lower rapport (β = −.137, p = .010; Table 4). Rater-level traits showed a weaker pattern: rater Extraversion showed a positive trend (β = .160, p = .076), while no other rater-level trait reached significance.
6.2.2 Group Rapport: Composition and Anchoring Effects. Neither a rater's own OCEAN traits nor the tetrad-level mean personality traits (composition) significantly predicted group rapport (χ2(10) = 8.43, p = .587; $R^2_m = .073$); indeed, no individual predictor in this model approached significance.
By contrast, the link-anchor model was highly significant (χ2(7) = 39.43, p < .001; $R^2_m = .283$). Both the maximum and minimum dyadic rapport a rater experienced within their tetrad significantly and positively predicted their group rapport rating (max: β = .388, p = .022; min: β = .248, p = .011; Table 4). The rater's own Conscientiousness showed a positive, non-significant trend in this model (β = .115, p = .185).
| Context / Model | AIC | BIC | $R^2_m$ | p |
|---|---|---|---|---|
| Dyadic Rapport | ||||
| OCEAN Full | 922.37 | 980.66 | .080 | .014* |
| Group Rapport | ||||
| Composition | 355.62 | 391.86 | .073 | .587 |
| Link-Anchor | 318.62 | 346.50 | .283 | < .001*** |
| Note. p-values from likelihood-ratio tests against a random-intercepts-only null model. *p < .05, **p < .01, ***p < .001. |
||||
| Model | Predictor | β | SE | p |
|---|---|---|---|---|
| A. Dyadic | Rater Extraversion | .160 | .090 | .076 |
| Rapport | Target Agreeableness | .126 | .050 | .013* |
| Target Conscientiousness | -.137 | .052 | .010* | |
| C. Group | Rater Conscientiousness | .115 | .086 | .185 |
| Link-Anchor | Max Dyadic Rapport | .388 | .167 | .022* |
| Min Dyadic Rapport | .248 | .096 | .011* | |
| *p < .05, **p < .01, ***p < .001. |
||||
6.3 Discussion
These results indicate that dyadic and group rapport are governed by distinct mechanisms. At the dyadic level, rapport is shaped primarily by the target's personality, as one might expect: participants high in Agreeableness may have facilitated rapport, consistent with Agreeableness's established role in fostering positive impressions in initial dyadic interactions [16]. Target Conscientiousness, on the other hand, limited rapport. Participants that scored higher in Conscientiousness thus may appear more rigid and therefore be evaluated less positively in this collaborative setting.
At the group level, personality, whether captured by a rater's traits or by the tetrad-level mean personality traits does not appear to explain group rapport, suggesting that group-level rapport cannot be straightforwardly mapped onto the aggregation of mean individual levels of OCEAN personality traits. Instead, group rapport is strongly shaped by the extremes of a rater's dyadic experiences within the tetrad: both their most positive and their most negative dyadic interaction independently and positively predict their overall sense of group rapport. This pattern is consistent with a bounded mental shortcut for evaluation [20]: group rapport appears to be bounded above by the best interaction experienced, and bounded below by whether any interaction fell beneath an acceptable relational threshold, rather than built up as an average of dyadic impressions.
The positive, though non-significant, association between rater Conscientiousness and group rapport may reflect rating the collective achievement of consensus goals over individual relational bonding, echoing evidence that Conscientiousness predicts positive team-level outcomes [52, 64].
This stands in contrast to Conscientiousness's negative association with dyadic rapport, where a conscientious partner may be perceived as task-focused rather than easy to bond with. This structural difference may imply that group-level evaluations are driven more by behavioral coordination than by targets’ personality traits.
7 Study 3: Personality and Non-verbal Behavioral Synchrony
Building on the validated rapport measures established in Study 1 and the personality effects identified in Study 2, we examine how non-verbal behavioral synchrony and individual personality differences might jointly contribute to rapport, and whether these associations generalize to unseen groups. Specifically, we asked three questions: (Q1) whether behavioral synchrony explains variance in rapport beyond personality traits, (Q2) whether personality traits retain predictive value once synchrony is considered, and (Q3) whether personality moderates the relationship between behavioral synchrony and rapport. The sample comprises 16 tetrads (64 participants, constituting 192 dyads), the subset of the full sample for which synchronized audio-visual recordings were processed.
7.1 Methodology
We modeled rapport as a function of both individual- and group-level predictors while accounting for the nested data structures. Visual and acoustic representations were extracted from the tetrads using both self-supervised embeddings and handcrafted features. VideoMAE [62] encoded spatiotemporal behavioral dynamics, while WavLM-Large [15] captured prosodic profiles. Facial Action Units and head pose were extracted using OpenFace [4], focusing on AU12 (smiling intensity) and nodding-related motion. Acoustic correlates of vocal arousal were obtained via fundamental frequency (F0) and intensity using Praat [11] through Parselmouth [35]. Synchrony was quantified via Pearson correlation on the behavioral time series (AU12, head pose, pitch, and intensity) and via cosine similarity for the embeddings, capturing actual coordination (see full Table 5). All six synchrony features were z-scored prior to modeling so that coefficients were directly comparable in scale to the OCEAN predictors.
For each outcome (dyadic and group rapport), four nested linear mixed-effects models were compared to address Q1 and Q2: a baseline model with random effects only (M0); a model adding the five OCEAN personality traits (M1); a model adding the six synchrony features (M2); and a full model combining both (M3). The dyadic models retained the crossed random-intercept structure for session, rater, and target. For the group models, each rater's own report of group rapport was retained as an individual observation, with a random intercept for tetrad, and each rater's own OCEAN scores entered as individual-level predictors. Synchrony, being a property of a dyad rather than of an individual, was instead aggregated at the tetrad level: each feature was averaged across all directed dyads within a session, giving a single tetrad-wide value per feature.
To address Q3, a fifth dyadic model (M4) extended the full model (M3) with interaction terms between pitch synchrony – the strongest and most consistent synchrony predictor across M2 and M3 – and the three OCEAN traits that emerged as significant predictors on their own in M1: rater Openness, target Agreeableness, and target Conscientiousness. This analysis was restricted to dyadic rapport; the group-level models, with only 16 tetrads, would not support interaction terms reliably.
Nested model comparisons were evaluated via likelihood-ratio tests, with all models fit via maximum likelihood for this purpose and refit via restricted maximum likelihood for the fixed-effect estimates reported below. To assess whether these associations generalize beyond the sample on which they were estimated, we additionally conducted leave-one-tetrad-out cross-validation (LOGO-CV): for each of the 16 tetrads in turn, all models were refit on the remaining 15 tetrads and used to predict rapport scores for the held-out tetrad. Prediction accuracy is summarized as the root mean squared error (RMSE) and the squared correlation between predicted and observed scores across all out-of-fold predictions.
| Feature | Description | Modality |
| Deep Video Sync | Cosine similarity of VideoMAE embeddings; captures movement coordination. | Visual |
| Deep Audio Sync | Cosine similarity of WavLM-Large embeddings; captures prosodic entrainment. | Audio |
| Facial Mimicry | Correlation of AU12 intensity; measures shared smiles. | Visual |
| Head Pose Sync | Correlation of head pitch (Rx); captures nodding alignment. | Visual |
| Pitch Sync | Correlation of F0 (Hz); indicates pitch alignment. | Audio |
| Intensity Sync | Correlation of energy (dB); indicates shared arousal. | Audio |
7.2 Results
7.2.1 Dyadic Rapport (Q1, Q2). Personality and synchrony features both improved model fit for dyadic rapport, though to different degrees (Table 6). The OCEAN-only model was a marginal improvement (χ2(10) = 16.70, p = .081; $R^2_m=.111$); the synchrony-only (χ2(6) = 22.71, p < .001; $R^2_m=.082$) and full (χ2(16) = 34.35, p = .005; $R^2_m=.162$) models were both significant improvements over baseline.
Adding synchrony to the personality-only model significantly improved fit (χ2(6) = 17.65, p = .007); adding personality to the synchrony-only model did not (χ2(10) = 11.64, p = .310) – synchrony captures variance in dyadic rapport that personality does not.
At the coefficient level (Table 7), the OCEAN-only model reproduced Study 2’s target-level pattern: target Agreeableness positively predicted rapport (β = .150, p = .030) and target Conscientiousness negatively predicted it (β = −.156, p = .036). In addition, within this 16-tetrad multimodal subset, rater Openness emerged as a significant negative predictor (β = −.272, p = .029), although this effect was not observed in the full 30-tetrad personality analysis. Pitch synchrony was the strongest and only significant synchrony predictor, negative both alone (β = −.383, p < .001) and in the full model (β = −.328, p = .002). Intensity synchrony showed a positive trend alone (p = .103). Adding synchrony weakened all three personality effects to trend level.
7.2.2 Personality-Synchrony Interactions (Q3). Adding the three personality-by-pitch-synchrony interactions significantly improved fit over the additive model (χ2(3) = 10.00, p = .019; $R^2_m$: .162 → .183; Table 8). Only rater Openness significantly moderated the pitch-synchrony effect (β = −.126, p = .046); target Agreeableness (p = .527) and target Conscientiousness (p = .149) did not. At mean Openness, pitch synchrony was only a trend (β = −.205, p = .062); the interaction shows this effect is roughly four times stronger for high-Openness raters (predicted slope ≈ −.33 at + 1 SD) than for low-Openness raters (≈ −.08 at − 1 SD).
7.2.3 Group Rapport (Q1, Q2). Personality traits did not improve model fit for group rapport (χ2(5) = 3.25, p = .661; $R^2_m=.050$; Table 6; no individual trait reached significance, Appendix C). Tetrad-level synchrony fared better, though still short of significance: it explained three times more variance than personality ($R^2_m=.154$ vs. .050) and approached significance against baseline (χ2(6) = 10.54, p = .104). The full model was not significant either (χ2(11) = 13.45, p = .265; $R^2_m=.192$), and neither incremental test reached significance (full vs. OCEAN: p = .117; full vs. synchrony: p = .713).
No individual predictor reached significance in either group model (Table 7), but two synchrony features were consistently the largest coefficients: deep video synchrony was negative (β = −.832, p = .153 alone; − .798, p = .202 in the full model) and head-pose synchrony was positive (β = .506, p = .232; .438, p = .329). Rater Conscientiousness showed a similarly consistent, non-significant positive trend (β = .196 in both models, p = .214 and .220).
7.2.4 Cross-Validation. Leave-one-tetrad-out cross-validation tested whether these associations generalize to unseen groups (Table 9). For dyadic rapport, no model beat the baseline (RMSE = .994, R2 = .154); the moderation model (M4) outperformed the additive full model, suggesting part of the interaction generalizes beyond the estimation sample. For group rapport, OCEAN and the full model again underperformed the baseline (RMSE = 1.149 and 1.175 vs. 1.105), while the synchrony-only model was the only model to beat the baseline on RMSE (1.075 vs. 1.105), though not on cross-validated R2 (.063 vs. .242). 4
| Context / Model | AIC | BIC | $R^2_m$ | χ2 (df) | p |
|---|---|---|---|---|---|
| Dyadic Rapport | |||||
| M0 Base | 495.28 | 511.57 | .000 | – | – |
| M1 OCEAN | 498.58 | 547.45 | .111 | 16.70 (10) | .081 |
| M2 Synchrony | 484.57 | 520.40 | .082 | 22.71 (6) | < .001*** |
| M3 Full | 492.93 | 561.34 | .162 | 34.35 (16) | .005** |
| Group Rapport | |||||
| M0 Base | 198.36 | 204.83 | .000 | – | – |
| M1 OCEAN | 205.10 | 222.37 | .050 | 3.25 (5) | .661 |
| M2 Synchrony | 199.82 | 219.25 | .154 | 10.54 (6) | .104 |
| M3 Full | 206.90 | 237.13 | .192 | 13.45 (11) | .265 |
| Note. χ2 and p from likelihood-ratio tests against the corresponding M0 baseline. Group-level synchrony predictors are tetrad-wide averages. *p < .05, **p < .01, ***p < .001. |
|||||
| Model | Predictor | β | SE | p |
|---|---|---|---|---|
| Dyadic M1 | Rater Openness | -.272 | .121 | .029* |
| OCEAN | Target Agreeableness | .150 | .067 | .030* |
| Target Conscientiousness | -.156 | .073 | .036* | |
| Dyadic M2 | Pitch Synchrony | -.383 | .099 | < .001*** |
| Synchrony | Intensity Synchrony | .187 | .114 | .103 |
| Dyadic M3 | Rater Openness | -.236 | .124 | .061 |
| Full | Target Agreeableness | .108 | .066 | .108 |
| Target Conscientiousness | -.106 | .071 | .141 | |
| Pitch Synchrony | -.328 | .103 | .002** | |
| Group M2 | Deep Video Sync | -.832 | .574 | .153 |
| Synchrony | Head Pose Sync | .506 | .419 | .232 |
| Group M3 | Rater Conscientiousness | .196 | .158 | .220 |
| Full | Deep Video Sync | -.798 | .618 | .202 |
| Head Pose Sync | .438 | .445 | .329 | |
| *p < .05, **p < .01, ***p < .001. |
||||
| Predictor | β | SE | p |
|---|---|---|---|
| Pitch Synchrony (main effect) | -.205 | .109 | .062 |
| Rater Openness × Pitch Synchrony | -.126 | .063 | .046* |
| Target Agreeableness × Pitch Synchrony | .036 | .057 | .527 |
| Target Conscientiousness × Pitch Synchrony | -.082 | .056 | .149 |
| Note. Model comparison: M4 (with interactions) vs. M3 (additive only), χ2(3) = 10.00, p = .019. $R^2_m$: M3 = .162, M4 = .183. Full parameter table (all main effects) in Appendix C. *p < .05. |
|||
| Outcome | Model | RMSE | R2 |
|---|---|---|---|
| Dyadic | M0 Base | .994 | .154 |
| M1 OCEAN | 1.026 | .005 | |
| M2 Synchrony | .999 | .008 | |
| M3 Full | 1.044 | .006 | |
| M4 Moderation | 1.033 | .013 | |
| Group | M0 Base | 1.105 | .242 |
| M1 OCEAN | 1.149 | .011 | |
| M2 Synchrony | 1.075 | .063 | |
| M3 Full | 1.175 | .024 | |
| Note. RMSE and R2 are computed across all out-of-fold predictions; R2 is the squared correlation between predicted and observed values. All models converged on all 16 folds for both outcomes. |
|||
7.3 Discussion
Behavioral synchrony, particularly pitch synchrony, is thus associated with dyadic rapport beyond what personality alone captures (Q1), and this is not symmetrical: personality adds little once synchrony is included (Q2), though the target-level effects replicated in M1 and lost significance rather than magnitude in M3. The prominence of vocal pitch synchrony is consistent with prior work on acoustic-prosodic entrainment [70], but its direction is not: higher pitch synchrony predicted lower dyadic rapport. Yet, this direction is not unique to our data: pitch entrainment lowered trust in conversational avatars [24], and a meta-analysis also found vocal pitch synchrony negatively associated with therapeutic alliance [36]. A possible explanation is that it reflects interactional mechanisms specific to multi-party collaboration. Additionally, this relationship is not uniform across individuals (Q3): within the multimodal subset, rater Openness significantly moderated the association between pitch synchrony and rapport. Because rater Openness was not a significant main effect in the full 30-tetrad personality analysis, this moderation should be interpreted as preliminary until replicated in larger multimodal datasets.
At the group level, personality showed no association with rapport, alone or combined with synchrony, though the rater Conscientiousness trend from Study 2 persisted unchanged. Tetrad-level synchrony fared better across two independent tests: it explained more variance than personality in-sample ($R^2_m=.154$ vs. .050, p = .104) and was the only model, dyadic or group, to beat its own baseline on cross-validated RMSE, though not on cross-validated R2. Deep video synchrony (negative) and head-pose synchrony (positive) were the largest individual contributors, though neither reached significance alone, suggesting that further research is needed.
Given the small number of tetrads (16), the group-level patterns are suggestive rather than conclusive. The dyadic moderation finding rests on a larger sample (192 dyads) and is backed by the likelihood-ratio test, providing stronger evidence than the group-level trends. Overall, dyadic rapport appears shaped jointly by personality and behavioral coordination, while group rapport looks more tied to tetrad-wide coordination than to raters’ own personality traits.
8 Conclusion
We introduced and validated the Self-reported Group and Dyadic Rapport (SGDR) scale, a compact instrument measuring rapport at both the dyadic and the group level in small-group interaction. Across three studies (30 tetrads, with 16 in study 3’s multimodal subset), the scale showed solid reliability and meaningful links to interpersonal closeness and task satisfaction at both levels, supporting its use in multimodal research. Findings suggest that dyadic and group rapport may function differently. Dyadic rapport was associated with personality traits and behavioral synchrony: higher target Agreeableness and lower target Conscientiousness predicted higher rapport; when behavioral synchrony was added, pitch synchrony became the strongest predictor. Moreover, these two classes of predictors interact rather than being merely additive: the association between pitch synchrony and rapport was roughly four times stronger for raters high in Openness, a moderation supported by the likelihood-ratio test. Group rapport, by contrast, was explained neither by a rater's own traits nor by tetrad-level mean personality traits. It was instead explained by the extremes of the rater's dyadic experiences (the best and the worst dyadic interactions independently predicted the group rating), while tetrad-wide behavioral coordination showed a suggestive, not-yet-conclusive association that nonetheless outperformed personality both in explained variance and in generalization to unseen groups. Notably, Conscientiousness was significantly negative for dyadic rapport, and trended positive at the group level — a further indication that the two levels may be influenced by different mechanisms. Taken together, these findings caution against treating group rapport as a scaled-up version of dyadic rapport and support modeling the two levels separately. The SGDR provides a practical tool for doing so, and a step toward multimodal systems able to track rapport in task-oriented multi-party interaction.
9 Limitations & Future Work
The present findings should be interpreted in light of several limitations. While studies 1 and 2 included 30 tetrads, Study 3 relied on only 16, limiting the ability to detect more than moderate effects. Group-level findings that include nonverbal and paraverbal behaviors should thus be considered exploratory. The Openness moderation identified in Study 3 should also be interpreted with caution, as it was observed only in the 16-tetrad multimodal subset and not in the full 30-tetrad personality analysis. This discrepancy, together with the comparatively modest reliability of the Openness scale (see Appendix C), suggests that this finding is preliminary and requires replication in larger multimodal datasets.
Results presented within these three studies, moreover, rely on self-report completed at the end of the sessions, while externally assessed rapport annotations, currently underway, will allow self-perceptions to be triangulated with observer judgments. Finally, rapport ratings were generally high with restricted variance, and the dataset covers a single linguistic and cultural context; the planned comparison with the MATRICS corpus [31, 48] will help establish cross-cultural generalizability, as will a planned comparison with our Korean colleagues, also collecting an identical dataset.
10 Generative AI Use Disclosure
Generative AI tools (ChatGPT, Claude) were used for editing and polishing the manuscript. All scientific content, experimental design, and results were produced by the authors.
11 Safe and Responsible Innovation Statement
This work follows ethical and responsible research practices for multimodal interaction. In compliance with GDPR, data are securely stored, pseudonymized, and anonymized where possible, with risks such as personal disclosure mitigated through screening and controlled access. The study protocol, recruitment, and materials were approved by a local Institutional Review Board, and all participants provided informed consent after full information. We support transparent, privacy-preserving, and socially beneficial human–agent interaction while minimizing misuse risks. We also consider bias, inclusivity, and cross-cultural validity in the dataset design.
Acknowledgments
This work was supported by the Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korean government (MSIT) (RS-2022-II220043, Adaptive Personality for Intelligent Agents). Warm thanks to our research collaborators at KETI (Korean Electronics Technology Institute), Jeehyeong Kim, Mira Lee, and Hyoseon Kye, who have inspired many aspects of this research and have been of invaluable help in designing the dataset and working with us to prepare for a cross-cultural comparison. We also thank Professor Yukiko Nakano for her inspirational work on the Japanese MATRICS dataset and for her invaluable advice. Finally, we are grateful for the support of the other members of the ArticuLab group in the ALMAnaCH project-team for their assistance, and particularly master's student Adriano Rivierez for his help with data collection and processing, and lab manager Sophie Etling, for her support with the logistics of data collection.
References
- Nalini Ambady and Robert Rosenthal. 1992. Thin slices of expressive behavior as predictors of interpersonal consequences: A meta-analysis.Psychological bulletin 111, 2 (1992), 256.
- Arthur Aron, Elaine N Aron, and Danny Smollan. 1992. Inclusion of other in the self scale and the structure of interpersonal closeness.Journal of personality and social psychology 63, 4 (1992), 596.
- Zachary G Baker, Emily M Watlington, and C Raymond Knee. 2020. The role of rapport in satisfying one's basic psychological needs. Motivation and emotion 44, 2 (2020), 329–343.
- Tadas Baltrušaitis, Peter Robinson, and Louis-Philippe Morency. 2016. Openface: an open source facial behavior analysis toolkit. In 2016 IEEE winter conference on applications of computer vision (WACV). IEEE, 1–10.
- Daryl J Bem. 1972. Self-perception theory. In Advances in experimental social psychology. Vol. 6. Elsevier, 1–62.
- Frank J Bernieri. 1988. Coordinated movement and rapport in teacher-student interactions. Journal of Nonverbal behavior 12, 2 (1988), 120–138.
- Frank J Bernieri, Janet M Davis, Robert Rosenthal, and C Raymond Knee. 1994. Interactional synchrony and rapport: Measuring synchrony in displays devoid of sound and facial affect. Personality and social psychology bulletin 20, 3 (1994), 303–311.
- Frank J Bernieri, John S Gillis, Janet M Davis, and Jon E Grahe. 1996. Dyad rapport and the accuracy of its judgment across situations: a lens model analysis.Journal of Personality and Social Psychology 71, 1 (1996), 110.
- Eva Bleckmann, Richard Rau, Oliver Lüdtke, Sascha Krause, and Jenny Wagner. 2026. How group personality composition affects person and group outcomes: An integrative analysis using the group actor–partner interdependence model.Journal of Personality and Social Psychology (2026).
- Godfred O Boateng, Torsten B Neilands, Edward A Frongillo, Hugo R Melgar-Quiñonez, and Sera L Young. 2018. Best practices for developing and validating scales for health, social, and behavioral research: a primer. Frontiers in public health 6 (2018), 149.
- Paul Boersma and David Weenink. 2020. Praat: Doing Phonetics by Computer [Computer program]. Version 6.1.38, retrieved from http://www.praat.org/.
- Justine Cassell, Alastair Gill, and Paul Tepper. 2007. Coordination in conversation and rapport. In Proceedings of the workshop on Embodied Language Processing. 41–50.
- Aleksandra Cerekovic, Oya Aran, and Daniel Gatica-Perez. 2016. Rapport with virtual agents: What do human social cues and personality explain?IEEE Transactions on Affective Computing 8, 3 (2016), 382–395.
- Tanya L Chartrand and John A Bargh. 1999. The chameleon effect: The perception-behavior link and social interaction.Journal of Personality and Social Psychology 76, 6 (1999), 893–910.
- Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al. 2022. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing 16, 6 (2022), 1505–1518.
- Ronen Cuperman and William Ickes. 2009. Big Five predictors of behavior and perceptions in initial dyadic interactions: Personality similarity helps extraverts and introverts, but hurts “disagreeables”.Journal of personality and social psychology 97, 4 (2009), 667.
- Sidney D'Mello, Ed Dieterle, and Angela Duckworth. 2017. Advanced, analytic, automated (AAA) measurement of engagement during learning. Educational psychologist 52, 2 (2017), 104–123.
- Misty C Duke, James M Wood, Brock Bollin, Matthew Scullin, and Julia LaBianca. 2018. Development of the Rapport Scales for Investigative Interviews and Interrogations (RS3i), Interviewee Version.Psychology, Public Policy, and Law 24, 1 (2018), 64.
- Gudrun Eisele, Hugo Vachon, Ginette Lafit, Peter Kuppens, Marlies Houben, Inez Myin-Germeys, and Wolfgang Viechtbauer. 2022. The effects of sampling frequency and questionnaire length on perceived burden, compliance, and careless responding in experience sampling data in a student population. Assessment 29, 2 (2022), 136–151.
- Barbara L Fredrickson and Daniel Kahneman. 1993. Duration neglect in retrospective evaluations of affective episodes.Journal of personality and social psychology 65, 1 (1993), 45.
- Kathryn A Fuller, Nilushi S Karunaratne, Som Naidu, Betty Exintaris, Jennifer L Short, Michael D Wolcott, Scott Singleton, and Paul J White. 2018. Development of a self-report instrument for measuring in-class student engagement reveals that pretending to engage is a significant unrecognized problem. PloS one 13, 10 (2018), e0205828.
- Amber A. Fultz. 2023. The Relationships Between Synchrony, Rapport, and Small Group Performance. Doctoral dissertation. Oregon State University. https://ir.library.oregonstate.edu/concern/graduate_thesis_or_dissertations/47429j58kAdvisor: Frank J. Bernieri.
- Simon Gächter, Chris Starmer, and Fabio Tufano. 2015. Measuring the closeness of relationships: a comprehensive evaluation of the'inclusion of the other in the self'scale. PloS one 10, 6 (2015), e0129478.
- Ramiro H Gálvez, Agustín Gravano, Štefan Beňuš, Rivka Levitan, Marian Trnka, and Julia Hirschberg. 2020. An empirical study of the effect of acoustic-prosodic entrainment on the perceived trustworthiness of conversational avatars. Speech Communication 124 (2020), 46–67.
- Erving Goffman. 1981. Forms of talk. University of Pennsylvania Press.
- Lewis R Goldberg et al. 1999. A broad-bandwidth, public domain, personality inventory measuring the lower-level facets of several five-factor models. Personality psychology in Europe 7, 1 (1999), 7–28.
- L. Gravel. 2001. French (Canadian) Translation of the IPIP Version of the NEO PI-R. https://ipip.ori.org/FrenchCanadian300-Item-IPIP-NEO.htm. Accessed: 2026-04-12.
- Juan Lorenzo Hagad, Roberto Legaspi, Masayuki Numao, and Merlin Suarez. 2011. Predicting Levels of Rapport in Dyadic Interactions through Automatic Detection of Posture and Posture Congruence. In 2011 IEEE Third International Conference on Privacy, Security, Risk and Trust and 2011 IEEE Third International Conference on Social Computing. 613–616. https://doi.org/10.1109/PASSAT/SocialCom.2011.143
- Judith A. Hall, Debra L. Roter, Danielle C. Blanch, and Richard M. Frankel. 2009. Observer-Rated Rapport in Interactions between Medical Students and Standardized Patients. Patient Education and Counseling 76, 3 (Sept. 2009), 323–327. https://doi.org/10.1016/j.pec.2009.05.009
- Takato Hayashi, Ryusei Kimura, Ryo Ishii, and Shogo Okada. 2025. Investigating Role of Big Five Personality Traits in Audio-Visual Rapport Estimation. In 2025 IEEE 19th International Conference on Automatic Face and Gesture Recognition (FG). IEEE, 1–10.
- Yuki Hayashi, Fumio Nihei, Yukiko Nakano, Hung-Hsuan Huang, and Shogo Okada. 2015. Construction of a Group Discussion Corpus and Analysis of its Relationship with Personality Traits. IPSJ Journal 56, 4 (2015), 1217–1227.
- Andreas Hinz, Dominik Michalski, Reinhold Schwarz, and Philipp Yorck Herzberg. 2007. The acquiescence effect in responding to a questionnaire. GMS Psycho-Social Medicine 4 (2007), Doc07.
- Lixing Huang, Louis-Philippe Morency, and Jonathan Gratch. 2011. Virtual Rapport 2.0. In Intelligent Virtual Agents, Hannes Högni Vilhjálmsson, Stefan Kopp, Stacy Marsella, and Kristinn R. Thórisson (Eds.). Springer, Berlin, Heidelberg, 68–79. https://doi.org/10.1007/978-3-642-23974-8_8
- Hayley Hung and Daniel Gatica-Perez. 2010. Estimating cohesion in small groups using audio-visual nonverbal behavior. IEEE Transactions on Multimedia 12, 6 (2010), 563–575.
- Yannick Jadoul, Bill Thompson, and Bart de Boer. 2018. Introducing Parselmouth: A Python interface to Praat. Journal of Phonetics (2018). praat-parselmouth version 0.4.7.
- Simone Jennissen, Julia Huber, Beate Ditzen, and Ulrike Dinger. 2025. Association between nonverbal synchrony, alliance, and outcome in psychotherapy: systematic review and meta-analysis. Psychotherapy Research 35, 7 (2025), 1213–1228.
- Krisztián Józsa and George A Morgan. 2017. Reversed items in Likert scales: Filtering out invalid responders. Journal of Psychological and Educational Research 25, 1 (2017), 7–25.
- Nale Lehmann-Willenbrock, Hayley Hung, and Joann Keyton. 2017. New frontiers in analyzing dynamic group interactions: Bridging social and computer science. Small group research 48, 5 (2017), 519–531.
- Ting-Han Lin, Guan Chen, Bilge Mutlu, J Gregory Trafton, and Sarah Sebo. 2026. The Reduced-Length Connection-Coordination Rapport (CCR) Scale. ACM Transactions on Human-Robot Interaction (2026).
- Ting-Han Lin, Hannah Dinner, Tsz Long Leung, Bilge Mutlu, J Gregory Trafton, and Sarah Sebo. 2025. Connection-Coordination rapport (CCR) scale: a dual-factor scale to measure human-robot rapport. In 2025 20th ACM/IEEE International Conference on Human-Robot Interaction (HRI). IEEE, 869–879.
- Nichola Lubold and Heather Pon-Barry. 2014. Acoustic-Prosodic Entrainment and Rapport in Collaborative Learning Dialogues. In Proceedings of the 2014 ACM Workshop on Multimodal Learning Analytics Workshop and Grand Challenge. ACM, Istanbul Turkey, 5–12. https://doi.org/10.1145/2666633.2666635
- Michael Madaio, Kun Peng, Amy Ogan, and Justine Cassell. 2018. A Climate of Support: A Process-Oriented Analysis of the Impact of Rapport on Peer Tutoring.Grantee Submission (2018).
- Darlene Magito McLaughlin and Edward G Carr. 2005. Quality of rapport as a setting event for problem behavior: Assessment and intervention. Journal of Positive Behavior Interventions 7, 2 (2005), 68–91.
- Jessica L Maples-Keller, Rachel L Williamson, Chelsea E Sleep, Nathan T Carter, W Keith Campbell, and Joshua D Miller. 2019. Using item response theory to develop a 60-item representation of the NEO PI–R using the International Personality Item Pool: Development of the IPIP–NEO–60. Journal of personality assessment 101, 1 (2019), 4–15.
- Robert R McCrae and Oliver P John. 1992. An introduction to the five-factor model and its applications. Journal of personality 60, 2 (1992), 175–215.
- Lynden K Miles, Louise K Nind, and C Neil Macrae. 2009. The rhythm of rapport: Interpersonal synchrony and social perception. Journal of experimental social psychology 45, 3 (2009), 585–589.
- Philipp Müller, Michael Xuelin Huang, and Andreas Bulling. 2018. Detecting low rapport during natural interactions in small groups from non-verbal behaviour. In Proceedings of the 23rd International Conference on Intelligent User Interfaces. 153–164.
- Fumio Nihei, Yukiko I. Nakano, Yuki Hayashi, Hung-Hsuan Huang, and Shogo Okada. 2014. Predicting Influential Statements in Group Discussions Using Speech and Head Motion Information. In Proceedings of the 16th International Conference on Multimodal Interaction (Istanbul, Turkey) (ICMI ’14). Association for Computing Machinery, New York, NY, USA, 136–143. https://doi.org/10.1145/2663204.2663248
- Catharine Oertel and Giampiero Salvi. 2013. A gaze-based method for relating group involvement to individual engagement in multimodal multiparty dialogue. In Proceedings of the 15th ACM on International conference on multimodal interaction. 99–106.
- Sam O'Connor Russell, Justine Reverdy, Benjamin Cowan, and Naomi Harte. 2025. Prediction of self-reported and external observations of conversational engagement in online group discussions. Journal on Multimodal User Interfaces (2025), 1–18.
- Bhargavi Paranjape, Zhen Bai, and Justine Cassell. 2018. Predicting the temporal and social dynamics of curiosity in small group learning. In International conference on artificial intelligence in education. Springer, 420–435.
- Miranda AG Peeters, Harrie FJM Van Tuijl, Christel G Rutte, and Isabelle MMJ Reymen. 2006. Personality and team performance: a meta-analysis. European journal of personality 20, 5 (2006), 377–396.
- Gerard J Puccio, Cyndi Burnett, Selcuk Acar, Jo A Yudess, Molly Holinger, and John F Cabra. 2020. Creative problem solving in small groups: The effects of creativity training on idea generation, solution creativity, and leadership effectiveness. The Journal of Creative Behavior 54, 2 (2020), 453–471.
- Er B Ravinder and AB Saraswathi. 2020. Literature review of Cronbach alpha coefficient (A) and Mcdonald's omega coefficient (Ω). European Journal of Molecular & Clinical Medicine 7, 6 (2020), 2943–2949.
- Justine Reverdy, Sam O'Connor Russell, Louise Duquenne, Diego Garaialde, Benjamin R Cowan, and Naomi Harte. 2022. RoomReader: A multimodal corpus of online multiparty conversational interactions. In Proceedings of the Thirteenth Language Resources and Evaluation Conference. 2517–2527.
- Tanja Schneeberger, Anna Lea Reinwarth, Robin Wensky, Manuel Silvio Anglet, Patrick Gebhard, and Janet Wessler. 2023. Fast friends: generating interpersonal closeness between humans and socially interactive agents. In Proceedings of the 23rd ACM international conference on intelligent virtual agents. 1–8.
- Helen Spencer-Oatey. 2000. Rapport management: A framework for analysis. Culturally speaking: Managing rapport through talk across cultures 1146 (2000).
- Jessica Sullivan. 2019. The primacy effect in impression formation: Some replications and extensions. Social Psychological and Personality Science 10, 4 (2019), 432–439.
- Mohsen Tavakol and Reg Dennick. 2011. Making sense of Cronbach's alpha. International journal of medical education 2 (2011), 53.
- Benjamin Thiry and Maëva Piolti. 2023. IPIP NEO 300, adaptation française européenne. https://benjaminthiry.netlify.app/posts/2023-02-12-ipipneo300fr/. Accessed: 2026-04-12.
- Linda Tickle-Degnen and Robert Rosenthal. 1990. The nature of rapport and its nonverbal correlates. Psychological inquiry 1, 4 (1990), 285–293.
- Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems 35 (2022), 10078–10093.
- Linda R Tropp and Stephen C Wright. 2001. Ingroup identification as the inclusion of ingroup in the self. Personality and Social Psychology Bulletin 27, 5 (2001), 585–600.
- Annelies EM Van Vianen and Carsten KW De Dreu. 2001. Personality in teams: Its relationship to social cohesion, task cohesion, and team performance. European journal of work and organizational psychology 10, 2 (2001), 97–120.
- Andreu Vigil-Colet, David Navarro-González, and Fabia Morales-Vives. 2020. To reverse or to not reverse Likert-type items: That is the question. Psicothema 32, 1 (2020), 108–114.
- Alessandro Vinciarelli, Maja Pantic, and Hervé Bourlard. 2009. Social signal processing: Survey of an emerging domain. Image and vision computing 27, 12 (2009), 1743–1759.
- Ning Wang and Jonathan Gratch. 2009. Rapport and facial expression. In 2009 3rd International Conference on Affective Computing and Intelligent Interaction and Workshops. IEEE, 1–6.
- Wenqing Wei, Sixia Li, Candy Olivia Mawalim, Xiguang Li, Kazunori Komatani, and Shogo Okada. 2025. Influence of Personality Traits and Demographics on Rapport Recognition Using Adversarial Learning. Multimodal Technologies and Interaction 9, 3 (March 2025), 18. https://doi.org/10.3390/mti9030018
- Janie H Wilson and Rebecca G Ryan. 2013. Professor–student rapport scale: Six items predict student outcomes. Teaching of Psychology 40, 2 (2013), 130–133.
- Camille J Wynn and Stephanie A Borrie. 2022. Classifying conversational entrainment of speech behavior: An expanded framework and review. Journal of Phonetics 94 (2022), 101173.
- Ran Zhao, Tanmay Sinha, Alan W Black, and Justine Cassell. 2016. Socially-aware virtual agents: Automatically assessing dyadic rapport from temporal patterns of behavior. In International conference on intelligent virtual agents. Springer, 218–233.
A Self-Reported Rapport Scales
This appendix presents the French version of the questionnaires given at the end of the group collaborative discussion. We also give the English translated version. We also note here that given the length of the questionnaire, we accounted for potential participant fatigue and attentional lapses by incorporating several reverse-coded items. To assess the impact of this, we initially conducted a sensitivity analysis to determine if these items introduced systematic response bias or inconsistent patterns. Following the application of an attentional filter—which excluded participants who failed to respond consistently to validated reverse-coded checks—we re-evaluated the factor stability. As the exclusion of these responses did not yield statistically significant differences in the final scores, we retained the full dataset to maintain statistical power, concluding that attention bias did not meaningfully distort the results.
| Code | Item | Échelle |
| Maintenant que la session est finie, évalue ton groupe | ||
| G1 | 1. Entoure l'image qui décrit le mieux la relation que tu as eue avec le groupe | 1–7 |
| GQ1 | Q1: L'atmosphère au sein de mon groupe était positive | 1–7 |
| GQ2 | Q2: Il y avait une forte connexion entre les membres de mon groupe | 1–7 |
| GQ3 | Q3: Les membres de mon groupe n’étaient pas attentifs les uns aux autres | 1–7 (R) |
| GQ4 | Q4: Les échanges au sein de mon groupe étaient fluides | 1–7 |
| GQ5 | Q5: Il n'y avait pas une bonne alchimie au sein du groupe | 1–7 (R) |
| GQ6 | Q6: Pendant l'activité, mon groupe était très investi dans la tâche | 1–7 |
| Task satisfaction | ||
| GQ7 | Q7: Étais-tu d'accord avec l'ordre de passage établi? | 1–7 |
| GQ8 | Q8: Étais-tu d'accord avec les aliments et emplacements choisis? | 1–7 |
| GQ9 | Q9: Étais-tu d'accord avec l'itinéraire choisi? | 1–7 |
| Et maintenant, évalue ton interaction avec chaque personne avec qui tu as discuté | ||
| d1 | 2.1 Entoure l'image qui décrit le mieux la relation que tu as eue avec | 1–7 |
| dq1 | Q1: Notre interaction était positive : | 1–7 |
| dq2 | Q2: Notre interaction était peu amicale | 1–7 (R) |
| dq3 | Q3: Notre interaction était fluide | 1–7 |
| dq4 | Q4: Notre interaction était peu coordonnée | 1–7 (R) |
| dq5 | Q5: Nous étions attentifs l'un à l'autre | 1–7 |
| dq6 | Q6: Notre interaction était captivante | 1–7 |
| dq7 | Q7: Notre interaction était chaleureuse | 1–7 |
| dq8 | Q8: Nous étions sur la même longueur d'onde | 1–7 |
| dq9 | Q9: Nous étions très impliqués l'un vers l'autre | 1–7 |
| Code | Item | Scale |
| Now that the session is over, please evaluate your group | ||
| G1 | 1. Circle the image that best describes the relationship you had with the group | 1–7 |
| GQ1 | Q1: The atmosphere within my group was positive | 1–7 |
| GQ2 | Q2: There was a strong connection between the members of my group | 1–7 |
| GQ3 | Q3: The members of my group were not attentive to one another | 1–7 (R) |
| GQ4 | Q4: Exchanges within my group were seamless | 1–7 |
| GQ5 | Q5: There was not a good chemistry within the group | 1–7 (R) |
| GQ6 | Q6: During the activity, my group was very invested in the task | 1–7 |
| Task satisfaction | ||
| GQ7 | Q7: Did you agree with the established order of appearance? | 1–7 |
| GQ8 | Q8: Did you agree with the food items and locations chosen? | 1–7 |
| GQ9 | Q9: Did you agree with the chosen itinerary? | 1–7 |
| And now, evaluate your interaction with each person you spoke with | ||
| d1 | 2.1 Circle the image that best describes the relationship you had with | 1–7 |
| dq1 | Q1: Our interaction was positive | 1–7 |
| dq2 | Q2: Our interaction was unfriendly | 1–7 (R) |
| dq3 | Q3: Our interaction was fluid | 1–7 |
| dq4 | Q4: Our interaction was poorly coordinated | 1–7 (R) |
| dq5 | Q5: We were attentive to one another | 1–7 |
| dq6 | Q6: Our interaction was engaging | 1–7 |
| dq7 | Q7: Our interaction was warm | 1–7 |
| dq8 | Q8: We were on the same wavelength | 1–7 |
| dq9 | Q9: We were very involved toward one another | 1–7 |
B Personality Questionnaire Item Selection and Linguistic Adaptation
Concerning the personality questionnaire used for trait estimation, the specific test was selected following a linguistic and psychometric sensitivity analysis conducted for each item by a bilingual researcher (native French speaker with extensive academic experience in English-speaking environments). This review process demonstrated that the Canadian-French version proposed by [27] maintained a higher semantic fidelity to the original English-language personality constructs and also retained neutrality as opposed to the ’European adaptation’ proposed by [60], which was excluded due to the introduction of regional Belgian idiomatic phrasings.
To ensure cross-cultural validity, the political orientation item (‘I tend to vote for liberal political candidates’) was inverted to a conservative-oriented construct, as the term ‘liberal’ possesses significantly different socio-political connotations in France than in the US context, where the test was originally designed. Furthermore, the item ‘I believe in the importance of art’ was translated using a standard, neutral construction (‘Je crois en l'importance de l'art’) to maintain internal consistency within the IPIP-NEO-60 framework. Finally, to prevent order effects and avoid block-related bias, all items were randomized across the survey instrument rather than grouped by factor. See Table below for the French personality questionnaire: items, OCEAN trait, item identifier, and scoring direction.
| Item | Énoncé | Trait | Id. | Cot. |
|---|---|---|---|---|
| Q1 | Je suis facilement stressé(e) | N | N2 | + |
| Q2 | J'aime aider les autres | A | A5 | + |
| Q3 | Je prends les choses en main | E | E5 | + |
| Q4 | Je ne m'aime pas | N | N6 | + |
| Q5 | J'aime rêvasser | O | O2 | + |
| Q6 | Je sais comment accomplir les choses | C | C2 | + |
| Q7 | Je me mets facilement en colère | N | N3 | + |
| Q8 | J'ai confiance dans les autres | A | A1 | + |
| Q9 | Je laisse du désordre dans ma chambre | C | C4 | − |
| Q10 | J'aime la vie | E | E12 | + |
| Q11 | Je profite des autres | A | A4 | − |
| Q12 | Je reste calme même dans les situations tendues | N | N11 | − |
| Q13 | J’évite la foule | E | E4 | − |
| Q14 | Je fixe des standards élevés pour moi et les autres | C | C8 | + |
| Q15 | Je ressens mes émotions très intensément | O | O5 | + |
| Q16 | Je dis la vérité | C | C5 | + |
| Q17 | Je préfère m'en tenir à ce que je connais | O | O7 | − |
| Q18 | Je me fais facilement des amis | E | E1 | + |
| Q19 | Je prends des décisions sans réfléchir | C | C12 | − |
| Q20 | J'ai une imagination vive | O | O1 | + |
| Q21 | Je ne tiens pas mes promesses | C | C6 | − |
| Q22 | J'aime les grandes fêtes | E | E3 | + |
| Q23 | Je parviens à contrôler mes envies | N | N10 | − |
| Q24 | Je me sens souvent déprimé(e) | N | N5 | + |
| Q25 | J'aime ranger | C | C3 | + |
| Q26 | Je crois en l'importance de l'art | O | O3 | + |
| Q27 | Je m'inquiète pour beaucoup de choses | N | N1 | + |
| Q28 | Je cherche l'aventure | E | E10 | + |
| Q29 | Je triche pour avancer | A | A3 | − |
| Q30 | Je suis à l'aise avec les autres | E | E2 | + |
| Q31 | J'ai du mal à commencer les tâches | C | C10 | − |
| Q32 | Je crois que les autres ont de bonnes intentions | A | A2 | + |
| Q33 | Je suis toujours en mouvement | E | E8 | + |
| Q34 | J'insulte les gens | A | A7 | − |
| Q35 | Je garde mon calme sous pression | N | N12 | − |
| Q36 | Je n'aime pas l'idée de changement | O | O8 | − |
| Q37 | J'ai une haute opinion de moi-même | A | A10 | − |
| Q38 | Je travaille dur | C | C7 | + |
| Q39 | Je m'amuse beaucoup | E | E11 | + |
| Q40 | Je me venge des autres | A | A8 | − |
| Q41 | J'ai de la sympathie pour ceux qui sont moins bien lotis que moi | A | A12 | + |
| Q42 | Je suis toujours occupé(e) | E | E7 | + |
| Q43 | Je perds facilement mon sang-froid | N | N4 | + |
| Q44 | J’évite les discussions philosophiques | O | O9 | − |
| Q45 | Je prends des décisions hâtives | C | C11 | − |
| Q46 | J'ai de la compassion pour les sans-abri | A | A11 | + |
| Q47 | Je gère les tâches avec fluidité | C | C1 | + |
| Q48 | Je n'aime pas l'art | O | O4 | − |
| Q49 | Je me sens facilement intimidé(e) | N | N8 | + |
| Q50 | J'essaie de diriger les autres | E | E6 | + |
| Q51 | Je m'inquiète pour les autres | A | A6 | + |
| Q52 | Je me crois supérieur(e) aux autres | A | A9 | − |
| Q53 | Je commence mes tâches immédiatement | C | C9 | + |
| Q54 | J'adore l'excitation | E | E9 | + |
| Q55 | Les discussions théoriques ne m'intéressent pas | O | O10 | − |
| Q56 | Je ne suis pas facilement influencé(e) par mes émotions | O | O6 | − |
| Q57 | Je me laisse rarement aller à l'excès | N | N9 | − |
| Q58 | Je crois en une seule vraie religion | O | O12 | − |
| Q59 | J'ai du mal à aller vers les autres | N | N7 | + |
| Q60 | J'ai tendance à voter pour des candidats conservateurs | O | O11 | − |
C Complementary Tables
C.1 Study 1
| Code | Item (component) | λ | h2 |
|---|---|---|---|
| dq1 | Positive (Positivity) | .52 | .27 |
| dq2 | UnfriendlyR (Positivity) | .40 | .16 |
| dq3 | Fluid (Coordination) | .86 | .74 |
| dq4 | Poorly coordinatedR (Coordination) | .57 | .32 |
| dq5 | Attentive (Attentiveness) | .66 | .44 |
| dq6 | Engaging (Attentiveness) | .68 | .47 |
| Proportion variance | .40 | ||
| N (dyadic ratings) | 72 | ||
| Code | Item (component) | λ | h2 |
|---|---|---|---|
| dq1 | Positive (Positivity) | .79 | .62 |
| dq2 | UnfriendlyR (Positivity) | .38 | .14 |
| dq3 | Fluid (Coordination) | .85 | .72 |
| dq4 | Poorly coordinatedR (Coordination) | .50 | .25 |
| dq5 | Attentive (Attentiveness) | .78 | .60 |
| dq6 | Engaging (Attentiveness) | .77 | .59 |
| dq7 | Warm (Positivity) | .81 | .66 |
| dq8 | Same wavelength (Coordination) | .73 | .54 |
| dq9 | Involved (Attentiveness) | .76 | .58 |
| Proportion variance | .52 | ||
| N (dyadic ratings) | 288 | ||
| Code | Item (component) | λ | h2 |
|---|---|---|---|
| gq1 | Positive atmosphere (Positivity) | .78 | .61 |
| gq2 | Strong connection (Positivity) | .76 | .57 |
| gq3 | Not attentiveR (Attention) | .54 | .29 |
| gq4 | Seamless exchanges (Coordination) | .80 | .65 |
| gq5 | Poor chemistryR (Chemistry) | .26 | .07 |
| gq6 | Invested in task (Attention) | .71 | .51 |
| Proportion variance | .45 | ||
| N (raters) | 120 | ||
| Code | Item (component) | λ | h2 |
|---|---|---|---|
| gq1 | Positive atmosphere (Positivity) | .79 | .62 |
| gq2 | Strong connection (Positivity) | .75 | .57 |
| gq3 | Not attentiveR (Attention) | .53 | .28 |
| gq4 | Seamless exchanges (Coordination) | .80 | .64 |
| gq6 | Invested in task (Attention) | .72 | .51 |
| Proportion variance | .53 | ||
| N (raters) | 120 | ||
C.2 Study 2
| Trait | Cronbach's α | McDonald's ω |
|---|---|---|
| Openness | .59 | .71 |
| Conscientiousness | .80 | .81 |
| Extraversion | .76 | .77 |
| Agreeableness | .77 | .78 |
| Neuroticism | .76 | .76 |
| Model | Predictor | β | SE | df | p |
|---|---|---|---|---|---|
| A. Dyadic Rapport | Intercept | .001 | .088 | 27.7 | .990 |
| Rater Agreeableness | .046 | .077 | 105.5 | .553 | |
| Rater Conscientiousness | .015 | .080 | 109.2 | .852 | |
| Rater Extraversion | .160 | .090 | 108.2 | .076 | |
| Rater Neuroticism | -.069 | .086 | 111.5 | .427 | |
| Rater Openness | -.094 | .083 | 113.8 | .257 | |
| Target Agreeableness | .126 | .050 | 94.3 | .013* | |
| Target Conscientiousness | -.137 | .052 | 95.6 | .010* | |
| Target Extraversion | .088 | .059 | 96.8 | .139 | |
| Target Neuroticism | .040 | .057 | 98.3 | .482 | |
| Target Openness | -.066 | .056 | 100.1 | .235 | |
| B. Group Composition | Intercept | .000 | .108 | 24.0 | 1.000 |
| Rater Agreeableness | .016 | .102 | 85.0 | .871 | |
| Rater Conscientiousness | .164 | .108 | 85.0 | .131 | |
| Rater Extraversion | .122 | .121 | 85.0 | .313 | |
| Rater Neuroticism | .092 | .118 | 85.0 | .441 | |
| Rater Openness | -.133 | .117 | 85.0 | .258 | |
| Tetrad-mean Agreeableness | .141 | .312 | 29.9 | .655 | |
| Tetrad-mean Conscientiousness | -.107 | .280 | 32.8 | .703 | |
| Tetrad-mean Extraversion | -.244 | .367 | 30.0 | .512 | |
| Tetrad-mean Neuroticism | -.432 | .281 | 35.1 | .133 | |
| Tetrad-mean Openness | .304 | .259 | 37.1 | .248 | |
| C. Group Link-Anchor | Intercept | -.038 | .146 | 83.7 | .798 |
| Rater Agreeableness | .020 | .084 | 105.3 | .808 | |
| Rater Conscientiousness | .115 | .086 | 110.0 | .185 | |
| Rater Extraversion | -.011 | .100 | 109.4 | .911 | |
| Rater Neuroticism | .067 | .093 | 111.6 | .474 | |
| Rater Openness | -.017 | .089 | 111.7 | .851 | |
| Max Dyadic Rapport | .388 | .167 | 111.5 | .022* | |
| Min Dyadic Rapport | .248 | .096 | 111.9 | .011* | |
| *p < .05, **p < .01, ***p < .001. df estimated via Satterthwaite approximation. |
|||||
C.3 Study 3
| Model | Predictor | β | SE | df | p |
|---|---|---|---|---|---|
| Dyadic – M1 OCEAN | Intercept | .090 | .108 | 64.3 | .408 |
| Rater Agreeableness | -.015 | .103 | 57.4 | .888 | |
| Rater Conscientiousness | .070 | .112 | 57.1 | .534 | |
| Rater Extraversion | .087 | .123 | 59.1 | .486 | |
| Rater Neuroticism | -.027 | .112 | 57.6 | .812 | |
| Rater Openness | -.272 | .121 | 60.6 | .029* | |
| Target Agreeableness | .150 | .067 | 50.5 | .030* | |
| Target Conscientiousness | -.156 | .073 | 49.3 | .036* | |
| Target Extraversion | .114 | .083 | 52.7 | .174 | |
| Target Neuroticism | .079 | .073 | 51.0 | .284 | |
| Target Openness | -.039 | .084 | 53.9 | .642 | |
| Dyadic – M2 Synchrony | Intercept | .111 | .108 | 69.0 | .311 |
| Deep Video Sync | -.008 | .052 | 135.3 | .873 | |
| Deep Audio Sync | -.065 | .068 | 136.2 | .344 | |
| Facial Mimicry (AU12) | .038 | .082 | 148.5 | .642 | |
| Head Pose Sync | .005 | .061 | 140.5 | .933 | |
| Pitch Sync | -.383 | .099 | 150.4 | < .001*** | |
| Intensity Sync | .187 | .114 | 139.6 | .103 | |
| Dyadic – M3 Full | Intercept | .097 | .108 | 62.0 | .371 |
| Rater Agreeableness | -.040 | .105 | 57.2 | .704 | |
| Rater Conscientiousness | .100 | .113 | 56.4 | .380 | |
| Rater Extraversion | .046 | .125 | 58.9 | .714 | |
| Rater Neuroticism | -.034 | .115 | 59.3 | .769 | |
| Rater Openness | -.236 | .124 | 60.6 | .061 | |
| Target Agreeableness | .108 | .066 | 50.4 | .108 | |
| Target Conscientiousness | -.106 | .071 | 49.2 | .141 | |
| Target Extraversion | .076 | .082 | 52.1 | .357 | |
| Target Neuroticism | .082 | .074 | 51.2 | .270 | |
| Target Openness | -.017 | .083 | 54.0 | .835 | |
| Deep Video Sync | -.001 | .053 | 128.8 | .991 | |
| Deep Audio Sync | -.045 | .071 | 126.5 | .534 | |
| Facial Mimicry (AU12) | .044 | .083 | 135.6 | .596 | |
| Head Pose Sync | -.011 | .062 | 137.3 | .858 | |
| Pitch Sync | -.328 | .103 | 150.2 | .002** | |
| Intensity Sync | .146 | .119 | 131.0 | .222 | |
| Dyadic – M4 Moderation | Intercept | .101 | .108 | 62.2 | .356 |
| Rater Agreeableness | -.047 | .106 | 57.6 | .659 | |
| Rater Conscientiousness | .102 | .114 | 56.7 | .375 | |
| Rater Extraversion | .093 | .127 | 60.4 | .468 | |
| Rater Neuroticism | -.027 | .116 | 59.6 | .814 | |
| Rater Openness | -.238 | .125 | 60.9 | .061 | |
| Target Agreeableness | .093 | .065 | 53.1 | .160 | |
| Target Conscientiousness | -.102 | .068 | 48.2 | .141 | |
| Target Extraversion | .100 | .080 | 52.0 | .216 | |
| Target Neuroticism | .078 | .071 | 49.7 | .273 | |
| Target Openness | -.016 | .080 | 52.4 | .839 | |
| Deep Video Sync | .032 | .053 | 125.1 | .541 | |
| Deep Audio Sync | -.041 | .070 | 119.3 | .555 | |
| Facial Mimicry (AU12) | .108 | .084 | 120.4 | .202 | |
| Head Pose Sync | .035 | .063 | 126.6 | .578 | |
| Pitch Sync | -.205 | .109 | 134.1 | .062 | |
| Intensity Sync | .133 | .117 | 127.3 | .257 | |
| Rater Openness × Pitch Sync | -.126 | .063 | 132.5 | .046* | |
| Target Agreeableness × Pitch Sync | .036 | .057 | 134.6 | .527 | |
| Target Conscientiousness × Pitch Sync | -.082 | .056 | 141.8 | .149 | |
| Group – M1 OCEAN | Intercept | -.035 | .140 | 58.0 | .805 |
| Rater Agreeableness | .052 | .142 | 58.0 | .716 | |
| Rater Conscientiousness | .196 | .155 | 58.0 | .214 | |
| Rater Extraversion | .034 | .169 | 58.0 | .842 | |
| Rater Neuroticism | .033 | .155 | 58.0 | .830 | |
| Rater Openness | -.120 | .164 | 58.0 | .468 | |
| Group – M2 Synchrony (tetrad-mean) | Intercept | -.029 | .133 | 57.0 | .829 |
| Deep Video Sync | -.832 | .574 | 57.0 | .153 | |
| Deep Audio Sync | .204 | .242 | 57.0 | .405 | |
| Facial Mimicry (AU12) | .129 | .233 | 57.0 | .584 | |
| Head Pose Sync | .506 | .419 | 57.0 | .232 | |
| Pitch Sync | -.248 | .424 | 57.0 | .561 | |
| Intensity Sync | -.050 | .281 | 57.0 | .860 | |
| Group – M3 Full (tetrad-mean sync) | Intercept | -.032 | .137 | 52.0 | .817 |
| Rater Agreeableness | .010 | .145 | 52.0 | .944 | |
| Rater Conscientiousness | .196 | .158 | 52.0 | .220 | |
| Rater Extraversion | .000 | .177 | 52.0 | .999 | |
| Rater Neuroticism | .061 | .156 | 52.0 | .696 | |
| Rater Openness | -.138 | .168 | 52.0 | .415 | |
| Deep Video Sync | -.798 | .618 | 52.0 | .202 | |
| Deep Audio Sync | .263 | .261 | 52.0 | .318 | |
| Facial Mimicry (AU12) | .110 | .244 | 52.0 | .653 | |
| Head Pose Sync | .438 | .445 | 52.0 | .329 | |
| Pitch Sync | -.104 | .456 | 52.0 | .820 | |
| Intensity Sync | -.086 | .305 | 52.0 | .779 | |
| *p < .05, **p < .01, ***p < .001. df estimated via Satterthwaite approximation. Conditional $R^2_c$ by model: Dyadic M1 = .655, M2 = .706, M3 = .697; Group M1 = .050, M2 = .154, M3 = .192 (equal to $R^2_m$ in the group models, reflecting near-zero session-level residual variance beyond the fixed effects). |
|||||
Footnote
1The dataset is planned for release once the full 30 tetrads are processed with transcriptions and externally assessed rapport annotations.
2These correspond to items G1 (IIS) and d1 (IOS) in the questionnaire in Appendix A.
3As described in section 4, three additional items were added partway through data collection (after tetrad 6) to improve reliability. The dyadic EFAs therefore use directed dyadic ratings from disjoint subsets (6 and 24 tetrads; N = 72 and 288), while the group EFAs use the 120 participant-level ratings. The 6-item dyadic solution rests on fewer independent raters and is treated as preliminary.
4The intercept-only baselines retain the highest cross-validated R2 for both outcomes. This reflects the nested design: out-of-fold, the random intercepts carrying most of the systematic variance revert to the grand mean, making RMSE the more informative criterion here.
This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
ICMI '26, Napoli, Italy
© 2026 Copyright held by the owner/author(s).
ACM ISBN 979-8-4007-2318-6/26/10.
DOI: https://doi.org/10.1145/3776574.3831156