Hidden Profile Decision Making in Multi-Agent LLM Groups
DOI: https://doi.org/10.1145/3785651.3831525
CSCW Companion '26: Companion of the Computer-Supported Cooperative Work and Social Computing, Salt Lake City, UT, USA, October 2026
The hidden profile paradigm is a foundational tool of CSCW research on computer-mediated group decision making: a group can identify the collectively optimal option only if its members surface their uniquely held information, and asynchronous communication is known to improve that pooling in human groups. As multi-agent LLM systems take on group decision-support roles, it is unclear whether the same channel-level interventions improve their performance. We replicate the hidden profile paradigm with groups of three LLM agents from four families (Kimi K2, GPT-OSS 120B, Gemini 2.5 Flash, Llama 4 Scout 17B) in a two-by-two design crossing communication mode (synchronous, asynchronous) with information distribution (shared, hidden). Across 219 valid sessions, only one model-by-mode cell exceeds the three-candidate chance baseline of 33 percent: Kimi K2 under asynchronous communication, at 87 percent correct (Fisher's exact, two-sided p equals 0.002), an early-stage result from 15 sessions that awaits larger-n replication; two other models share private information at rates above 75 percent but never aggregate it into the correct decision. In this task and these model families, information sharing and information integration appear to be dissociated capabilities, and we develop the implications for CSCW systems that scaffold AI-mediated distributed decision making.
ACM Reference Format:
Carlos Toxtli and Manuel Delaflor. 2026. Hidden Profile Decision Making in Multi-Agent LLM Groups. In Companion of the Computer-Supported Cooperative Work and Social Computing (CSCW Companion '26), October 10--14, 2026, Salt Lake City, UT, USA. ACM, New York, NY, USA 5 Pages. https://doi.org/10.1145/3785651.3831525
1 Introduction
The hidden profile paradigm of Stasser and Titus [29] demonstrates that groups frequently fail to identify the collectively optimal option when the information needed to identify it is distributed across members rather than shared. The paradigm is foundational to CSCW because the choice of collaborative technology systematically changes whether the relevant information ever reaches the group's common ground: Dennis [10] showed that asynchronous electronic channels improve information pooling by reducing production blocking [11, 21], an effect corroborated by meta-analyses of CMC group performance [7, 22].
CSCW systems for distributed work inherited that finding as a design principle: give a distributed group a channel that surfaces uniquely held information and the group will approach the optimal collective decision. The principle implicitly assumes a human group whose limits on shared cognition can be relieved by the right channel; that assumption is being silently broken as multi-agent large-language-model (LLM) systems [12, 16, 19, 23] take on group decision-support roles [31]. Two empirical questions remain open: does the asynchronous-CMC advantage transfer to LLM-agent groups, and do LLM agents even reproduce the hidden-profile effect that justifies the channel-level intervention?
We address these questions with a controlled two-by-two factorial replication of the hidden profile paradigm across four LLM families. This is an exploratory, poster-stage study (one task domain, one group size, and four model checkpoints), so its findings are offered as early evidence and effect-size baselines rather than settled generalisations. Within that scope, our central finding is that information sharing and information integration were dissociated capabilities in the models we tested: the classical CSCW pattern of “provide a channel that surfaces unshared information” is necessary but not sufficient for agentic decision quality. We contribute (i) the first systematic replication of the hidden profile paradigm with multi-agent LLM groups under CMC conditions, (ii) evidence that sharing and integration are decoupled in the tested LLMs on this task, and (iii) design implications for CSCW systems that scaffold AI-mediated distributed decision making.
2 Background and Related Work
Hidden profiles and CMC. Stasser and Titus [29] showed that groups under-discuss uniquely held information; the effect is robust across group composition, time pressure, and discussion structure [13, 25, 30, 35], and information sharing predicts team performance only when the task requires integrating distinct pieces [22]. Channel characteristics substantially modulate the effect [7, 9, 27, 33]; Dennis [10] showed that asynchronous channels reduce production blocking by removing the single-speaker constraint, favouring the surfacing of unique information, while cautioning that better information exchange did not by itself produce better decisions.
LLMs and group-decision paradigms. Single-agent LLM replications [2, 5, 6, 15, 17] establish that an LLM can stand in for a human participant in constrained one-shot tasks, but not whether LLM groups reproduce collective phenomena; multi-agent LLM work [12, 16, 19, 23] documents non-trivial interaction dynamics but is not grounded in classical group-decision paradigms with effect-size benchmarks. A related line documents the gap between information availability and information use [8, 18, 34, 36]: models attend to relevant context yet fail to combine it correctly, and the hidden profile task is precisely the design that distinguishes those capabilities. Recent CHI and CSCW work on LLM-mediated group decision making [4, 20, 24, 26, 32] documents AI deliberation facilitators, structured turn-taking interfaces, and agent-augmented brainstorming; we fill the multi-agent hidden-profile gap and provide an effect-size baseline these systems can compare against [1, 9, 33].
Why transfer a human-cognition paradigm to LLM agents? The paradigm was built to probe human limits (working memory, evaluation apprehension, production blocking) that LLM agents do not obviously share [11, 29]. It transfers anyway because it operationalises the integration requirement at the task level: it separates whether the relevant information is in the group's joint context from whether the decision rule weights it appropriately, a distinction independent of working-memory limits. It thus diagnoses a system property (is the decision rule integration-faithful?) that CSCW builders care about even when human cognitive constraints no longer apply [1, 31]; the same construct-validity logic underwrites recent social-psychology probes of LLM behaviour [2, 15].
3 Method
Design and task. We use a two-by-two factorial design crossing communication mode (synchronous versus asynchronous) with information distribution (shared versus hidden profile); the shared-profile cells are a control in which every agent has every candidate's full attribute set. Groups of three LLM agents form a hiring committee with three candidates; the optimal candidate is identifiable only by combining all members’ private information. The setup follows the classical hiring-committee instantiation [13, 28, 29]: an agent reading the final transcript has access to everything needed for the optimal choice, even though no single agent had it at the start.
Session workflow. Figure 1 traces one session end to end: each agent privately receives its candidate profiles (in the hidden condition, shared attributes identical for all three agents plus unique attributes only it holds), the group discusses over three rounds under the assigned communication mode, surfaced facts accumulate in the shared transcript, each agent privately casts a vote in a separate deterministic call, and the modal vote is the group decision, scored against the designed optimum.
Channels. In the synchronous channel, agents take turns producing one short message at a time; the conversation is strictly serialised, mirroring human turn-taking. In the asynchronous channel each agent receives the full set of prior messages and produces its full contribution without the turn-taking constraint, operationalising Dennis's [10] production-blocking-relieving property. Single-shot parallel composition is one of several legitimate mappings of the manipulation [7, 27, 33], and it conflates the removal of turn-taking with two co-varying factors: asynchronous contributions are typically longer, and they arrive with different context available at composition time. A post-review length-only control (below) bounds the length factor. The shared-profile cells are the construct-validity check: decision quality there is near-ceiling under both modes, ruling out a degenerate operationalisation.
Length-only control (post-review). To separate message length from turn-taking removal, we added a sync-long condition after review: the synchronous protocol kept verbatim, but agents instructed to compose full-length contributions matching asynchronous volume. It covers the two checkpoints still servable at camera-ready time, GPT-OSS 120B (institutional vLLM endpoint) and Llama 4 Scout (Groq), 15 sessions per condition per model alongside fresh synchronous and asynchronous anchors, on a faithful reconstruction of the task materials; we report it as an informal bound.
Models, configuration, and measures. We test four instruction-tuned chat models: Kimi K2 (Moonshot), GPT-OSS 120B (OpenAI), Gemini 2.5 Flash (Google), and Llama 4 Scout 17B (Meta). Each cell is replicated 15 times per model (240 sessions, 219 valid; 21 GPT-OSS async sessions failed during message-buffer construction with provider-side errors). Models run at temperature 0.7, top-p 0.95, a 1024-token output cap, and a 1500-token role prompt; the buffer is appended verbatim, never summarised; vote extraction uses a separate temperature-0 structured-output call. Dependent measures follow hidden-profile conventions [29, 35]: decision quality (optimal-candidate selection), information-sharing rate (proportion of unique items surfaced), and discussion depth. Statistics: 95% Wilson CIs, Fisher's exact within-model contrasts, and a model × mode logistic regression as the omnibus test.
Prompts and reproducibility. Each agent's role prompt frames it as a hiring-committee member, supplies its candidate profile, and asks it to discuss and then vote; prompts are deliberately neutral with respect to the hidden-profile mechanism, so the experiment measures baseline rather than prompt-engineered integration. The group decision is the modal vote, ties broken by mention order (under four percent of sessions). All prompts, model identifiers, and substrate code are released; agents are stateless chat-style calls, so the replication transfers to any conversational LLM by swapping the model name.
4 Results
Strikingly model-dependent. Results are highly model-dependent (Table 1). In plain terms: of the fifteen hidden-profile discussions per cell, only Kimi K2 groups under asynchronous communication found the optimal candidate more often than guessing would predict (13 of 15 times), while the same model synchronously succeeded only 4 of 15 times, and no other model's groups beat guessing under either mode. The omnibus statistics are in the caption of Table 1; the only significant within-model sync-versus-async contrast is Kimi K2’s (Fisher's exact p = .002, + 60.0 percentage points), and Kimi K2 async is the only cell whose 95% CI lower bound (62.1%) clears the three-candidate chance baseline of 33.3%, replicating Dennis's human finding [10]. We flag the evidential weight explicitly: this headline cell rests on 15 sessions, so we report it as an early-stage result awaiting larger-n replication (the checkpoint was no longer servable at camera-ready time), and the GPT-OSS asynchronous cell retains only n = 5 valid sessions, so its zero rate should be read with particular caution. GPT-OSS 120B and Gemini 2.5 Flash never select the optimal candidate (0/15 sync hidden, 95% one-sided upper bound 20.4%); Llama 4 Scout 17B remains below chance in both modes. Across the failing cells, information-sharing rates are nonetheless high (GPT-OSS 78.2%, Gemini 76.4%, vs. Kimi K2’s 81.9%), which puts the surfaced information above the task's identifiability threshold in the typical session: the optimal candidate becomes identifiable once 3 of its 6 unique positives are on the table (5 in the worst case where both distractors’ unique positives are also surfaced), and the observed sharing rates imply an expected 4.6 to 4.9 of those 6 items surfaced per session. The aggregate hidden-profile optimal-decision rate is 21.8% (24/110), driven entirely by Kimi K2.
| Model | Sync (n = 15) | Async (n) | Sync → Async |
|---|---|---|---|
| Kimi K2 | 26.7% [10.9, 52.0] | 86.7% [62.1, 96.3] (n = 15) | + 60.0 pp** |
| Llama 4 Scout 17B | 20.0% [7.0, 45.2] | 26.7% [10.9, 52.0] (n = 15) | $+\phantom{0}6.7$ pp |
| GPT-OSS 120B | 0.0% [0.0, 20.4] | 0.0% [0.0, 43.4] (n = 5) | — |
| Gemini 2.5 Flash | 0.0% [0.0, 20.4] | 0.0% [0.0, 20.4] (n = 15) | — |
| Overall hidden-profile rate | 21.8% (24/110) | ||
| ** p < .01 Fisher's exact within-model contrast. Read: of 15 hidden-profile sessions per cell, the share whose modal vote selected the optimal candidate. |
|||
Sharing is dissociated from integration. GPT-OSS sessions that surface every private item end with a decision ignoring them, while Kimi K2 sessions with comparable sharing rates end correctly. The pattern mirrors the human hidden-profile finding [29, 35] but the mechanism differs: the LLM has no working-memory limit yet reweights reiterated information independent of its diagnostic value, consistent with broader LLM information-use failures [8, 15, 34]. The asynchronous advantage tracks this: for Kimi K2 asynchrony substantially improves decision quality, consistent with production blocking [10, 11], while for models that already fail synchronously it produces no improvement, so the channel effect operates conditional on baseline integration. Information-sharing rate is a misleading indicator of decision quality; integration is a distinct capability that systems builders must evaluate separately.
Anatomy of a failure. Table 2 makes the failure concrete: candidate A is optimal once unique information is pooled (8 positives against 6 and 6) but appears weakest on shared information alone (2 against 4 and 4), while candidate B pairs the highest shared-positive density with the mildest negatives. A failing session looks like this lightly condensed excerpt from a post-review GPT-OSS 120B synchronous session on the reconstructed materials, in which the group surfaced five of A’s six unique positives:
Round 3, agent Morgan: “[A:] great execution record (the $2M launch, 18 percent supply-cost cut, data-driven budgeting still in use) and solid mentorship (four promotions) [...] My current preference: Candidate B, the strongest mix of strategic education, proven cross-functional execution, and references.” All three agents vote B.
The group surfaces the decisive unique evidence, acknowledges it item by item (the same agent had raised A’s supply-cost and mentorship records itself in round 2), and then votes for the distractor anyway: the sharing step succeeds while the integration step fails. Transcript inspection adds a mechanism-level observation: failing sessions invent exculpatory details for the distractor (an unprompted restructuring excuse for B’s missed promotion, fabricated performance ratings), rationalising the shared-information prior rather than weighing the surfaced evidence.
| Candidate | Shared view | Pooled view | Designed role | Chosen in failures |
|---|---|---|---|---|
| A | 2 + / 2 − | 8 + / 2 − | optimal | 0% |
| B | 4 + / 2 − | 6 + / 2 − | shared-positive distractor | 78.3% |
| C | 4 + / 3 − | 6 + / 3 − | dominated | 21.7% |
Failure modes are structured, not random. Across the 60 failing GPT-OSS and Gemini sessions, the decision lands on B 47 times (78.3%, 95% CI [66.4, 86.9]) and on C 13 times, never on A; a χ2 test against the uniform baseline rejects random failure (χ2(2) = 42.9, p < .001). Llama 4 Scout shows the same direction (17 of 23 failures on B, 73.9%). These models fail by weighting the discussion's positive-cue density rather than the cues’ diagnostic value [29, 35]; designers of LLM-mediated decision support must discount redundant positive information explicitly at the interaction-design layer.
Length-only control bounds the message-length confound. In both models, sync-long raised message volume past asynchronous levels (Scout: 890 to 4,581 characters per message, vs. 3,093 async; GPT-OSS: 1,583 to 11,767, vs. 7,875) and pushed unique-item sharing to or past asynchronous levels (Scout .54 to .86, vs. .81; GPT-OSS .56 to .97, vs. .92), yet decision quality did not move relative to the short-message synchronous anchor (Scout 8/15 to 6/15, Fisher's exact p = .72; GPT-OSS 0/15 to 2/15, p = .48). Where an asynchronous advantage appears on these materials (Scout: 14/15 correct), it persists over the length-matched control (14/15 vs. 6/15, p = .005): message length, and the extra sharing it buys, does not produce the asynchronous effect; independent composition does. For GPT-OSS no channel manipulation rescues integration (0 to 3 of 15 in all conditions, all below chance), and its failures again converge on the distractor (B in 31 of 40, 77.5%, matching the original 78.3%). Three caveats: the control runs on reconstructed materials, on which Scout integrates better than in the original run (8/15 sync vs. 20% published), a reminder that hidden-profile difficulty is instantiation-sensitive; GPT-OSS was served by an institutional vLLM endpoint (its partial original-provider synchronous anchor, 2/14, is consistent); and the length-only bound for the Kimi K2 advantage, the headline effect, awaits that checkpoint being servable.
Control condition and failure inspection. The shared-profile cells serve as a construct-validity control: all four models reach near-perfect rates under both modes, isolating the hidden-profile losses as integration failures across distributed sources rather than ranking failures. The 21 GPT-OSS async API failures occurred during message-buffer construction, not at decision time; results are unchanged whether we drop these sessions or impute at either extreme (GPT-OSS rate remains under 5% in both counterfactuals).
5 Discussion
A socio-technical gap and a two-step model. The dissociation is, in Ackerman's terms [1], a textbook socio-technical gap: the system reliably supports the articulation work of getting distributed information into a shared transcript, but not the cooperative cognition the task requires. This motivates a two-step model of agentic decision quality, P(correct) = P(surface) · P(integrate | surfaced): decades of CMC research [7, 9, 10] target the first step and treat the second as a downstream consequence, but for the LLM groups we tested the two were independent and must be designed for separately.
Design implications: two CSCW interaction patterns. (1) Synthesiser-agent role. Separate the discussant agents from a dedicated synthesiser whose only job is to integrate the transcript at decision time, producing a structured candidate-by-attribute matrix (Table 2’s pooled view) before any vote is cast. Giving integration its own role makes it independently evaluable: the synthesiser can be benchmarked on transcript-to-matrix fidelity in isolation, exactly the capability our results show cannot be assumed from discussion quality [1, 22]. The failure analysis says what it must resist: weighting a candidate by the density of positive mentions rather than their diagnostic value. (2) Integration-audit UI. Render a panel listing, per candidate, the discussion-derived attributes and which agent surfaced them, so a human reviewer can see whether the final ranking matches the cumulative evidence [1, 15]. In our failing sessions this panel would have shown, at a glance, eight surfaced positives for the candidate the group did not choose; the audit surface converts a silent integration failure into a visible mismatch a human can veto. Two further directions, separately auditable surfaced-versus-weighted artefacts [1] and channel-matched model selection benchmarked on decision quality [19, 26, 32], we leave as future work.
Alternative explanations. Training-data contamination: our scenarios are not literal reproductions of any published stimulus set, and the failing models converge on the distractor rather than the optimal candidate, the opposite of what contamination-driven recognition would predict. Prompt sensitivity: we deliberately use neutral prompts, because prompt-engineered integration is precisely the design intervention CSCW systems need to build in, not something to assume away. Sampling variance: the 15-fold replication and the statistics in Table 1 bound it; the GPT-OSS and Gemini zero-rates have one-sided 95% upper bounds below 21%, excluding “lucky” integration.
6 Limitations and Conclusion
The study is exploratory in scope: a single hidden-profile task instance, a fixed group size of three, four model families, and 15 sessions per cell (5 in the GPT-OSS asynchronous cell after API failures). The headline Kimi K2 asynchronous advantage therefore awaits larger-n replication, and general claims about LLM integration capability would require validation across task domains, information structures, group sizes, and prompting strategies. The asynchronous operationalisation is one of several defensible mappings; the length-only control bounds its message-length confound, but alternative implementations might produce different outcomes. Hybrid human-agent groups [1, 14, 23], longer-horizon integration, retrieval augmentation, and adversarial members are left to future work. You can lead an LLM agent to information, but you cannot make it think. In the groups we tested, information sharing was necessary but not sufficient for collective decision quality, and the asynchronous-CMC advantage Dennis [10] established for human groups appeared only conditionally on baseline integration capability. CSCW systems that scaffold LLM-mediated distributed decision making should instrument integration as a first-class design target, not as a downstream consequence of better channels.
Generative AI Use Disclosure. All text in this paper was written by the authors; an AI writing assistant (e.g., Grammarly, Claude) was used solely for grammar, spelling, and clarity checks. The content and intellectual contributions are entirely the human authors’. The multi-agent LLM groups described in the methodology are the experimental subjects of the study, not authors of any portion of the manuscript.
Acknowledgments
This research used in part resources on the Palmetto 2 cluster at Clemson University [3] under National Science Foundation awards MRI 1228312, II NEW 1405767, MRI 1725573, and MRI 2018069. The views expressed in this article do not necessarily represent the views of NSF or the United States government.
References
- Mark S. Ackerman. 2000. The Intellectual Challenge of CSCW: The Gap Between Social Requirements and Technical Feasibility. Human–Computer Interaction 15, 2–3 (2000), 179–203. https://doi.org/10.1207/S15327051HCI1523_5
- Gati V. Aher, Rosa I. Arriaga, and Adam Tauman Kalai. 2023. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. In Proceedings of the 40th International Conference on Machine Learning (ICML 2023)(Proceedings of Machine Learning Research, Vol. 202). PMLR, Honolulu, HI, USA, 337–371.
- Asher Antao, James Daly Burton, Douglas Dawson, Jill Gemmill, Zachary Gerstener, Ben Godfrey, Scott Groel, Zach Jordan, Becky Ligon, Dane Smith, et al. 2024. Modernizing Clemson University's Palmetto Cluster: Lessons Learned from 17 Years of HPC Administration. In Practice and Experience in Advanced Research Computing 2024: Human Powered Computing. ACM, New York, NY, USA, 1–9.
- Lisa P. Argyle, Christopher A. Bail, Ethan C. Busby, Joshua R. Gubler, Thomas Howe, Christopher Rytting, Taylor Sorensen, and David Wingate. 2023. Leveraging AI for Democratic Discourse: Chat Interventions Can Improve Online Political Conversations at Scale. Proceedings of the National Academy of Sciences 120, 41 (2023), e2311627120. https://doi.org/10.1073/pnas.2311627120
- Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023. Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis 31, 3 (2023), 337–351. https://doi.org/10.1017/pan.2023.2
- Christopher A. Bail. 2024. Can Generative Artificial Intelligence Improve Social Science?Proceedings of the National Academy of Sciences 121, 21 (2024), e2314021121. https://doi.org/10.1073/pnas.2314021121
- Boris B. Baltes, Marcus W. Dickson, Michael P. Sherman, Cara C. Bauer, and Jacqueline S. LaGanke. 2002. Computer-Mediated Communication and Group Decision Making: A Meta-Analysis. Organizational Behavior and Human Decision Processes 87, 1 (2002), 156–179. https://doi.org/10.1006/obhd.2001.2961
- Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models Are Few-Shot Learners. In Advances in Neural Information Processing Systems 33 (NeurIPS 2020). Curran Associates, Inc., Red Hook, NY, USA, 1877–1901.
- Richard L. Daft and Robert H. Lengel. 1986. Organizational Information Requirements, Media Richness and Structural Design. Management Science 32, 5 (1986), 554–571. https://doi.org/10.1287/mnsc.32.5.554
- Alan R. Dennis. 1996. Information Exchange and Use in Group Decision Making: You Can Lead a Group to Information, but You Can't Make It Think. MIS Quarterly 20, 4 (1996), 433–457. https://doi.org/10.2307/249563
- Michael Diehl and Wolfgang Stroebe. 1987. Productivity Loss in Brainstorming Groups: Toward the Solution of a Riddle. Journal of Personality and Social Psychology 53, 3 (1987), 497–509. https://doi.org/10.1037/0022-3514.53.3.497
- Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2024. Improving Factuality and Reasoning in Language Models Through Multiagent Debate. In Proceedings of the 41st International Conference on Machine Learning (ICML 2024)(Proceedings of Machine Learning Research, Vol. 235). PMLR, Vienna, Austria, 11733–11763.
- Tobias Greitemeyer and Stefan Schulz-Hardt. 2003. Preference-Consistent Evaluation of Information in the Hidden Profile Paradigm: Beyond Group-Level Explanations for the Dominance of Shared Information in Group Decisions. Journal of Personality and Social Psychology 84, 2 (2003), 322–339. https://doi.org/10.1037/0022-3514.84.2.322
- Jonathan Grudin. 1994. Computer-Supported Cooperative Work: History and Focus. Computer 27, 5 (1994), 19–26. https://doi.org/10.1109/2.291294
- Thilo Hagendorff. 2023. Machine Psychology: Investigating Emergent Capabilities and Behavior in Large Language Models Using Psychological Methodology. arXiv preprint arXiv:2303.13988. https://doi.org/10.48550/arXiv.2303.13988
- Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. 2024. MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework. In The Twelfth International Conference on Learning Representations (ICLR 2024). OpenReview.net, Vienna, Austria, 28 pages.
- John J. Horton. 2023. Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?NBER Working Paper 31122. National Bureau of Economic Research, Cambridge, MA. https://doi.org/10.3386/w31122
- Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems 33 (NeurIPS 2020). Curran Associates, Inc., Red Hook, NY, USA, 9459–9474.
- Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. CAMEL: Communicative Agents for “Mind” Exploration of Large Language Model Society. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023). Curran Associates, Inc., Red Hook, NY, USA, 51991–52008.
- Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2024. Encouraging Divergent Thinking in Large Language Models Through Multi-Agent Debate. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024). Association for Computational Linguistics, Miami, FL, USA, 17889–17904.
- Lin Lu, Y. Connie Yuan, and Poppy Lauretta McLeod. 2012. Twenty-Five Years of Hidden Profiles in Group Decision Making: A Meta-Analysis. Personality and Social Psychology Review 16, 1 (2012), 54–75. https://doi.org/10.1177/1088868311417243
- Jessica R. Mesmer-Magnus and Leslie A. DeChurch. 2009. Information Sharing and Team Performance: A Meta-Analysis. Journal of Applied Psychology 94, 2 (2009), 535–546. https://doi.org/10.1037/a0013773
- Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST 2023). ACM, New York, NY, USA, Article 2, 22 pages. https://doi.org/10.1145/3586183.3606763
- Joon Sung Park, Lindsay Popowski, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2022. Social Simulacra: Creating Populated Prototypes for Social Computing Systems. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology (UIST 2022). ACM, New York, NY, USA, Article 74, 18 pages. https://doi.org/10.1145/3526113.3545616
- Stefan Schulz-Hardt, Felix C. Brodbeck, Andreas Mojzisch, Rudolf Kerschreiter, and Dieter Frey. 2006. Group Decision Making in Hidden Profile Situations: Dissent as a Facilitator for Decision Quality. Journal of Personality and Social Psychology 91, 6 (2006), 1080–1093. https://doi.org/10.1037/0022-3514.91.6.1080
- Omar Shaikh, Valentino Chai, Michele J. Gelfand, Diyi Yang, and Michael S. Bernstein. 2024. Rehearsal: Simulating Conflict to Teach Conflict Resolution. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI 2024). ACM, New York, NY, USA, Article 920, 20 pages. https://doi.org/10.1145/3613904.3642159
- Lee Sproull and Sara Kiesler. 1986. Reducing Social Context Cues: Electronic Mail in Organizational Communication. Management Science 32, 11 (1986), 1492–1512. https://doi.org/10.1287/mnsc.32.11.1492
- Garold Stasser and Dennis Stewart. 1992. Discovery of Hidden Profiles by Decision-Making Groups: Solving a Problem Versus Making a Judgment. Journal of Personality and Social Psychology 63, 3 (1992), 426–434. https://doi.org/10.1037/0022-3514.63.3.426
- Garold Stasser and William Titus. 1985. Pooling of Unshared Information in Group Decision Making: Biased Information Sampling During Discussion. Journal of Personality and Social Psychology 48, 6 (1985), 1467–1478. https://doi.org/10.1037/0022-3514.48.6.1467
- Garold Stasser and William Titus. 1987. Effects of Information Load and Percentage of Shared Information on the Dissemination of Unshared Information During Group Discussion. Journal of Personality and Social Psychology 53, 1 (1987), 81–93. https://doi.org/10.1037/0022-3514.53.1.81
- Cass R. Sunstein. 2006. Infotopia: How Many Minds Produce Knowledge. Oxford University Press, New York.
- Michael Henry Tessler, Michiel A. Bakker, Daniel Jarrett, Hannah Sheahan, Martin J. Chadwick, Raphael Koster, Georgina Evans, Lucy Campbell-Gillingham, Tantum Collins, David C. Parkes, Matthew Botvinick, and Christopher Summerfield. 2024. AI Can Help Humans Find Common Ground in Democratic Deliberation. Science 386, 6719 (2024), eadq2852. https://doi.org/10.1126/science.adq2852
- Joseph B. Walther. 1996. Computer-Mediated Communication: Impersonal, Interpersonal, and Hyperpersonal Interaction. Communication Research 23, 1 (1996), 3–43. https://doi.org/10.1177/009365096023001001
- Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen. 2024. A Survey on Large Language Model Based Autonomous Agents. Frontiers of Computer Science 18, 6 (2024), 186345. https://doi.org/10.1007/s11704-024-40231-1
- Gwen M. Wittenbaum, Andrea B. Hollingshead, and Isabel C. Botero. 2004. From Cooperative to Motivated Information Sharing in Groups: Moving Beyond the Hidden Profile Paradigm. Communication Monographs 71, 3 (2004), 286–310. https://doi.org/10.1080/0363452042000299894
- Jintian Zhang, Xin Xu, Ningyu Zhang, Ruibo Liu, Bryan Hooi, and Shumin Deng. 2023. Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View. arXiv preprint arXiv:2310.02124. https://doi.org/10.48550/arXiv.2310.02124
This work is licensed under a Creative Commons Attribution 4.0 International License.
CSCW Companion '26, Salt Lake City, UT, USA
© 2026 Copyright held by the owner/author(s).
ACM ISBN 979-8-4007-2378-0/26/10.
DOI: https://doi.org/10.1145/3785651.3831525