메뉴
HN
Hacker News 12일 전

EEG 연구: 뇌, 두 음성 동시 처리 가능

IMP
7/10
핵심 요약

여러 사람이 동시에 말하는 환경에서 우리가 주의를 전환할 때, 뇌가 이전 화자와 새 화자의 음성을 일시적으로 동시에 처리한다는 사실이 뇌파(EEG) 기록을 통해 밝혀졌습니다. 연구진은 주의 전환 시 알파파(alpha wave) 감소를 통해 인지적 노력이 증가함을 확인했고, 대형 언어 모델(LLM)을 활용해 청각적 주의 전환 후 어휘 맥락이 초기화되는 시의적 메커니즘을 규명했습니다. 이는 복잡한 청취 환경 속에서 뇌가 어떻게 유연하게 음성 정보를 처리하는지 이해하는 데 중요한 기초 자료가 됩니다.

번역된 본문

본문: 논문 저자, 지표, 댓글, 언론 보도, 동료 평가, 독자 댓글, 그림, 그림, 초록: 여러 화자가 동시에 말하는 상황에서 성공적인 음성 의사소통을 위해서는 지속적인 주의력 유지와 빠른 주의 전환의 능숙한 결합이 필요합니다. 신경생리학 문헌은 지속적인 주의력의 신경학적 기반에 대해 상세한 통찰력을 제공하지만, 주의 전환이 어떻게 일어나는지에 대해서는 여전히 상당한 불확실성이 남아있습니다. 이 연구에서 우리는 몰입형 다중 화자 환경에서 청력이 정상인 성인의 뇌파(EEG) 기록을 사용하여 배경 소음 속에서 경쟁하는 두 음성 스트림의 신경적 인코딩을 측정했습니다. 참가자들은 15~30초마다 스트림 간의 주의를 전환하라는 큐를 받았습니다. 시간 응답 함수(Temporal Response Functions, TRF)를 통해 신경 추적을 평가하여 주의 초점의 안정적인 해독을 확인했습니다. 우리의 결과는 주의 전환 중 비대칭적인 이탈 및 참여 과정을 나타내며, 이전 대상에서 벗어나기 전에 새 대상 스트림의 신경 추적이 나타나 두 음성 스트림의 일시적인 동시 인코딩을 보여줍니다. 이러한 전환은 뇌파 알파파 강도의 감소와 밀접하게 일치하며, 이는 주의 전환의 다른 단계 동안 인지적 노력에 대한 정보를 제공합니다. 그런 다음 우리는 대형 언어 모델(Large Language Models)을 사용하여 구축된 네 가지 문맥 축적 전략을 비교하여 어휘 예측 메커니즘을 반영하는 피질 활동을 분리하여 주의 전환 후 어휘 맥락이 어떻게 업데이트되는지 결정했습니다. 우리의 연구 결과는 청각적 주의 전환의 기반이 되는 시간적, 맥락적 메커니즘을 모두 규명하며, 청취자가 주의를 전환한 후 어휘 맥락에서 재설정을 수행할 가능성을 시사합니다. 이 연구는 동적 주의 재할당에 초점을 맞추어 복잡한 청취 환경에서 유연한 음성 처리를 위한 뇌의 능력에 대한 통찰력을 제공합니다. 인용: Carta S, Aličković E, Zaar J, López Valdés A, Di Liberto GM (2026) 주의 전환 중 인간 대뇌 피질에서 경쟁하는 음성 스트림이 동시에 표현됨. PLoS Biol 24(7): e3003876. https://doi.org/10.1371/journal.pbio.3003876 학술 편집자: Manuel S. Malmierca, 살라망카 대학교, 스페인 접수일: 2025년 7월 3일; 승인일: 2026년 6월 12일; 게재일: 2026년 7월 16일 저작권: © 2026 Carta et al. 이 글은 크리에이티브 커먼즈 저작자표시 라이선스에 따라 배포되는 오픈 액세스 기사로, 원저작자와 출처가 표시되는 한 모든 매체에서 무제한 사용, 배포, 복제가 허용됩니다. 데이터 가용성: 본 원고에 보고된 연구 결과를 뒷받침하는 모든 데이터는 제한 없이 무료로 액세스할 수 있습니다. 전처리된 뇌파(EEG) 데이터 세트, 결과 분석 파일 및 분석 코드는 공개 저장소인 Zenodo(https://zenodo.org/records/20569817)에서 공개적으로 사용할 수 있습니다. 뇌파 기록은 연속 이벤트 신경 데이터(CND) 형식 표준에 따라 제공됩니다. 관련 음성 자극은 동일한 저장소의 STIMULI 폴더에서도 찾을 수 있습니다. 자금 지원: S.C., A.L.V., 및 G.D.L.은 윌리엄 뎀판트 재단(https://www.williamdemantfonden.dk/, 보조금 21-0628 및 22-0552)과 아일랜드 연구 위원회(https://www.researchireland.ie/, 보조금 번호 18/CRT/6223)의 지원을 받았습니다. G.D.L.은 추가적으로 트리니티 칼리지 더블린의 ADAPT, 아일랜드 연구 위원회 인공지능 기반 디지털 콘텐츠 기술 센터(https://www.adaptcentre.ie/)에서 아일랜드 연구 위원회의 재정적 지원을 받아 이 연구를 수행했습니다 [보조금 13/RC/2106_P2]. 자금 지원자는 연구 설계, 데이터 수집 및 분석, 출판 결정 또는 원고 작성에 어떠한 역할도 하지 않았습니다. 이해상충: 저자는 이해상충이 없음을 선언했습니다. 약어: 뇌전도(EEG), 안전도(EOG), 근전도(EMG), 사건 관련 스펙트럼 교란(ERSP), 개두부 내 뇌전도(iEEG), 거짓 발견율(FDR), 기능적 자기공명영상(fMRI)

원문 보기
원문 보기 (영어)
Article Authors Metrics Comments Media Coverage Peer Review Reader Comments Figures Figures Abstract Successful speech communication in multi-talker scenarios requires a skillful combination of sustained attention and rapid attention switching. While the neurophysiology literature offers detailed insights into the neural underpinnings of sustained attention, there remains considerable uncertainty on how attention switching takes place. In this study, using EEG recordings from normal-hearing adults in an immersive multi-talker environment, we measured the neural encoding of two competing speech streams amid background babble. Participants were cued to switch attention between streams every 15–30 s. Neural tracking was assessed via Temporal Response Functions (TRF), confirming reliable decoding of attentional focus. Our results indicate asymmetric disengagement and engagement processes during attention switches, where the neural tracking of the new target stream emerges before disengaging from the previous target, revealing a transient simultaneous encoding of two speech streams. That transition was closely mirrored by a reduction in EEG alpha power, informing on the cognitive effort during different phases of the attention switch. We then isolated cortical activity reflecting lexical prediction mechanisms to determine how lexical context is updated after an attention switch, comparing four context-accumulation strategies that were constructed using Large Language Models. Our findings elucidate both the temporal and contextual mechanisms underlying auditory attention shifts, pointing to the possibility that listeners carry out a reset in lexical context after switching attention. By focusing on dynamic attentional reallocation, this study offers insights into the brain’s capacity for flexible speech processing in complex listening environments. Citation: Carta S, Aličković E, Zaar J, López Valdés A, Di Liberto GM (2026) Competing speech streams are simultaneously represented in the human cortex during attention switching. PLoS Biol 24(7): e3003876. https://doi.org/10.1371/journal.pbio.3003876 Academic Editor: Manuel S. Malmierca, Universidad de Salamanca, SPAIN Received: July 3, 2025; Accepted: June 12, 2026; Published: July 16, 2026 Copyright: © 2026 Carta et al. This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Data Availability: All data supporting the findings reported in this manuscript are freely accessible without restriction. The EEG pre-processed dataset, the resulting analysis files, and the analysis code are publicly available on the open repository Zenodo ( https://zenodo.org/records/20569817 ). The EEG recordings are provided following the Continuous-event Neural Data (CND) format standard. The associated speech stimuli can also be found in the same repository, within the STIMULI folder. Funding: S.C., A.L.V., and G.D.L. were supported by the William Demant Fonden ( https://www.williamdemantfonden.dk/ ), under grants 21-0628 and 22-0552, and by Taighde Éireann – Research Ireland ( https://www.researchireland.ie/ ) under grant No. 18/CRT/6223. G.D.L. additionally conducted this research with the financial support of Research Ireland at ADAPT, the Research Ireland Centre for AI-Driven Digital Content Technology ( https://www.adaptcentre.ie/ ) at Trinity College Dublin [grant 13/RC/2106_P2]. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Competing interests: The authors have declared that no competing interests exist. Abbreviations: EEG, electroencephalography; EOG, electro-oculography; EMG, electro-myography; ERSP, event-related spectral perturbation; iEEG, intra-cranial electroencephalography; FDR, false discovery rate; fMRI, functional magnetic resonance imaging; ICA, Independent Component Analysis; IQR, interquartile range; LLM, large language model; MEG, magnetoencephalography; PSD, power spectral density; RMS, root-mean-squared; SE, standard error; SEM, standard error of the mean; SNR, signal-to-noise ratio; SPL, sound pressure level; TRF, Temporal Response Functions Introduction To understand speech in multi-talker environments, listeners single out the target speaker from competing sound streams [ 1 – 3 ]. The neurophysiology of this selective attention process has been widely studied with simulated cocktail-party scenarios [ 4 , 5 ], shedding light on how our brains segregate a target stream from competing speech streams, and enabling the transformation of the target speech into linguistic meaning. While the extent to which masker speech streams are processed remains highly debated [ 6 – 8 ], there is no doubt that there are considerable differences between the processing of target and masker speech, which have been measured with various technologies, such as non-invasive electroencephalography (EEG) [ 1 , 9 ], intra-cranial electroencephalography (iEEG) [ 10 ], magnetoencephalography (MEG) [ 3 , 11 ] and functional magnetic resonance imaging (fMRI) [ 12 , 13 ]. That work could pinpoint precise loci in the auditory cortical areas where that segregation emerges [ 14 ] as well as measuring the substantial (but not total) suppression of linguistic processing for the masker speech [ 1 , 15 – 17 ]. However, neurophysiology literature in this field has almost entirely focused on sustained attention tasks [ 2 , 10 ], leaving considerable uncertainty on the neural underpinnings of attention switching. Dynamic switching paradigms have been widely used in the domain of cognitive control studies to probe for cognitive flexibility and cognitive stability [ 18 ]. In those experiments, participants are often required to flexibly adapt their behavioral response depending on new instructions, initiating a task-switch [ 19 – 21 ]. For example, given a single digit, they are required to classify it either based on parity, i.e., whether it is even or odd, or based on relative magnitude, i.e., whether the digit is greater than or less than 5 [ 22 ]. In these paradigms, the switch-cost is the increase in reaction time or error rate when switching from one task to the other. Similar behavioral paradigms have also involved simple speech stimuli in multi-talker settings [ 23 – 25 ]. However, the main interest of those tightly controlled experiments was to model the process of target speech selection as one particular instance of a task-switching problem, i.e., target stream selection could either depend on spatial location or voice identity [ 23 ], rather than focusing on the dynamic aspect of attention re-allocation per se in naturalistic multi-talker scenarios. As such, very little is known on how a flexible reorienting of attention might impact speech processing of continuous competing streams. In recent speech neurophysiology research, experimental paradigms have started to include switches of attention as a tool towards tailored EEG/MEG methodological advances in the domain of attention decoding [ 26 , 27 ], or to investigate how sustained speech attention unfolds for moving auditory objects [ 28 ]. However, to the best of our knowledge, only one previous study has specifically focused on the neurophysiology of attention switching in multi-talker scenarios, relating the neural encoding of speech during attentional re-orienting with EEG alpha activity and pupil dilation dynamics [ 29 ]. Those findings proved that the neurophysiology of attention switching can be studied non-invasively. Building on that work, our study sheds light on the exact neural dynamics supporting the steering of attention between two competing speech streams, disengaging from the previous target stream while engaging to the new one. In this study, we measure the neural encoding of speech using a range of encoding wind