Scale matters: Large language models with billions (rather than millions) of parameters better match neural representations of natural language.

Journal

bioRxiv : the preprint server for biology

ISSN: 2692-8205

Titre abrégé: bioRxiv

Pays: United States

ID NLM: 101680187

Informations de publication

Date de publication:
02 Jul 2024

Historique:

medline: 15 7 2024

pubmed: 15 7 2024

entrez: 15 7 2024

Statut: epublish

Résumé

Recent research has used large language models (LLMs) to study the neural basis of naturalistic language processing in the human brain. LLMs have rapidly grown in complexity, leading to improved language processing capabilities. However, neuroscience researchers haven't kept up with the quick progress in LLM development. Here, we utilized several families of transformer-based LLMs to investigate the relationship between model size and their ability to capture linguistic information in the human brain. Crucially, a subset of LLMs were trained on a fixed training set, enabling us to dissociate model size from architecture and training set size. We used electrocorticography (ECoG) to measure neural activity in epilepsy patients while they listened to a 30-minute naturalistic audio story. We fit electrode-wise encoding models using contextual embeddings extracted from each hidden layer of the LLMs to predict word-level neural signals. In line with prior work, we found that larger LLMs better capture the structure of natural language and better predict neural activity. We also found a log-linear relationship where the encoding performance peaks in relatively earlier layers as model size increases. We also observed variations in the best-performing layer across different brain regions, corresponding to an organized language processing hierarchy.

Identifiants

DOI: 10.1101/2024.06.12.598513 PMID: 39005394 PMC: PMC11244877

pubmed: 39005394

doi: 10.1101/2024.06.12.598513

pmc: PMC11244877

pii:

doi:

Types de publication

Journal Article Preprint

Langues

eng

Scale matters: Large language models with billions (rather than millions) of parameters better match neural representations of natural language.

Journal

Informations de publication

Résumé

Identifiants

Types de publication

Langues

Auteurs

Zhuoqiao Hong (Z)

Haocheng Wang (H)

Zaid Zada (Z)

Harshvardhan Gazula (H)

David Turner (D)

Bobbi Aubrey (B)

Leonard Niekerken (L)

Werner Doyle (W)

Sasha Devore (S)

Patricia Dugan (P)

Daniel Friedman (D)

Orrin Devinsky (O)

Adeen Flinker (A)

Uri Hasson (U)

Samuel A Nastase (SA)

Ariel Goldstein (A)

Classifications MeSH