Artificial Intelligence and Human Language Cognition
An Integrative Comparative Framework
Keywords:
artificial intelligence; human language cognition; large language models; grounding; multimodality; intentionality; conceptual reviewAbstract
This article presents an integrative conceptual review comparing contemporary artificial intelligence language systems with human language cognition. A targeted, non-systematic search of foundational and recent literature published through July 2026 covered linguistics, psycholinguistics, cognitive science, philosophy of mind, natural language processing, multimodal learning, interpretability, and embodied robotics. Sources were selected based on their direct relevance to five recurring comparative dimensions: representation, acquisition and adaptation, processing, grounding and interaction, and intentionality, consciousness, and responsibility. The synthesis distinguishes functional linguistic competence from representational grounding, phenomenal consciousness, and moral responsibility, thereby avoiding the assumption that fluent performance either proves or refutes human-like understanding. The resulting framework treats AI-human differences as continua rather than fixed binaries and differentiates text-only language models from multimodal, tool-using, interactive, and embodied systems. Current evidence shows substantial convergence in formal language performance, prediction, abstraction, and some internally represented, world-relevant structures. However, it also reveals persistent differences in developmental history, sensorimotor coupling, online social learning, reliability profiles, and the evidence available for subjective experience or autonomous responsibility. These differences support a cautious model of cognitive complementarity: task allocation should depend on capability, verifiability, stakes, contextual sensitivity, and accountability rather than on broad claims that either humans or AI are categorically superior. The framework is interpretive and revisable. It would require revision if future systems demonstrate robust causal grounding, durable autonomous goals, cross-context learning from embodied interaction, or empirically credible indicators of consciousness.
References
Barsalou, L. W. (1999). Perceptual symbol systems. Behavioral and Brain Sciences, 22(4), 577-660.
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610-623. https://doi.org/10.1145/3442188.3445922
Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5185-5198. https://doi.org/10.18653/v1/2020.acl-main.463
Binz, M., Akata, E., Bethge, M., et al. (2025). A foundation model to predict and capture human cognition. Nature, 644, 1002-1009. https://doi.org/10.1038/s41586-025-09215-4
Brohan, A., Brown, N., Carbajal, J., et al. (2023). RT-2: Vision-language-action models transfer web knowledge to robotic control. arXiv. https://doi.org/10.48550/arXiv.2307.15818
Brown, T. B., Mann, B., Ryder, N., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877-1901.
Butlin, P., Long, R., Elmoznino, E., et al. (2023). Consciousness in artificial intelligence: Insights from the science of consciousness. arXiv. https://doi.org/10.48550/arXiv.2308.08708
Dentella, V., Gunther, F., Murphy, E., Marcus, G., & Leivada, E. (2024). Testing AI on language comprehension tasks reveals insensitivity to underlying meaning. Scientific Reports, 14, 28083. https://doi.org/10.1038/s41598-024-79531-8
Driess, D., Xia, F., Sajjadi, M. S. M., et al. (2023). PaLM-E: An embodied multimodal language model. Proceedings of the 40th International Conference on Machine Learning, 202, 8469-8488.
Du, Z., Zeng, A., Dong, Y., & Tang, J. (2024). Understanding emergent abilities of language models from the loss perspective. Advances in Neural Information Processing Systems, 37, 53138-53167. https://doi.org/10.52202/079017-1683
Fedorenko, E., Piantadosi, S. T., & Gibson, E. A. F. (2024). Language is primarily a tool for communication rather than thought. Nature, 630, 575-586. https://doi.org/10.1038/s41586-024-07522-w
Feng, J., & Steinhardt, J. (2024). How do language models bind entities in context? International Conference on Learning Representations.
Ferreira, F., Bailey, K. G. D., & Ferraro, V. (2002). Good-enough representations in language comprehension. Current Directions in Psychological Science, 11(1), 11-15. https://doi.org/10.1111/1467-8721.00158
Firth, J. R. (1957). A synopsis of linguistic theory, 1930-1955. In Studies in Linguistic Analysis (pp. 1-32). Philological Society.
Glenberg, A. M., & Kaschak, M. P. (2002). Grounding language in action. Psychonomic Bulletin & Review, 9(3), 558-565.
Grice, H. P. (1957). Meaning. The Philosophical Review, 66(3), 377-388.
Gurnee, W., & Tegmark, M. (2024). Language models represent space and time. International Conference on Learning Representations.
Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1-3), 335-346.
Harris, Z. S. (1954). Distributional structure. Word, 10(2-3), 146-162.
Hu, J., Mahowald, K., Lupyan, G., Ivanova, A., & Levy, R. (2024). Language models align with human judgments on key grammatical constructions. Proceedings of the National Academy of Sciences, 121(36), e2400917121. https://doi.org/10.1073/pnas.2400917121
Johnson, M. (1987). The body in the mind: The bodily basis of meaning, imagination, and reason. University of Chicago Press.
Just, M. A., & Carpenter, P. A. (1992). A capacity theory of comprehension: Individual differences in working memory. Psychological Review, 99(1), 122-149.
Lakoff, G. (1987). Women, fire, and dangerous things: What categories reveal about the mind. University of Chicago Press.
Lakoff, G., & Johnson, M. (1980). Metaphors we live by. University of Chicago Press.
Langacker, R. W. (1987). Foundations of cognitive grammar: Volume 1, Theoretical prerequisites. Stanford University Press.
Levelt, W. J. M. (1989). Speaking: From intention to articulation. MIT Press.
Loftus, E. F. (2005). Planting misinformation in the human mind: A 30-year investigation of the malleability of memory. Learning & Memory, 12(4), 361-366. https://doi.org/10.1101/lm.94705
Mahowald, K., Ivanova, A. A., Blank, I. A., Kanwisher, N., Tenenbaum, J. B., & Fedorenko, E. (2024). Dissociating language and thought in large language models. Trends in Cognitive Sciences, 28(6), 517-540. https://doi.org/10.1016/j.tics.2024.01.011
Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv. https://doi.org/10.48550/arXiv.1301.3781
Mon-Williams, R., Li, G., Long, R., Du, W., & Lucas, C. G. (2025). Embodied large language models enable robots to complete complex tasks in unpredictable environments. Nature Machine Intelligence, 7, 592-601. https://doi.org/10.1038/s42256-025-01005-x
Nisbett, R. E., & Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231-259. https://doi.org/10.1037/0033-295X.84.3.231
OpenAI. (2023). GPT-4 technical report. arXiv. https://doi.org/10.48550/arXiv.2303.08774
Ouyang, L., Wu, J., Jiang, X., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730-27744.
Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative agents: Interactive simulacra of human behavior. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, Article 2, 1-22. https://doi.org/10.1145/3586183.3606763
Schaeffer, R., Miranda, B., & Koyejo, S. (2023). Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems, 36.
Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3(3), 417-424.
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36.
Tanenhaus, M. K., Spivey-Knowlton, M. J., Eberhard, K. M., & Sedivy, J. C. (1995). Integration of visual and linguistic information in spoken language comprehension. Science, 268(5217), 1632-1634.
Tomasello, M. (2003). Constructing a language: A usage-based theory of language acquisition. Harvard University Press.
Tomasello, M., Carpenter, M., Call, J., Behne, T., & Moll, H. (2005). Understanding and sharing intentions: The origins of cultural cognition. Behavioral and Brain Sciences, 28(5), 675-691.
Turing, A. M. (1950). Computing machinery and intelligence. Mind, 59(236), 433-460.
Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124-1131. https://doi.org/10.1126/science.185.4157.1124
Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
Wei, J., Tay, Y., Bommasani, R., et al. (2022). Emergent abilities of large language models. Transactions on Machine Learning Research.
Weizenbaum, J. (1976). Computer power and human reason: From judgment to calculation. W. H. Freeman.
Yao, S., Zhao, J., Yu, D., et al. (2023). ReAct: Synergizing reasoning and acting in language models. International Conference on Learning Representations.
Yin, Y., Jia, N., & Wakslak, C. J. (2024). AI can help people feel heard, but an AI label diminishes this impact. Proceedings of the National Academy of Sciences, 121(14), e2319112121. https://doi.org/10.1073/pnas.2319112121