| تعداد نشریات | 32 |
| تعداد شمارهها | 591 |
| تعداد مقالات | 5,801 |
| تعداد مشاهده مقاله | 8,739,929 |
| تعداد دریافت فایل اصل مقاله | 6,403,742 |
The Wug Test Revisited: A Cross-Linguistic Protocol for Comparing Morphological Productivity in Humans and Large Language Models | ||
| Interdisciplinary Studies in English Language Teaching | ||
| مقالات آماده انتشار، پذیرفته شده، انتشار آنلاین از تاریخ 03 مهر 1405 | ||
| نوع مقاله: Original Article | ||
| شناسه دیجیتال (DOI): 10.22080/iselt.2026.32544.1218 | ||
| نویسندگان | ||
| Zhaleh Gholami1؛ Hamed Zarabi* 2 | ||
| 1Ph.D. Candidate of Linguistics, Islamic Azad University of Isfahan (Khorasgan) Branch, Isfahan, Iran | ||
| 2Ph.D. Candidate of TEFL, Hakim Sabzevari University, Sabzevar, Iran. | ||
| تاریخ دریافت: 05 مرداد 1405، تاریخ بازنگری: 27 شهریور 1405، تاریخ پذیرش: 29 شهریور 1405 | ||
| چکیده | ||
| The Wug Test is often treated as a compact demonstration that speakers can extend morphological patterns beyond memorized words. Its apparent simplicity, however, conceals a difficult inferential problem when the same task is administered to large language models (LLMs): a formally correct nonce-word response may arise from productive abstraction, lexical analogy, subword completion, prompt-local imitation, or contamination from training data. This article develops a falsifiable comparative protocol for distinguishing these mechanisms across adult native speakers, children, second language learners, and contemporary LLMs. English is paired with Persian, whose productive and restricted plural systems, derivational resources, script, and uneven representation in model training create a stringent cross-linguistic test. The design orthogonally manipulates phonotactic probability, lexical-neighborhood density, morphological family support, tokenizer alignment, script, prompt regime, and inflectional or derivational complexity. Responses are evaluated not only for target accuracy, but also for error topology, calibration, prompt stability, and transfer to held-out morphological families. Hierarchical models estimate group-specific sensitivities while contamination audits and repeated model sampling constrain claims about novelty. The central argument is methodological: human-like aggregate accuracy is neither necessary nor sufficient evidence of human-like productivity. A model warrants that attribution only when it reproduces the structured generalizations, dissociations, and error distributions that make nonce-word evidence theoretically informative in human speakers. If implemented, the protocol would provide a reproducible standard for evaluating morphological claims about LLMs and for designing multilingual benchmarks that distinguish robust generalization from prompt- or tokenizer-dependent success. | ||
| کلیدواژهها | ||
| Large Language Models؛ Morphological Productivity؛ Persian Morphology؛ Tokenization؛ Wug Test | ||
| مراجع | ||
|
Albright, A., & Hayes, B. (2003). Rules vs. analogy in English past tenses: A computational/experimental study. Cognition, 90(2), 119–161. https://doi.org/10.1016/S0010-0277(03)00146-X
Alegre, M., & Gordon, P. (1999). Frequency effects and the representational status of regular inflections. Journal of Memory and Language, 40(1), 41–61. https://doi.org/10.1006/jmla.1998.2607
Anh, D., Raviv, L., & Galke, L. (2024). Morphology matters: Probing the cross-linguistic morphological generalization abilities of large language models through a Wug Test. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics (pp. 177–188). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.cmcl-1.15
Aranovich, R. (2026). Charles Hockett’s “Two Models of Grammatical Description”: A neo-Bloomfieldian searching for alternatives. WORD, 72(1), 5–17. https://doi.org/10.1080/00437956.2026.2621441
Baayen, R. H. (1992). Quantitative aspects of morphological productivity. In G. E. Booij & J. van Marle (Eds.), Yearbook of morphology 1991 (pp. 109–149). Kluwer.
Baayen, R. H. (1993). On frequency, transparency and productivity. In G. E. Booij & J. van Marle (Eds.), Yearbook of morphology 1992 (pp. 181–208). Kluwer.
Baayen, R. H., & Lieber, R. (1991). Productivity and English derivation: A corpus-based study. Linguistics, 29(5), 801–843. https://doi.org/10.1515/ling.1991.29.5.801
Baayen, R. H., & Renouf, A. (1996). Chronicling the Times: Productive lexical innovations in an English newspaper. Language, 72(1), 69–96. https://doi.org/10.2307/416794
Barr, D. J., Levy, R., Scheepers, C., & Tily, H. J. (2013). Random effects structure for confirmatory hypothesis testing: Keep it maximal. Journal of Memory and Language, 68(3), 255–278. https://doi.org/10.1016/j.jml.2012.11.001
Bates, D., Mächler, M., Bolker, B., & Walker, S. (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1–48. https://doi.org/10.18637/jss.v067.i01
Bauer, L. (2001). Morphological productivity. Cambridge University Press.
Berko, J. (1958). The child’s learning of English morphology. WORD, 14(2–3), 150–177. https://doi.org/10.1080/00437956.1958.11659661
Blevins, J. P. (2006). Word-based morphology. Journal of Linguistics, 42(3), 531–573. https://doi.org/10.1017/S0022226706004191
Blevins, J. P. (2016). Word and paradigm morphology. Oxford University Press.
Booij, G. (2010). Construction morphology. Oxford University Press.
Bostrom, K., & Durrett, G. (2020). Byte pair encoding is suboptimal for language model pretraining. In Findings of the Association for Computational Linguistics: EMNLP 2020 (pp. 4617–4624). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.findings-emnlp.414
Bozorgian, H., & Rahimi, H. (2025). Peer e-feedback and ChatGPT-4o in EFL writing: A cognitive-interpersonal comparison based on EFL students. Journal of Language and Education, 11(4), 51–65. https://doi.org/10.17323/jle.2025.27195
Brysbaert, M., & Stevens, M. (2018). Power analysis and effect size in mixed effects models: A tutorial. Journal of Cognition, 1(1), Article 9. https://doi.org/10.5334/joc.10
Bybee, J. L. (1985). Morphology: A study of the relation between meaning and form. John Benjamins.
Bybee, J. (1995). Regular morphology and the lexicon. Language and Cognitive Processes, 10(5), 425–455. https://doi.org/10.1080/01690969508407111
Bybee, J. L., & Moder, C. L. (1983). Morphological classes as natural categories. Language, 59(2), 251–270. https://doi.org/10.2307/413574
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., & Zhang, C. (2023). Quantifying memorization across neural language models. In International Conference on Learning Representations.
Ciaccio, C., Miaschi, A., & Dell’Orletta, F. (2025). Evaluating lexical proficiency in neural language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics: Volume 1, Long Papers (pp. 1267–1286). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.64
Clahsen, H. (1999). Lexical entries and rules of language: A multidisciplinary study of German inflection. Behavioral and Brain Sciences, 22(6), 991–1013. https://doi.org/10.1017/S0140525X99002228
Cotterell, R., Kirov, C., Sylak-Glassman, J., Yarowsky, D., Eisner, J., & Hulden, M. (2016). The SIGMORPHON 2016 shared task—Morphological reinflection. In Proceedings of the 14th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology (pp. 10–22). Association for Computational Linguistics. https://doi.org/10.18653/v1/W16-2002
Daland, R., Hayes, B., White, J., Garellek, M., Davis, A., & Norrmann, I. (2011). Explaining sonority projection effects. Phonology, 28(2), 197–234. https://doi.org/10.1017/S0952675711000145
Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M., & Gardner, M. (2021). Documenting large webtext corpora: A case study on the Colossal Clean Crawled Corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 1286–1305). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.98
Ernestus, M., & Baayen, R. H. (2003). Predicting the unpredictable: Interpreting neutralized segments in Dutch. Language, 79(1), 5–38. https://doi.org/10.1353/lan.2003.0076
Ghomeshi, J. (1997). Non-projecting nouns and the Ezafe construction in Persian. Natural Language & Linguistic Theory, 15(4), 729–788. https://doi.org/10.1023/A:1005886709040
Ghomeshi, J. (2003). Plural marking, indefiniteness, and the noun phrase. Studia Linguistica, 57(2), 47–74. https://doi.org/10.1111/1467-9582.00099
Gouskova, M., & Becker, M. (2013). Nonce words show that Russian yer alternations are governed by the grammar. Natural Language & Linguistic Theory, 31(3), 735–765. https://doi.org/10.1007/s11049-013-9197-5
Hay, J., & Baayen, R. H. (2005). Shifting paradigms: Gradient structure in morphology. Trends in Cognitive Sciences, 9(7), 342–348. https://doi.org/10.1016/j.tics.2005.04.002
Hayes, B., & Londe, Z. C. (2006). Stochastic phonological knowledge: The case of Hungarian vowel harmony. Phonology, 23(1), 59–104. https://doi.org/10.1017/S0952675706000765
Hayes, B., & Wilson, C. (2008). A maximum entropy model of phonotactics and phonotactic learning. Linguistic Inquiry, 39(3), 379–440. https://doi.org/10.1162/ling.2008.39.3.379
Hockett, C. F. (1954). Two models of grammatical description. WORD, 10(2–3), 210–234. https://doi.org/10.1080/00437956.1954.11659524
Ismayilzada, M., Circi, D., Sälevä, J., Sirin, H., Köksal, A., Dhingra, B., Bosselut, A., Ataman, D., & van der Plas, L. (2025). Evaluating morphological compositional generalization in large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies: Volume 1, Long Papers (pp. 1270–1305). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.naacl-long.59
Jaeger, T. F. (2008). Categorical data analysis: Away from ANOVAs (transformation or not) and towards logit mixed models. Journal of Memory and Language, 59(4), 434–446. https://doi.org/10.1016/j.jml.2007.11.007
Kandpal, N., Wallace, E., & Raffel, C. (2022). Deduplicating training data mitigates privacy risks in language models. In Proceedings of the 39th International Conference on Machine Learning (pp. 10697–10707). PMLR.
Kann, K., & Schütze, H. (2016). MED: The LMU system for the SIGMORPHON 2016 shared task on morphological reinflection. In Proceedings of the 14th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology (pp. 62–70). Association for Computational Linguistics. https://doi.org/10.18653/v1/W16-2010
Kann, K., Cotterell, R., & Schütze, H. (2017). Neural multi-source morphological reinflection. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (Vol. 1, pp. 514–524). Association for Computational Linguistics. https://doi.org/10.18653/v1/E17-1049
Karimi-Doostan, G. (2005). Light verbs and structural case. Lingua, 115(12), 1737–1756. https://doi.org/10.1016/j.lingua.2004.08.002
Keuleers, E., & Brysbaert, M. (2010). Wuggy: A multilingual pseudoword generator. Behavior Research Methods, 42(3), 627–633. https://doi.org/10.3758/BRM.42.3.627
Kirov, C., & Cotterell, R. (2018). Recurrent neural networks in linguistic theory: Revisiting Pinker and Prince (1988) and the past tense debate. Transactions of the Association for Computational Linguistics, 6, 651–665. https://doi.org/10.1162/tacl_a_00247
Kuczaj, S. A., II. (1977). The acquisition of regular and irregular past tense forms. Journal of Verbal Learning and Verbal Behavior, 16(5), 589–600. https://doi.org/10.1016/S0022-5371(77)80021-2
Kudo, T. (2018). Subword regularization: Improving neural network translation models with multiple subword candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 66–75). Association for Computational Linguistics. https://doi.org/10.18653/v1/P18-1007
Kudo, T., & Richardson, J. (2018). SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (pp. 66–71). Association for Computational Linguistics. https://doi.org/10.18653/v1/D18-2012
Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1), Article 33267. https://doi.org/10.1525/collabra.33267
Lazard, G. (1992). A grammar of contemporary Persian. Mazda Publishers.
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., & Carlini, N. (2022). Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 8424–8445). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.acl-long.577
Lerner, P., & Yvon, F. (2025). Unlike “likely,” “unlike” is unlikely: BPE-based segmentation hurts morphological derivations in LLMs. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 5181–5190). Association for Computational Linguistics.
Liu, L., & Hulden, M. (2022). Can a transformer pass the Wug Test? Tuning copying bias in neural morphological inflection models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Vol. 2, pp. 739–749). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.acl-short.82
Mahootian, S. (1997). Persian. Routledge.
Makarov, P., & Clematide, S. (2018). UZH at CoNLL–SIGMORPHON 2018 shared task on universal morphological reinflection. In Proceedings of the CoNLL–SIGMORPHON 2018 Shared Task (pp. 69–75). Association for Computational Linguistics. https://doi.org/10.18653/v1/K18-3008
Marchman, V. A., & Bates, E. (1994). Continuity in lexical and morphological development: A test of the critical mass hypothesis. Journal of Child Language, 21(2), 339–366. https://doi.org/10.1017/S0305000900009302
Marcus, G. F., Pinker, S., Ullman, M., Hollander, M., Rosen, T. J., & Xu, F. (1992). Overregularization in language acquisition. Monographs of the Society for Research in Child Development, 57(4), i–178. https://doi.org/10.2307/1166115
McClelland, J. L., & Patterson, K. (2002). Rules or connections in past-tense inflections: What does the evidence rule out? Trends in Cognitive Sciences, 6(11), 465–472. https://doi.org/10.1016/S1364-6613(02)01993-9
McCoy, R. T., Smolensky, P., Linzen, T., Gao, J., & Celikyilmaz, A. (2023). How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN. Transactions of the Association for Computational Linguistics, 11, 652–670. https://doi.org/10.1162/tacl_a_00567
Mehranirad, M., Arabpour, M., & Emami Neyshaburi, E. (2025). Developing a Persian version of the Wug Test to assess children’s morphological knowledge. Journal of Child Language Acquisition and Development, 1198–1214. https://doi.org/10.5281/zenodo.17008773
Munafò, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers, C. D., Percie du Sert, N., Simonsohn, U., Wagenmakers, E.-J., Ware, J. J., & Ioannidis, J. P. A. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1, Article 0021. https://doi.org/10.1038/s41562-016-0021
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606. https://doi.org/10.1073/pnas.1708274114
Pantelidou, N., Leivada, E., Montero, R., & Morosi, P. (2026). Community size rather than grammatical complexity better predicts large language model accuracy in a novel Wug Test. PLOS ONE, 21(3), e0343164. https://doi.org/10.1371/journal.pone.0343164
Pinker, S., & Prince, A. (1988). On language and connectionism: Analysis of a parallel distributed processing model of language acquisition. Cognition, 28(1–2), 73–193. https://doi.org/10.1016/0010-0277(88)90032-7
Plag, I. (1999). Morphological productivity: Structural constraints in English derivation. Mouton de Gruyter.
Prasada, S., & Pinker, S. (1993). Generalisation of regular and irregular morphological patterns. Language and Cognitive Processes, 8(1), 1–56. https://doi.org/10.1080/01690969308406948
Razeghi, Y., Logan, R. L., IV, Gardner, M., & Singh, S. (2022). Impact of pretraining term frequencies on few-shot numerical reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2022 (pp. 840–854). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.findings-emnlp.59
Rumelhart, D. E., & McClelland, J. L. (1986). On learning the past tenses of English verbs. In J. L. McClelland, D. E. Rumelhart, & the PDP Research Group (Eds.), Parallel distributed processing: Explorations in the microstructure of cognition (Vol. 2, pp. 216–271). MIT Press.
Rust, P., Pfeiffer, J., Vulić, I., Ruder, S., & Gurevych, I. (2021). How good is your tokenizer? On the monolingual performance of multilingual language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Vol. 1, pp. 3118–3135). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.acl-long.243
Sagot, B., & Walther, G. (2010). A morphological lexicon for the Persian language. In Proceedings of the Seventh International Conference on Language Resources and Evaluation (pp. 3707–3714). European Language Resources Association.
Samvelian, P. (2012). The syntax of the Persian noun phrase. John Benjamins.
Sennrich, R., Haddow, B., & Birch, A. (2016). Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 1715–1725). Association for Computational Linguistics. https://doi.org/10.18653/v1/P16-1162
Silfverberg, M., Wiemerslage, A., Liu, L., & Mao, L. J. (2017). Data augmentation for morphological reinflection. In Proceedings of the CoNLL SIGMORPHON 2017 Shared Task: Universal Morphological Reinflection (pp. 90–99). Association for Computational Linguistics. https://doi.org/10.18653/v1/K17-2012
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. https://doi.org/10.1177/0956797611417632
Truong, T. H., Otmakhova, Y., Verspoor, K., Cohn, T., & Baldwin, T. (2024). Revisiting subword tokenization: A case study on affixal negation in large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 5082–5095). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.naacl-long.284
Ullman, M. T. (1999). Acceptability ratings of regular and irregular past-tense forms: Evidence for a dual-system model of language from word frequency and phonological neighborhood effects. Language and Cognitive Processes, 14(1), 47–67. https://doi.org/10.1080/016909699386383
Vylomova, E., White, J., Salesky, E., Mielke, S. J., Wu, S., Ponti, E. M., Maudslay, R. H., Zmigrod, R., Valvoda, J., Toldova, S., Tyers, F., Klyachko, E., Yegorov, I., Krizhanovsky, N., Czarnowska, P., Nikkarinen, I., Krizhanovsky, A., Pimentel, T., Hennigen, L. T., … Hulden, M. (2020). SIGMORPHON 2020 shared task 0: Typologically diverse morphological inflection. In Proceedings of the 17th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology (pp. 1–39). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.sigmorphon-1.1
Weissweiler, L., Hofmann, V., Kantharuban, A., Cai, A., Dutt, R., Hengle, A., Kabra, A., Kulkarni, A., Vijayakumar, A., Yu, H., Schütze, H., Oflazer, K., & Mortensen, D. R. (2023). Counting the bugs in ChatGPT’s wugs: A multilingual investigation into the morphological capabilities of a large language model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 6508–6524). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.401
Westfall, J., Kenny, D. A., & Judd, C. M. (2014). Statistical power and optimal design in experiments in which samples of participants respond to samples of stimuli. Journal of Experimental Psychology: General, 143(5), 2020–2045. https://doi.org/10.1037/xge0000014
Wilson, C., & Li, J. S. Y. (2021). Were we there already? Applying minimal generalization to the SIGMORPHON-UniMorph shared task on cognitively plausible morphological inflection. In Proceedings of the 18th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology (pp. 283–291). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.sigmorphon-1.31
Wu, S., Cotterell, R., & Hulden, M. (2021). Applying the transformer to character-level transduction. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics (pp. 1901–1907). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.eacl-main.163
Xiong, Y., & Wu, S. (2025). Do large language models learn like humans? Interleaved and spaced practice in morphological learning. Acta Psychologica, 260, 105518. https://doi.org/10.1016/j.actpsy.2025.105518
Yang, C. (2016). The price of linguistic productivity: How children learn to break the rules of language. MIT Press.
| ||
|
آمار تعداد مشاهده مقاله: 1 |
||