Text Linguistics and the Use of Linguistic Data in Modern Technologies: Prospects for Development
DOI:
https://doi.org/10.57125/FS.2023.06.20.02Keywords:
linguistic text analysis, corpus linguistics, lexical models, syntactic models, unstressed pronounAbstract
This study explores text linguistics as a branch of linguistics that analyses the structure, meaning, and other aspects of texts. The paper aims to analyse the use of linguistic data in modern technologies, in particular, to develop applications that can understand and generate linguistic content. To achieve this goal, the research uses the methodology of qualitative analysis and is based on the work of textual theorists. The study is based on a qualitative analysis of examples collected from the BNC corpus. The instructions for obtaining data are to use textual forms located at specific positions in the co-text. The use of qualitative analysis for data acquisition is useful for studying phonology, morphology, morphosyntax, and general linguistic facts. The results confirm the promising possibilities of text linguistics and the use of language data in technology, in particular the growth in the volume and availability of language data and the application of new methods of text analysis. The conclusions point to the connection of text linguistics with other fields and the expansion of the capabilities of semantic search engines and automatic text processing systems. In addition, the article emphasises the importance of digital resources and the study of phrases as effective tools for linguistic analysis. The analysis of three illustrative examples taken from the British linguistic corpus is carried out with the aim of studying textual and discourse phenomena in the use of language data. A certain inconsistency between the factors has been revealed, which leads to a blurring of the linguistic facts that the authors of the article tried to investigate. It should be noted that the selected examples, which were created to support the theory, are artificial, which makes it difficult to analyse the context. This makes it difficult for native English speakers to evaluate these examples. This is especially true for the first example, where the syntactic constructions do not match the context. This deviation was observed at the level of the sentence's information structure. As for the second example, the discomfort observed can be explained by the lack of motivation to use the referent that was introduced into the context as a reference marker that is independently associated with it in the place where the phrase appears in the text. There is also a clear polyphonic inconsistency in the origin of the point of view that characterises each segment of the sentence. This inconsistency is related to the choice of the reference marker.
References
Chomsky, N., Roberts, I., & Watumull, J. (2023). Noam Chomsky: The False Promise of ChatGPT. The New York Times, 8. https://www.nytimes.com/2023/03/08/opinion/noam-chomsky-chatgpt-ai.html
Cooper, C. R. (2023). The identification of YouTube videos that feature the linguistic features of English informal speech. Applied Corpus Linguistics, 3(3), 100068. https://doi.org/10.1016/j.acorp.2023.100068
Dai, Z. (2023). Statistics in Corpus Linguistics: A New Approach by Sean Wallis. New York/Oxon: Routledge, 2021. ISBN 9781138589384 (PB: 44.95), ISBN 9781138589377 (HB: 160.00), ISBN 9780429491696 (eBook: 44.95), xxvi+ 382 pages. Natural Language Engineering, 29(1), 177-180. https://doi.org/10.1017/s1351324921000413
Espada, J. P., Martínez, J. S., Rico, I. C., & Sánchez, L. E. V. (2023). Extracting keywords of educational texts using a novel mechanism based on linguistic approaches and evolutive graphs. Expert Systems with Applications, 213, 118842. https://doi.org/10.1016/j.eswa.2022.118842
Fox, J. M., Jackson, R. A., & Crawford, K. R. (2023). News of Noam: Unpacking Media Coverage of Chomsky. International Journal of Languages, Literature and Linguistics, 9, 318-325. https://doi.org/10.18178/ijlll.2023.9.5.425
Friginal, E., Cox, A., & Udell, R. (2023). Corpus Linguistics and Writing Instruction. In Demystifying Corpus Linguistics for English Language Teaching (pp. 79-97). Cham: Springer International Publishing. https://link.springer.com/chapter/10.1007/978-3-031-11220-1_5
Gillings, M., & Mautner, G. (2023). Concordancing for CADS: Practical challenges and theoretical implications. International Journal of Corpus Linguistics. https://doi.org/10.1075/ijcl.21168.gil
Hashimoto, B. (2023). Corpus of Founding Era American English: designing a corpus for interpreting the United States Constitution. Corpora, 18(1), 1-14. https://www.euppublishing.com/doi/abs/10.3366/cor.2023.0270
Jalilbayli, O. B. (2022). Forecasting the prospects for innovative changes in the development of future linguistic education for the XXI century: the choice of optimal strategies. Futurity Education, 2(4), 36-43. https://doi.org/10.57125/FED.2022.25.12.0.4
Laske, C. (2022). Corpus linguistics: the digital tool kit for analysing language and the law. Comparative Legal History, 10(1), 3-32. https://doi.org/10.1080/2049677X.2022.2063510
Lin, P., & Adolphs, S. (2023). Corpus linguistics. In The Routledge Handbook of Applied Linguistics (pp. 296-308). Routledge. https://www.taylorfrancis.com/chapters/edit/10.4324/9781003082644-25/corpus-linguistics-phoebe-lin-svenja-adolphs
Liu, C., Jhang, S., & Lee, S. From Gender-Biased to Gender-Specific and Gender-Inclusive Words: A Corpus-Based Study. https://doi.org/10.24303/lakdoi.2023.31.1.85
Mautner, G. (2016). Checks and balances: How corpus linguistics can contribute to CDA. Methods of critical discourse studies, 154-179. https://www.torrossa.com/en/resources/an/5018239#page=165
McCoy, R. T., Smolensky, P., Linzen, T., Gao, J., & Celikyilmaz, A. (2023). How much do language models copy from their training data? evaluating linguistic novelty in text generation using raven. Transactions of the Association for Computational Linguistics, 11, 652-670. https://doi.org/10.1162/tacl_a_00567
Meyer, C. F. (2023). English corpus linguistics: An introduction. Cambridge University Press. https://assets.cambridge.org/052180/8790/sample/0521808790ws.pdf
Pan, Y. (2022). Intensification for discursive evaluation: a corpus-pragmatic view. Text & Talk, 42(3), 391-417. https://doi.org/10.1515/text-2020-0046
Piantadosi, S. (2023). Modern language models refute Chomsky’s approach to language. Lingbuzz Preprint, lingbuzz, 7180. https://lingbuzz.net/lingbuzz/007180
Raković, M., Iqbal, S., Li, T., Fan, Y., Singh, S., Surendrannair, S., ... & Gašević, D. (2023). Harnessing the potential of trace data and linguistic analysis to predict learner performance in a multi‐text writing task. Journal of Computer Assisted Learning, 39(3), 703-718. https://doi.org/10.1111/jcal.12769
Raphael, E. B. (2023). Gendered Representations in Language: A Corpus-Based Comparative Study of Adjective-Noun Collocations for Marital Relationships. Theory and Practice in Language Studies, 13(5), 1191-1196. https://doi.org/10.17507/tpls.1305.12
Römer, U. (2004). Comparing real and ideal language learner input: The use of an EFL textbook corpus in corpus linguistics and language teaching. Corpora and language learners, 151-168. https://www.torrossa.com/en/resources/an/5001951#page=158
Sarvazyan, A. M., González, J. Á., Franco-Salvador, M., Rangel, F., Chulvi, B., & Rosso, P. (2023). Overview of autextification at iberlef 2023: Detection and attribution of machine-generated text in multiple domains. arXiv preprint arXiv:2309.11285. https://doi.org/10.48550/arXiv.2309.11285
Schleppegrell, M. J., & Oteíza, T. (2023). Systemic functional linguistics: Exploring meaning in language. In The Routledge handbook of discourse analysis (pp. 156-169). Routledge. https://www.researchgate.net/publication/348994754_The_Routledge_Handbook_of_Systemic_Functional_Linguistics
Schneider, G. (2022). Recent changes in spoken British English in verbal and nominal constructions. Broadening the Spectrum of Corpus Linguistics: New approaches to variability and change, 105, 173. https://www.torrossa.com/en/resources/an/5495945#page=180
Schneider, G., Hundt, M., & Schreier, D. (2019). Pluralized non-count nouns across Englishes: A corpus-linguistic approach to variety types. Corpus Linguistics and Linguistic Theory, 16(3), 515-546. https://doi.org/10.1515/cllt-2018-0068
Tomas, F., Dodier, O., & Demarchi, S. (2022). Computational measures of deceptive language: prospects and issues. Frontiers in Communication, 7, 792378. https://doi.org/10.3389/fcomm.2022.792378
Udomphol, P. (2023). A Comparative Corpus-Based Analysis of the Cross-Cultural Lexico-Grammatical Differences Between Master’s Level Academic Writing in New Zealand and the United States (Doctoral dissertation, Auckland University of Technology). https://openrepository.aut.ac.nz/handle/10292/15834
Vogel, F., & Hamann, H. (2023). Legal Linguistics in times of language models and text automation: JLL Call for Abstracts (Deadline 31 March 2023). International Journal of Language & Law (JLL), 12, 1-7. https://doi.org/10.14762/jll.2023.001
Yang, Y., & Liu, X. (2023). Review of Multifunctionality in English: Corpora, Language and Academic Literacy Pedagogy: Zihan Yin and Elaine Vine (Eds.), Routledge, London and New York, 2022 (Paperback), ISBN: 978-0-367-72512-9. https://link.springer.com/article/10.1007/s41701-023-00156-9
ZAHLER, S. (2023). Some Issues in Usage‐Based Methods: Contributions from Corpus Linguistics, Psycholinguistics, and Variationist Sociolinguistics. The Handbook of Usage‐Based Linguistics, 73-90. https://doi.org/10.1002/9781119839859.ch4
Zinn, J. O., & Müller, M. (2022). Understanding discourse and language of risk. Journal of Risk Research, 25(3), 271-284. https://doi.org/10.1080/13669877.2021.2020883
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2023 author

This work is licensed under a Creative Commons Attribution 4.0 International License.