Text Linguistics and the Use of Linguistic Data in Modern Technologies: Prospects for Development

Authors

DOI:

https://doi.org/10.57125/FS.2023.06.20.02

Keywords:

linguistic text analysis, corpus linguistics, lexical models, syntactic models, unstressed pronoun

Abstract

This study explores text linguistics as a branch of linguistics that analyses the structure, meaning, and other aspects of texts. The paper aims to analyse the use of linguistic data in modern technologies, in particular, to develop applications that can understand and generate linguistic content. To achieve this goal, the research uses the methodology of qualitative analysis and is based on the work of textual theorists. The study is based on a qualitative analysis of examples collected from the BNC corpus. The instructions for obtaining data are to use textual forms located at specific positions in the co-text. The use of qualitative analysis for data acquisition is useful for studying phonology, morphology, morphosyntax, and general linguistic facts. The results confirm the promising possibilities of text linguistics and the use of language data in technology, in particular the growth in the volume and availability of language data and the application of new methods of text analysis. The conclusions point to the connection of text linguistics with other fields and the expansion of the capabilities of semantic search engines and automatic text processing systems. In addition, the article emphasises the importance of digital resources and the study of phrases as effective tools for linguistic analysis. The analysis of three illustrative examples taken from the British linguistic corpus is carried out with the aim of studying textual and discourse phenomena in the use of language data. A certain inconsistency between the factors has been revealed, which leads to a blurring of the linguistic facts that the authors of the article tried to investigate. It should be noted that the selected examples, which were created to support the theory, are artificial, which makes it difficult to analyse the context. This makes it difficult for native English speakers to evaluate these examples. This is especially true for the first example, where the syntactic constructions do not match the context. This deviation was observed at the level of the sentence's information structure. As for the second example, the discomfort observed can be explained by the lack of motivation to use the referent that was introduced into the context as a reference marker that is independently associated with it in the place where the phrase appears in the text. There is also a clear polyphonic inconsistency in the origin of the point of view that characterises each segment of the sentence. This inconsistency is related to the choice of the reference marker.

References

Chomsky, N., Roberts, I., & Watumull, J. (2023). Noam Chomsky: The False Promise of ChatGPT. The New York Times, 8. https://www.nytimes.com/2023/03/08/opinion/noam-chomsky-chatgpt-ai.html

Cooper, C. R. (2023). The identification of YouTube videos that feature the linguistic features of English informal speech. Applied Corpus Linguistics, 3(3), 100068. https://doi.org/10.1016/j.acorp.2023.100068

Dai, Z. (2023). Statistics in Corpus Linguistics: A New Approach by Sean Wallis. New York/Oxon: Routledge, 2021. ISBN 9781138589384 (PB: 44.95), ISBN 9781138589377 (HB: 160.00), ISBN 9780429491696 (eBook: 44.95), xxvi+ 382 pages. Natural Language Engineering, 29(1), 177-180. https://doi.org/10.1017/s1351324921000413

Espada, J. P., Martínez, J. S., Rico, I. C., & Sánchez, L. E. V. (2023). Extracting keywords of educational texts using a novel mechanism based on linguistic approaches and evolutive graphs. Expert Systems with Applications, 213, 118842. https://doi.org/10.1016/j.eswa.2022.118842

Fox, J. M., Jackson, R. A., & Crawford, K. R. (2023). News of Noam: Unpacking Media Coverage of Chomsky. International Journal of Languages, Literature and Linguistics, 9, 318-325. https://doi.org/10.18178/ijlll.2023.9.5.425

Friginal, E., Cox, A., & Udell, R. (2023). Corpus Linguistics and Writing Instruction. In Demystifying Corpus Linguistics for English Language Teaching (pp. 79-97). Cham: Springer International Publishing. https://link.springer.com/chapter/10.1007/978-3-031-11220-1_5

Gillings, M., & Mautner, G. (2023). Concordancing for CADS: Practical challenges and theoretical implications. International Journal of Corpus Linguistics. https://doi.org/10.1075/ijcl.21168.gil

Hashimoto, B. (2023). Corpus of Founding Era American English: designing a corpus for interpreting the United States Constitution. Corpora, 18(1), 1-14. https://www.euppublishing.com/doi/abs/10.3366/cor.2023.0270

Jalilbayli, O. B. (2022). Forecasting the prospects for innovative changes in the development of future linguistic education for the XXI century: the choice of optimal strategies. Futurity Education, 2(4), 36-43. https://doi.org/10.57125/FED.2022.25.12.0.4

Laske, C. (2022). Corpus linguistics: the digital tool kit for analysing language and the law. Comparative Legal History, 10(1), 3-32. https://doi.org/10.1080/2049677X.2022.2063510

Lin, P., & Adolphs, S. (2023). Corpus linguistics. In The Routledge Handbook of Applied Linguistics (pp. 296-308). Routledge. https://www.taylorfrancis.com/chapters/edit/10.4324/9781003082644-25/corpus-linguistics-phoebe-lin-svenja-adolphs

Liu, C., Jhang, S., & Lee, S. From Gender-Biased to Gender-Specific and Gender-Inclusive Words: A Corpus-Based Study. https://doi.org/10.24303/lakdoi.2023.31.1.85

Mautner, G. (2016). Checks and balances: How corpus linguistics can contribute to CDA. Methods of critical discourse studies, 154-179. https://www.torrossa.com/en/resources/an/5018239#page=165

McCoy, R. T., Smolensky, P., Linzen, T., Gao, J., & Celikyilmaz, A. (2023). How much do language models copy from their training data? evaluating linguistic novelty in text generation using raven. Transactions of the Association for Computational Linguistics, 11, 652-670. https://doi.org/10.1162/tacl_a_00567

Meyer, C. F. (2023). English corpus linguistics: An introduction. Cambridge University Press. https://assets.cambridge.org/052180/8790/sample/0521808790ws.pdf

Pan, Y. (2022). Intensification for discursive evaluation: a corpus-pragmatic view. Text & Talk, 42(3), 391-417. https://doi.org/10.1515/text-2020-0046

Piantadosi, S. (2023). Modern language models refute Chomsky’s approach to language. Lingbuzz Preprint, lingbuzz, 7180. https://lingbuzz.net/lingbuzz/007180

Raković, M., Iqbal, S., Li, T., Fan, Y., Singh, S., Surendrannair, S., ... & Gašević, D. (2023). Harnessing the potential of trace data and linguistic analysis to predict learner performance in a multi‐text writing task. Journal of Computer Assisted Learning, 39(3), 703-718. https://doi.org/10.1111/jcal.12769

Raphael, E. B. (2023). Gendered Representations in Language: A Corpus-Based Comparative Study of Adjective-Noun Collocations for Marital Relationships. Theory and Practice in Language Studies, 13(5), 1191-1196. https://doi.org/10.17507/tpls.1305.12

Römer, U. (2004). Comparing real and ideal language learner input: The use of an EFL textbook corpus in corpus linguistics and language teaching. Corpora and language learners, 151-168. https://www.torrossa.com/en/resources/an/5001951#page=158

Sarvazyan, A. M., González, J. Á., Franco-Salvador, M., Rangel, F., Chulvi, B., & Rosso, P. (2023). Overview of autextification at iberlef 2023: Detection and attribution of machine-generated text in multiple domains. arXiv preprint arXiv:2309.11285. https://doi.org/10.48550/arXiv.2309.11285

Schleppegrell, M. J., & Oteíza, T. (2023). Systemic functional linguistics: Exploring meaning in language. In The Routledge handbook of discourse analysis (pp. 156-169). Routledge. https://www.researchgate.net/publication/348994754_The_Routledge_Handbook_of_Systemic_Functional_Linguistics

Schneider, G. (2022). Recent changes in spoken British English in verbal and nominal constructions. Broadening the Spectrum of Corpus Linguistics: New approaches to variability and change, 105, 173. https://www.torrossa.com/en/resources/an/5495945#page=180

Schneider, G., Hundt, M., & Schreier, D. (2019). Pluralized non-count nouns across Englishes: A corpus-linguistic approach to variety types. Corpus Linguistics and Linguistic Theory, 16(3), 515-546. https://doi.org/10.1515/cllt-2018-0068

Tomas, F., Dodier, O., & Demarchi, S. (2022). Computational measures of deceptive language: prospects and issues. Frontiers in Communication, 7, 792378. https://doi.org/10.3389/fcomm.2022.792378

Udomphol, P. (2023). A Comparative Corpus-Based Analysis of the Cross-Cultural Lexico-Grammatical Differences Between Master’s Level Academic Writing in New Zealand and the United States (Doctoral dissertation, Auckland University of Technology). https://openrepository.aut.ac.nz/handle/10292/15834

Vogel, F., & Hamann, H. (2023). Legal Linguistics in times of language models and text automation: JLL Call for Abstracts (Deadline 31 March 2023). International Journal of Language & Law (JLL), 12, 1-7. https://doi.org/10.14762/jll.2023.001

Yang, Y., & Liu, X. (2023). Review of Multifunctionality in English: Corpora, Language and Academic Literacy Pedagogy: Zihan Yin and Elaine Vine (Eds.), Routledge, London and New York, 2022 (Paperback), ISBN: 978-0-367-72512-9. https://link.springer.com/article/10.1007/s41701-023-00156-9

ZAHLER, S. (2023). Some Issues in Usage‐Based Methods: Contributions from Corpus Linguistics, Psycholinguistics, and Variationist Sociolinguistics. The Handbook of Usage‐Based Linguistics, 73-90. https://doi.org/10.1002/9781119839859.ch4

Zinn, J. O., & Müller, M. (2022). Understanding discourse and language of risk. Journal of Risk Research, 25(3), 271-284. https://doi.org/10.1080/13669877.2021.2020883

Downloads

Published

2023-03-20

How to Cite

Aliyeva, G. B. (2023). Text Linguistics and the Use of Linguistic Data in Modern Technologies: Prospects for Development. Futurity of Social Sciences, 1(2), 18–29. https://doi.org/10.57125/FS.2023.06.20.02