Author
Abstract
Background. This article examines the types and specific features of the most well‑known German linguistic corpora as tools for studying the German language. The corpora are described to inform the scientific community about the opportunities they offer. The objectives of the study include: description of the electronic portals, their structure, where these German corpora are hosted; presentation of the data — text base and volume, as well as their structure. The conditions and capabilities of linguistic search in these resources are also discussed. The history of the appearance of the first corpus is mentioned, its modern definition is given. There is a review of the scientific literature about the use of corpus data in linguistics and related sciences too. Purpose. The description of the volume, structure, and features of German‑language corpora, as well as the possibilities of automating the process of extracting material taking into account the needs of researchers. Materials and methods. The material for the study are German‑language electronic resources, two of which are examined in detail: the electronic dictionary of the German language — Digitales Wörterbuch der deutschen Sprache (DWDS) and the corpora of the Institute of the German Language — IDS‑Korpora: Corpora of Written Language (LIMAS), the DeReKo project and COSMAS II, Datenbank für Gesprochenes Deutsch (DGD). The paper also references the NEGRA project, a special corpus of the University of Saarbrücken in the Federal State of Saarland, and the Deutscher Wortschatz (German Vocabulary) — corpus dictionary project of the University of Leipzig. The primary research method is structural and descriptive. Quantitative indicators were used to describe and present the data. Results. Along with the corpora descriptions, the authors present recommendations on the application areas of each German‑language resource and the automated search and processing tools available on its portal. The descriptions also mention the volume, structure, and features of the linguistic markup supported by the corpora, which determines the specifics of the extracted material. The authors emphasize the potential of using corpus tools for linguists to test their hypotheses and process empirical data.
Suggested Citation
Handle:
RePEc:cxm:russhs:17:4:2025:137-160
DOI: https://doi.org/10.12731/3033-5981-2025-17-4-540
Download full text from publisher
References listed on IDEAS
- repec:cxm:russhs:17:2:2025:81-102 is not listed on IDEAS
Full references (including those not matched with items on IDEAS)
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:cxm:russhs:17:4:2025:137-160. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Yan Maksimov (email available below). General contact details of provider: http://nkras.ru/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.