Scottish Corpus of Texts and Speech
||This article's factual accuracy may be compromised due to out-of-date information. (November 2009)|
The Scottish Corpus of Texts & Speech (SCOTS) is an ongoing project to build a corpus of modern-day (post-1940) written and spoken texts in Scottish English and varieties of Scots. SCOTS has been available online since November 2004, and can be freely searched and browsed. By the end of the project, in mid-2007, SCOTS aims to increase the size of the text collection to 4 million words.
SCOTS contains texts in Scottish English and varieties of broad Scots, including Doric, Lallans, urban varieties such as Glaswegian and Insular Scots. SCOTS contains a geographical spread of texts as well as a demographic spread. Each text is accompanied by extensive metadata, including such information as author’s decade of birth, gender, occupation, birthplace and place of residence, and details about the text such as publication information, audience, date and genre.
Genre and mode
SCOTS is a multimedia corpus, containing written texts and spoken texts, available as orthographic transcriptions, accompanied by source audio or video files. SCOTS includes a large number of genres and text types, including prose fiction, poetry, business and personal correspondence, religious texts, parliamentary and administrative documents, emails, conversations and interviews.
Search and analysis
SCOTS can be investigated in various ways, depending on the user’s interest. The corpus can be browsed, for example by the author’s name or date of the text, and all texts can be downloaded in plain text format.
Transcriptions are synchronised with audio / video files, which are streamed and may also be downloaded.
An Advanced Search facility allows the user to build up more complex queries, choosing from all the fields available in the metadata. Geographical results are plotted on an interactive map, so regional variation may be investigated.