References
ISO 24624:2016 Language resource management — Transcription of spoken language
This page is work in progress.
Papers
These are publications addressing in some way or other the standard, its use, or more general aspects of standardisation of spoken language transcription.
Arkhangelskiy, Timofey, Ferger, Aanne & Hedeland, Hanna (2019): Uralic multimedia corpora: ISO/TEI corpus data in the project INEL. 5th International Workshop on Computational Linguistics for Uralic Languages, 115–124, Tartu, Estonia. Association for Computational Linguistics
Bański, Piotr & Hedeland, Hanna (2022): Standards in CLARIN. In: CLARIN: The Infrastructure for Language Resources. Darja Fišer and Andreas Witt (eds). De Gruyter (open access).
Boas, Hans, Schmidt, Thomas & Blevins, Margaret (2026): TGDA 2.0 - A corpus platform for the Texas German Dialect Archive. Language Resources & Evaluation. Springer.
Ecker, Jennifer (2026): Standardising language data through the conversion pipeline TEIWorLD In: Online-Only Publikationen Des Leibniz-Instituts für Deutsche Sprache, 15.
Frick, Elena & Schmidt, Thomas (2025): Querying spoken language data. In: Bański, P./Heid, U./Herzberg, L. (eds.): Standards for language data and infrastructures. Series: Digital Linguistics. Boston: de Gruyter.
Fisseni, Bernhard & Schmidt, Thomas (2020): CLARIN Web Services for TEI-annotated Transcripts of Spoken Language. In: Selected Papers from the CLARIN Annual Conference 2019. Linköping University Electronic Press: Linköping, pp. 12–22.
Hedeland, Hanna & Schmidt, Thomas (2022): The TEI-based ISO Standard ‘Transcription of spoken language’ as an Exchange Format within CLARIN and beyond. In: Selected Papers from the CLARIN Annual Conference 2021. Linköping University Electronic Press: Linköping, pp. 34–45.
Maarten, Janssen (2016): TEITOK: Text-Faithful Annotated Corpora. In: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 4037–4043, Portorož, Slovenia. European Language Resources Association (ELRA).
Parisse, Christophe, Etienne, Carole & Liégeois, Loïc (2020): TEICORPO: A Conversion Tool for Spoken Language Transcription with a Pivot File in TEI In: Journal of the Text Encoding Initiative 13. 2020.
Schmidt, Thomas (2005): Time-based data models and the Text Encoding Initiative’s guidelines for transcription of speech. Working Papers in Multilingualism, Series B (62). Hamburg.
Schmidt, Thomas, Duncan, Susan, Ehmer, Oliver, Hoyt, Jeffrey, Kipp, Michael, Loehr, Dan, Magnusson, Magnus, Rose, Travis & Sloetjes, Han (2009). An exchange format for multimodal annotations. In Kipp, M., Martin, J., Paggio, P. & Heylen, D. (eds.): Multimodal corpora: from models of natural interaction to systems and applications. Berlin/Heidelberg: Springer, 2009, 207-221.
Schmidt, Thomas (2011): A TEI-based approach to standardising spoken language transcription. In: Journal of the Text Encoding Initiative 1. 2011.
Schmidt, Thomas, Hedeland, Hanna, & Frick, Elena (2021): Ein Standard in der Praxis: ISO 24624:2016. Transcription of spoken language. FORGE 2021: Forschungsdaten in den Geisteswissenschaften - Mapping the Landscape - Geisteswissenschaftliches Forschungsdatenmanagement zwischen lokalen und globalen, generischen und spezifischen Lösungen (FORGE2021), Cologne.
Schmidt, Thomas (2025) Représenter et accéder à la parole dans les corpus oraux : diversification et adaptation des méthodes et technologies. In: Kanaan-Caillol, L./Dugua, C./Abouda, L./Gerstenberg, A. (Hrsg.): Représenter la parole. Berlin/Boston: de Gruyter.
Schmidt, Thomas, Ferger, Anne & Frick, Elena (2026): Putting things on top of other things: The ZuMult platform for multimodal corpora and its ecosystem. To appear in: Grisot, C. et al.: Selected Papers of the CLARIN Annual Conference 2025.
Schmidt, Thomas (2026): Chapter 3. Typology of Non-Textual Language Data. Accepted, to appear in: Lenz, A. / Witt, A. / Kamocki, P. (eds.): Routledge Handbook of Digital Linguistics.
Werthmann, Antonina (2025): From spoken language data to TEI-based ISO standard. In: Bański, Piotr, Heid, Ulrich, Herzberg, Laura (eds.): Harmonizing language data. Standards for linguistic resources. (= Digital Linguistics 4). Berlin / Boston: de Gruyter. S. 145-168.
Corpora
These are spoken language corpora making transcripts available in the standard.
- The EXMARaLDA demo corpus will have short examples of ISO/TEI transcripts in eight languages.
- The Oral-History.Digital (OHD) portal provides transcripts from its oral history archive in the ISO/TEI format. This is one of the results of the Text+oh.d cooperation project.
- Many of the corpora of the Archive for Spoken German have been transformed to an ISO/TEI version for inclusion in the archive's ZuMult instance
- For the corpus Enquêtes Sociolinguistiques à Orléans (ELSO), ISO/TEI versions of the transcripts were generated for inclusion in the project's ZuMult platform.
- The language documentation project INEL provides ISO/TEI transcriptions for the spoken data of all of its corpora.
- Version 1.1 of the Training corpus of spoken Slovenian ROG has ISO/TEI versions of all transcripts.
- Transcripts of the narrative tasks of the Equatorial Guinea Spanish corpus have been deposited on LaRS@SWISSUbase.
- The TIGR corpus of spoken Italian contains ISO/TEI transcripts of its data. The corpus will be deposited with LaRS@SWISSUbase.
- Transcripts of two corpora from the Texas German Dialect Project were converted to ISO/TEI for inclusion into the project's ZuMult platform.
- The transcripts of five collections from the DOBES archive were converted to ISO/TEI for inclusion in the repository of the Institute for the German Language .