SIST ISO 24624:2018 PDF
Language resource management -- Transcription of spoken language
Language resource management -- Transcription of spoken language
- Статус документа:
- Действующий
- Формат:
- Электронный (PDF)
- Количество страниц:
- 39
- Дата публикации:
- 5 сентября 2018 г.
- ICS:
- 01.140.10
- Технический комитет:
- IDT - Information, documentation, language and terminology
ISO 24624:2016 specifies rules for representing transcriptions of audio- and video-recorded spoken interactions in XML documents based on the guidelines of the TEI. As a secondary objective, the document aims to relate transcribed data with standards for annotated corpora. It is applicable to transcription data for studies in sociolinguistics, conversation analysis, dialectology, corpus linguistics, corpus lexicography, language technology, qualitative social studies and other transcription data of recorded spoken language. It is not applicable to other forms of transcription, most importantly transcriptions of hand-written manuscripts. Annex A gives a fully encoded example and Annex B provides an element index and an attribute index.
Abstract
Overview
ISO 24624:2016, "Language resource management - Transcription of spoken language", defines a TEI- and XML-based framework for representing orthography-based transcriptions of audio- and video-recorded spoken interactions. Designed to improve interoperability, the standard specifies how to encode transcription macrostructure, microstructure, and basic metadata so recordings and annotations can be exchanged, searched and processed across tools and corpora. It applies to research and development in linguistics, corpus construction, language technology and qualitative social studies. The standard does not cover non-audio transcriptions (e.g., handwritten manuscripts).
Key topics and technical requirements
- TEI / XML encoding: Uses TEI guidelines as the reference framework to produce XML transcription documents that are interoperable with TEI-aware tools.
- Metadata (TEI header): Prescribes description of the electronic file, recording and distribution info, participant and setting descriptions, and source encoding to ensure reproducible, discoverable transcriptions.
- Macrostructure elements: Defines timeline markers and utterance grouping using elements such as ****, ****, ****, **** to map speech segments to time and context.
- Microstructure elements: Covers tokenization and fine-grained annotation with elements like **** (tokens), ****, ****, ****, **** (punctuation), ****, ****, ****, and ****.
- Dependent and stand-off annotations: Supports dependent annotations (annotations that reference other annotations) and grouping of stand-off annotations for multilayer corpora.
- Compatibility goals: Developed to be compatible with common transcription tools (ANVIL, CLAN, ELAN, EXMARaLDA, FOLKER, Transcriber) and transcription conventions; intended both as a target format for legacy conversion and a robust format for future processing.
- Examples and indexes: Annex A provides a fully encoded example; Annex B lists element and attribute indexes to aid implementation.
Applications and users
ISO 24624:2016 is practical for:
- Corpus linguists, sociolinguists, conversation analysts and dialectologists building annotated speech corpora.
- Language technologists and NLP practitioners who need standardized, time-aligned transcriptions for speech recognition, training data, or evaluation.
- Qualitative social researchers using recorded interactions requiring systematic transcription and metadata.
- Archivists and lexicographers preparing speech data for long-term reuse and cross-tool interoperability.
The standard helps unify formats, reduce conversion overhead, and improve data longevity and machine-readability.
Related standards and resources
- TEI Guidelines (reference framework used by ISO 24624)
- ISO 24611 (MAF) - token representation compatibility
- CMDI / IMDI / ISOCAT - related metadata initiatives
- W3C specifications (e.g., SSML, EMMA) - complementary; ISO 24624 does not address speech synthesis or semantic multimodal interpretation
Keywords: ISO 24624:2016, transcription of spoken language, TEI, XML transcription, language resource management, spoken language corpora, speech annotation, linguistic corpus interoperability.
Технические детали
- SKU
- SIST ISO 24624:2018
Похожие стандарты
Стандарты, упомянутые в описании
ISO 24624:2016
ДействующийLanguage resource management — Transcription of spoken language
Overview ISO 24624:2016, "Language resource management - Transcription of spoken language", defines a TEI- and XML-based framework for representing orthography-based transcriptions of audio- and vide…
BS ISO 24624:2016
ДействующийLanguage resource management. Transcription of spoken language.
SIST ISO 24611:2013
ДействующийLanguage resource management -- Morpho-syntactic annotation framework (MAF)
Overview SIST ISO 24611:2013 (ISO 24611:2012) - Morpho-syntactic annotation framework (MAF) defines a standardized framework for morpho-syntactic annotation of word-forms in texts. It provides a meta…