Overview
ISO/TR 21636-2:2023 – Language Coding: A Framework for Language Varieties – Part 2: Description of the Framework provides the foundational structure for the identification and description of language varieties within individual human languages. Developed by ISO Technical Committee 37, this standard serves as a key reference for stakeholders dealing with linguistic resources and technologies, ensuring a consistent approach to language metadata and the classification of varieties across modalities, time periods, regions, and social groups. It explicitly excludes artificial communication systems such as programming languages and isolated symbols or gestures that cannot be freely combined into complex linguistic expressions.
Key Topics
ISO/TR 21636-2:2023 focuses on the following major aspects:
- Linguistic Variation: Explains how and why languages differ internally, addressing the complexities involved in dialects, sociolects, and other varieties.
- Framework Dimensions: Establishes eight primary dimensions by which language varieties can be identified and described:
- Space (geographic variation, e.g., dialects)
- Time (historical epochs and periods)
- Social Group (sociolects, professional jargons)
- Medium (spoken, written, signed, multimodal, etc.)
- Situation (contextual factors affecting register and genre)
- Individual Speaker (idiolects)
- Proficiency (levels of language mastery)
- Communicative Functioning (pragmatic role in communication)
- Metadata for Language Varieties: Provides guidelines for structuring necessary metadata to facilitate interoperability, reusability, and accurate archiving of language resources.
Applications
The framework described in ISO/TR 21636-2:2023 delivers substantial value across various sectors that handle linguistic data:
- Language Resources (LRs): Supports precise cataloguing and retrieval of spoken, written, and signed language materials, critical for language archives and digital repositories.
- Language Technologies (LTs): Enhances the development, evaluation, and customization of speech recognition, machine translation, and linguistic analysis tools.
- Linguistic Research & Documentation: Offers a structured approach for researchers and archivists to describe and compare varieties within and across languages, supporting sociolinguistic, historical, and typological studies.
- Cultural Heritage and Preservation: Ensures the accurate identification of language varieties in archives, contributing to the preservation of linguistic diversity for future generations.
- eGovernment, eLearning, eHealth: Facilitates tailored communication and user experiences by accommodating multiple language varieties relevant to specific audiences or contexts.
- Translation and Interpretation: Assists professionals in selecting and working with the appropriate language variety for target audiences, enhancing the accuracy and cultural relevance of translations.
Related Standards
- ISO 639 Series: International standard for language codes, referenced for the identification of existing languages but not for internal varieties.
- ISO/IEC 11179: Metadata registry standard, providing the broader context for the metadata approach in ISO/TR 21636-2:2023.
- IETF BCP 47: Best Current Practice for language tag extension, relevant for indicating complex or imitated varieties.
- ISO 3166, ISO 6709, ISO 19112: Standards for geographical identifiers, useful when specifying regions relevant to language varieties.
This framework provides a robust, standardized method for anyone working with language data to accurately describe, classify, and exchange information about language varieties. By adopting the ISO/TR 21636-2:2023 guidelines, organizations and researchers can ensure consistent language coding, enabling better interoperability and delivering improved outcomes across all applications dependent on nuanced linguistic information.