ISO 12199:2022
Alphabetical ordering of multilingual terminological and lexicographical data represented in the Latin alphabet
Alphabetical ordering of multilingual terminological and lexicographical data represented in the Latin alphabet
- Статус документа:
- Отменён
- Формат:
- Электронный (PDF)
- Количество страниц:
- 52
- Дата публикации:
- 14 июня 2022 г.
- Издание:
- ISO IS 12199 edition 2 version 1
- ICS:
- 01.020
This document specifies the sequence of characters to be used in the alphabetical ordering of multilingual terminological and lexicographical data (terms, term elements, or words) represented in the Latin alphabet. Character sets of languages represented in the Latin alphabet are taken into account insofar as terminological or lexicographical data have been recorded. Character sets used in internationally standardized transliteration into Latin script are also taken into account. The sequence of alphabetical characters given is intended for multilingual purposes only and is not intended to affect the alphabetical order of any specific language. The main part of this document specifies letter-by-letter ordering of character strings. Annex A treats word-by-word ordering, which is a widely used alternative to this system. Annex B gives two additional rules that can be useful for lexicographical and terminological ordering. Annex C gives ordering rules for chemical names. Annex D lists the character repertoire of the Latin alphabet. Annex E lists languages using the Latin alphabet. Annex F gives alphabetical sequences derived from the sequence specified in this document for a number of languages that use the Latin alphabet. Annex G gives a formal description of the rules laid down in the main part of this document conforming with ISO/IEC 14651.
Abstract
Overview
ISO 12199:2022 specifies a practical, language-neutral method for alphabetical ordering (collation) of multilingual terminological and lexicographical data represented in the Latin alphabet. It defines the sequence of characters and multi-level ordering rules to ensure consistent sorting of terms, term elements and words across systems, databases and printed lists. The standard is intended for multilingual environments and does not replace language‑specific collation rules.
Key topics and technical requirements
- Character sequence and basic order: Defines ordering of digits (0–9) and basic Latin letters (a–z, with case equivalence). The sequence is designed to minimize conflicts between languages in multilingual resources.
- Multi-level ordering: Specifies first to fourth ordering levels:
- First level: primary letter-by-letter order (letters and digits).
- Second level: diacritical marks and special Latin letters treated relative to base letters.
- Third level: capitalization differences.
- Fourth level: special characters and punctuation.
- Equivalence mappings: Special Latin letters and letters with diacritics are mapped to corresponding basic Latin letters for primary ordering (see Table 1 in the standard).
- Word‑by‑word alternative: Annex A provides a normative word‑by‑word ordering method commonly used instead of pure letter‑by‑letter sorting.
- Additional rules and domain-specific handling:
- Annex B: extra lexicographical and terminological rules.
- Annex C: ordering rules for chemical names.
- Annex D: Latin alphabet character repertoire.
- Annex E/F: lists of languages using Latin script and derived alphabetical sequences for selected languages.
- Annex G: formal rule description conforming with ISO/IEC 14651.
- Preparatory procedures: The standard notes pre-sorting steps (e.g., case folding, numeral padding, handling polygraphs) but does not mandate extraction or normalization methods.
- Language sensitivities: Special handling examples (e.g., Turkish dotless/dotted I) are described to support correct multilingual ordering.
Practical applications and typical users
ISO 12199:2022 is used where consistent, language-agnostic ordering is required:
- Terminologists and lexicographers compiling multilingual glossaries, dictionaries and terminological databases.
- Software engineers and database designers implementing collation/sorting for multilingual search, index and UI lists.
- Localization and internationalization specialists ensuring consistent sort order across locales.
- Libraries, archives and content managers producing multilingual catalogues and indexes.
- Chemical database curators applying Annex C for name ordering.
Benefits include improved data interchange, predictable user experience in multilingual indexes, and compatibility with other standards-based sorting (e.g., ISO/IEC 14651).
Related standards
- ISO/IEC 14651 (formal collation specification) - informatively referenced and used as a formal model in Annex G.
- ISO 1087 (terminology vocabulary) - normative reference.
- ISO 10241-1 - complementary for terminological documentation.
Keywords: ISO 12199:2022, alphabetical ordering, multilingual collation, Latin alphabet, terminological data, lexicographical ordering, diacritics, character sequence, localization, data interchange.
Технические детали
- Технический комитет
- ISO/TC 37/SC 2 - Terminology workflow and language coding
- SKU
- ISO 12199:2022
Похожие стандарты
Стандарты, упомянутые в описании