ISO 24614-2:2011
Language resource management — Word segmentation of written texts — Part 2: Word segmentation for Chinese, Japanese and Korean
Language resource management — Word segmentation of written texts — Part 2: Word segmentation for Chinese, Japanese and Korean
- Статус документа:
- Действующий
- Формат:
- Электронный (PDF)
- Количество страниц:
- 49
- Дата публикации:
- 25 августа 2011 г.
- Издание:
- ISO IS 24614 edition 1 version 1
- ICS:
- 01.140.10
The basic concepts and general principles of word segmentation as defined in ISO 24614-1 apply to Chinese, Japanese and Korean. Text needs to be segmented into tokens, words, phrases or some other types of smaller textual units in order to perform certain computational applications on language resources, such as natural language processing, information retrieval and machine translation. ISO 24614-2:2011 is restricted to the segmentation of a text into words or other word segmentation units (WSUs). This task is distinct from morphological or syntactic analysis per se, although it greatly depends on morphosyntactic analysis. It is also different from the task of laying out a framework for constructing a lexicon and identifying its lexical entries, namely lemmas and lexemes. The frameworks for the latter tasks are provided by ISO 24611, ISO 24613 and ISO 24615. ISO 24614-2:2011 specifies rules for delineating WSUs for Chinese, Japanese and Korean. Some rules are common to all three languages, though each language also has its own distinct rules for identifying WSUs. The common features are discussed, then the distinct rules are laid out for Chinese, for Japanese and for Korean.
Abstract
Overview
ISO 24614-2:2011 - Language resource management - Word segmentation of written texts - Part 2: Word segmentation for Chinese, Japanese and Korean - defines rules for segmenting running text into word segmentation units (WSUs) for CJK languages. The standard focuses on delineating WSUs (tokens, words, sub-word units) required for computational processing and is explicitly limited to segmentation - not full morphological or syntactic analysis, nor lexicon design. It documents features common to the three languages and then provides language-specific rules (Chinese, Japanese, Korean), plus terms, markup conventions and examples (e.g., Japanese bunsetsu, Korean eojeol).
Key topics and technical requirements
- Scope and purpose: Applies ISO 24614-1 basic concepts and general principles to CJK word segmentation; restricted to segmentation tasks needed by NLP, IR and MT.
- Common rules: General criteria for identifying WSUs shared across Chinese, Japanese and Korean (compound formation, derivation, fixed expressions, abbreviations, loanwords, foreign-character strings).
- Language‑specific rules: Detailed segmentation rules for:
- Chinese (including handling of suffixes like 儿 (r) and parts of speech such as nouns, verbs, adjectives, numerals and measure words)
- Japanese (concepts such as bunsetsu, particles, verb/adjective endings)
- Korean (eojeol handling, grammatical affixes and agglutination)
- Terminology & markup: Definitions (adnoun, bunsetsu, eojeol, WSU) and recommended markup conventions to annotate segmentation consistently.
- Annex and bibliography: Comparative tables and references to support interoperable annotation.
Practical applications
ISO 24614-2:2011 provides a standardized basis for:
- Training and evaluating tokenizers and segmenters for NLP, machine translation, speech processing, and information retrieval.
- Creating interoperable, annotated corpora and language resources (corpus linguistics, lexicon alignment).
- Improving search indexing, text mining and data preparation in CJK languages where white-space tokenization is not sufficient.
- Ensuring consistent preprocessing in multilingual pipelines that include Chinese, Japanese or Korean.
Who should use this standard
- NLP engineers and data scientists building tokenization and preprocessing modules.
- Corpus linguists, annotators and standards implementers producing CJK annotated datasets.
- MT and IR system developers, lexicographers and localization teams requiring consistent segmentation.
- Academic researchers and software vendors working on language resources for Chinese, Japanese or Korean.
Related standards
- ISO 24614-1:2010 - Basic concepts and general principles (word segmentation)
- ISO 24611 - Morpho-syntactic annotation framework
- ISO 24613:2008 - Lexical markup framework (LMF)
- ISO 24615 - (related lexical/morpho-syntactic standards)
Keywords: ISO 24614-2:2011, word segmentation, Chinese word segmentation, Japanese word segmentation, Korean word segmentation, language resource management, WSU, NLP, machine translation, information retrieval.
Технические детали
- Технический комитет
- ISO/TC 37/SC 4 - Language resource management
- SKU
- ISO 24614-2:2011
Похожие стандарты
Стандарты, упомянутые в описании
ISO 24614-1:2010
ДействующийLanguage resource management — Word segmentation of written texts — Part 1: Basic concepts and general princi…
Overview - ISO 24614-1:2010 (Word segmentation, basic concepts) ISO 24614-1:2010 is an international standard in language resource management that defines the basic concepts and general principles fo…
ISO 24613:2008
ОтменёнLanguage resource management - Lexical markup framework (LMF)
BS ISO 24614-1:2010
ДействующийLanguage resource management. Word segmentation of written texts. Basic concepts and general principles.
ISO 24611:2012
ОтменёнLanguage resource management — Morpho-syntactic annotation framework (MAF)
Overview ISO 24611:2012 - Morpho-syntactic annotation framework (MAF) defines a standardized framework for representing morpho-syntactic annotations of word-forms in texts. It provides a meta-model t…