ISO 24615-1:2014
Language resource management — Syntactic annotation framework (SynAF) — Part 1: Syntactic model
Language resource management — Syntactic annotation framework (SynAF) — Part 1: Syntactic model
- Статус документа:
- Действующий
- Формат:
- Электронный (PDF)
- Количество страниц:
- 25
- Дата публикации:
- 5 февраля 2014 г.
- Издание:
- ISO IS 24615 edition 1 version 1
- ICS:
- 01.020
ISO 24615-1:2014 describes the syntactic annotation framework (SynAF), a high level model for representing the syntactic annotation of linguistic data, with the objective of supporting interoperability across language resources or language processing components. ISO 24615-1:2014 is complementary and closely related to ISO 24611 (MAF, morpho-syntactic annotation framework) and provides a metamodel for syntactic representations as well as reference data categories for representing both constituency and dependency information in sentences or other comparable utterances and segments.
Abstract
Overview
ISO 24615-1:2014 - Language resource management: Syntactic annotation framework (SynAF) - Part 1: Syntactic model defines a high-level metamodel and reference data categories for representing syntactic annotation of linguistic data. Its primary goal is to support interoperability between language resources (e.g., treebanks) and language processing components (parsers, taggers) by providing a common conceptual model for both constituency and dependency information.
Keywords: ISO 24615-1:2014, SynAF, syntactic annotation, syntactic model, language resource management, interoperability, constituency, dependency, data categories.
Key topics and technical requirements
- SynAF metamodel (UML-based): Describes classes and relationships used to model syntactic structure. Key classes include SyntacticNode, T_Node (terminal nodes / word forms and empty elements), NT_Node (non-terminal/phrasal nodes), SyntacticEdge (relations between nodes), and Annotation.
- Constituency and dependency coverage: The model explicitly supports hierarchical constituency (phrases, clauses, sentences) and dependency relations (head–modifier, subject–predicate), allowing both representations and their interrelation.
- Multi-layered annotation strategy: Encourages layered annotations so constituency and dependency information can coexist and reference morpho-syntactic annotations (integration with MAF).
- Spans and discontinuous constituents: Terminal nodes are defined over spans (including multiple spans) to account for discontinuous or non-contiguous elements.
- Data categories and DCS: Annex A lists normative data categories; implementers are expected to define a Data Category Selection (DCS) using the specified categories (available via ISOCat).
- Interoperability requirements: The standard is designed to be complementary to ISO 24611 (MAF) and compatible with other language resource frameworks; it mandates use of the provided metamodel and data category registry for consistent annotation exchange.
Practical applications
- Building and sharing treebanks that combine constituency and dependency annotations.
- Designing annotation schemes and tools (annotation editors, conversion utilities) that must interoperate with other language resources.
- Standardizing parser output formats to be compatible with corpora and downstream NLP components.
- Integrating morpho-syntactic (MAF) and syntactic annotations in multi-layer corpora for research and production NLP pipelines.
Who should use this standard
- NLP engineers and parser developers
- Corpus linguists and annotation teams
- Language technology vendors and research groups
- Tool and format designers working on syntactic annotation, treebanking, or resource conversion
Related standards
- ISO 24611:2012 (MAF - Morpho-syntactic Annotation Framework)
- ISO 24612 (LAF - Linguistic Annotation Framework)
- ISO 24613 (LMF - Lexical Markup Framework)
- ISO 12620:2009 (data category registry / ISOCat)
Using ISO 24615-1:2014 helps ensure consistent syntactic annotation, improves resource reusability, and simplifies integration across language-processing components.
Технические детали
- Технический комитет
- ISO/TC 37/SC 4 - Language resource management
- SKU
- ISO 24615-1:2014
Похожие стандарты
Стандарты, упомянутые в описании
ISO 24611:2012
ОтменёнLanguage resource management — Morpho-syntactic annotation framework (MAF)
Overview ISO 24611:2012 - Morpho-syntactic annotation framework (MAF) defines a standardized framework for representing morpho-syntactic annotations of word-forms in texts. It provides a meta-model t…
ISO 12620:2009
ОтменёнTerminology and other language and content resources — Specification of data categories and management of a D…
ISO 24612:2012
ДействующийLanguage resource management — Linguistic annotation framework (LAF)
Overview ISO 24612:2012 - Language resource management - Linguistic annotation framework (LAF) - defines a standard framework for representing linguistic annotations of primary language data (text, s…