Overview
ISO 24611-1:2025 - Language resource management - Morphosyntactic annotation framework (MAF) - Part 1: Core model defines a standardized metamodel and XML serialization for representing morphosyntactic annotations of word-sized units in texts. The standard separates surface-level tokens (text segmentation) from higher-level word-forms (lexical abstractions) and specifies how morphosyntactic properties (e.g., part-of-speech, number, gender) are encoded, referenced and exchanged. The first edition (2025) replaces ISO 24611:2012 and introduces a full TEI XML serialization and updated conformance rules.
Key technical topics and requirements
- MAF metamodel: Clear distinction between token-level segmentation and word-form-level abstractions; supports n-to-n relationships and local graph representations for multi-token and ambiguous constructions.
- Token segmentation strategies: Inline markup () and stand-off annotation options; guidance on adjacent, overlapping and discontinuous tokens; normalization and script conversion considerations.
- Word-form representation: `` constructs, token attachment patterns (one-to-one, one-to-many, many-to-one, discontinuous, zero-token cases), compound word-forms and lexical references.
- Morphosyntactic content: Use of feature structures, compact morphosyntactic tags and FSR (feature structure representation) libraries; design guidance for tagsets (without prescribing specific tag compositions).
- Ambiguity handling: Mechanisms for representing alternative analyses and lexical/word-form ambiguities (structural ambiguities are out of scope and deferred to other parts).
- Metadata and conformance: Metadata recommendations and conformance clauses for implementers to ensure interoperability.
- Serialization and interoperability: XML serialization aligned with the TEI Guidelines; ability to reference external data categories in a repository conforming to ISO 12620-2 (data category repository).
Practical applications
- Standardized POS (part-of-speech) annotation and morphosyntactic tagging for corpora.
- Interchange format for linguistic resources used by NLP pipelines, corpus linguistics, digital humanities and lexicography.
- Integration of annotated corpora with lexical databases and dictionaries via explicit references to lexical entries.
- Annotation tool development: enables consistent inline and stand-off annotation strategies and interoperability across tools.
- Data archival and sharing: promotes stable semantics by linking annotations to persistent data categories.
Who should use this standard
- Computational linguists, corpus engineers and NLP developers
- Annotation tool and language-resource platform vendors
- Digital humanities researchers and lexicographers
- Standards bodies and data curators aiming for interoperable morphosyntactic resources
Related standards
- ISO 12620-2 (data category repository - referencing of data categories)
- TEI Guidelines (for XML encoding conventions)
- ISO 24619 (persistent identifiers for data categories)
- Other parts of the ISO 24611 series (planned extensions, e.g., word lattices)
Keywords: ISO 24611-1:2025, MAF, morphosyntactic annotation, POS annotation, tokenization, TEI XML, ISO 12620-2, language resource management, feature structures.