Overview
ISO 24616:2012 - Multilingual information framework (MLIF) defines a generic platform for modelling and managing multilingual information across domains such as localization, translation, multimedia annotation, document management, digital libraries, and information or business modelling. MLIF provides a UML-specified metamodel plus a set of generic data categories (per ISO 12620:2009) and an XML serialization to enable consistent representation, linking and interoperability between formats such as XLIFF, TMX, smilText and ITS.
Key topics and technical requirements
- Metamodel (7 core components): MLDC (Multilingual Data Collection), GI (Global Information), GroupC (Grouping), MultiC (Multilingual Component), MonoC (Monolingual Component), HistoC (History), SegC (Segmentation).
- UML-based specification: The metamodel is defined using UML principles (subset relevant to MLIF).
- XML serialization: MLIF prescribes XML elements and attributes for serializing the metamodel, enabling machine-readable interchange.
- Data categories & adornment: MLIF uses ISO 12620 data categories to “adorn” model components (e.g., translationStatus, creationDate, matchQuality).
- Mandatory W3C attributes: xml:lang (mandatory on MonoC to indicate working language) and xml:id for unique identifiers.
- Segmentation & inline markup: SegC supports recursive segmentation and inline annotations (beginPairedTag, placeholder, genericPlaceholder) to preserve presentational features.
- Versioning & provenance: HistoC captures author, version, transaction and date for change tracking.
- Compliance modes: Implement MLIF fully from or embed MLIF-compliant elements (, , ) within other models.
Practical applications and who uses it
- Localization and translation tool vendors - to build interoperable translation memories and CAT tools that exchange TMX/XLIFF content reliably.
- Content managers and digital libraries - to manage multilingual collections with consistent metadata, provenance and segmentation.
- Multimedia and captioning teams - to synchronize text with audio/video (temporal synchronization: duration, begin, next).
- NLP and corpus engineers - to annotate corpora (morphology, POS, lemmas) and enhance translation prediction (see Annex A CAT example).
- Standards architects and integrators - to map or link domain models and ensure interoperability between language-resource formats.
Related standards
- ISO 12620:2009 - data category registry for language resources (used to adorn MLIF elements)
- ISO 24611 (MAF), ISO 24615 (SynAF), ISO 16642 (TMF) - complementary frameworks for morphological, syntactic and terminological description referenced for finer-grained annotation
- Interoperability targets: XLIFF, TMX, smilText, ITS
MLIF is an extensible, standards-based framework focused on interoperability and practical reuse of multilingual resources across localization, translation memory, multimedia annotation and digital content management workflows.