Overview
ISO 1951:2007 - Presentation/representation of entries in dictionaries - specifies a formal, media‑independent model for representing and presenting dictionary entries in both print and electronic formats. Focused on monolingual and multilingual, general and specialized dictionaries, the standard follows a lemma‑oriented approach (not concept‑oriented) and aims to facilitate the production, merging, comparison, extraction, exchange, dissemination and retrieval of lexicographical data.
Keywords: ISO 1951:2007, dictionary entry structure, lexicographical data, electronic dictionaries, XML encoding, lexicography.
Key topics and technical requirements
- Formal generic structure: Defines a computable model for dictionary entries that is independent of output media to ensure consistent representation across print and digital products.
- Data elements and compositional elements: Standardizes units such as lexical units (headwords, variants, translations) and compositional elements (blocks, containers, groups) to encode relationships unambiguously.
- Formal grammar: Uses formal syntax (Extended Backus‑Naur Form in the XmLex model) to describe high‑level structures and content models so entries can be parsed and validated automatically.
- Content models and qualifiers: Specifies content elements, embedded elements and general qualifiers to control permitted values and element relationships.
- Presentation aids: Recommends layout aids and compacting mechanisms (abbreviations, repeated headword handling, nesting, lexicographical symbols) to support readable print and electronic presentations.
- XML implementation examples: Informative annexes provide XML/XSL/XHTML examples (XmLex_V00) and validation approaches to guide practical encoding and transformation.
- Objective compliance language: Uses "shall"/"should"/"can" to indicate requirements, recommendations and informative guidance.
Applications and practical value
- Enables single‑source lexicography where one structured data repository can generate multiple outputs (print, web, mobile apps).
- Facilitates interoperability and data exchange between publishers, lexicographical toolchains, translation memory systems and lexical databases.
- Supports automated tasks: merging dictionaries, inverting bilingual dictionaries, extracting datasets for NLP, searching and retrieval of lexicographical data.
- Guides publishers and developers in implementing consistent presentation and encoding practices for electronic dictionaries and hybrid publishing workflows.
Who should use ISO 1951:2007
- Lexicographers and dictionary publishers
- Language resource engineers and NLP developers
- Software developers building dictionary editors, converters or publishing pipelines
- Terminologists, librarians and digital humanities professionals involved in lexical data management
Related standards
ISO 1951:2007 is a practical standard for anyone needing standardized, exchangeable, and media‑independent dictionary entry structures and presentation methods.