Overview
SIST ISO 24620-5:2024 - part of the ISO 24620 series on Language resource management (Controlled Human Communication, CHC) - defines lexico‑morpho‑syntactic principles and a formal methodology for recognizing and protecting personal data in free text. The standard targets texts in multiple languages (agglutinating, inflectional, isolating) and across countries and domains (e.g. law, finance, health). It focuses on formal, rule‑based approaches (intension‑based representations) rather than statistical methods and does not cover automated image processing.
Key technical topics and requirements
- Scope and applicability
- Applies to human and automated processing of free text containing personal data.
- Designed for cross‑language and cross‑country use; supports different language families and varying national formats (e.g. addresses).
- Linguistic building blocks
- Uses lexical, morphological and syntactic indicants to detect personal data (names, addresses, phone numbers, IDs, bank account numbers, etc.).
- Introduces the notion of seme (Saussurean signified and its instantiations in text) and indicants (occurrences of lexical/morphological/syntactic phenomena that signal personal data).
- Formal methodology
- Specifies intension‑based formal representations and a meta‑language/grammar to express recognition rules.
- Requires an ordered system of constraint rules and an associated algorithm that, when applied, extracts or flags personal data instances.
- Emphasizes explainability and extensibility so new semes and languages can be added.
- Protection options
- After detection, personal data can be masked, removed, anonymized or pseudonymized following organizational or legal requirements.
- Limitations
- Excludes automated image processing and excludes statistical/machine‑learning methods (the standard focuses on formal rule‑based techniques).
Practical applications and users
Who benefits:
- Data Protection Officers and compliance teams ensuring GDPR/CCPA alignment when sharing or processing text.
- NLP engineers and language technologists building rule‑based redaction or entity‑recognition systems.
- Software vendors delivering text anonymization, pseudonymization, and redaction tools.
- Public bodies, legal, finance and healthcare organizations that must protect personal data in documents and communications.
Practical uses:
- Automated and semi‑automated redaction workflows for document sharing and litigation.
- Sanitizing training corpora and logs before machine learning or analytics.
- Pre‑processing text for cross‑border data transfer while meeting regulatory safeguards.
- Creating explainable, language‑agnostic pipelines for personal data recognition.
Related standards (for context)
- ISO 24620 series (other CHC parts)
- ISO/IEC 27701 (privacy information management guidance)
- GDPR and other regional privacy regulations (as motivating examples for protection measures)
Keywords: SIST ISO 24620-5:2024, personal data recognition, personal data protection, lexico-morpho-syntactic, language resource management, CHC, text anonymization, pseudonymization, GDPR, rule-based redaction.