Overview
ISO 24617-10:2024 - Language resource management - Semantic annotation framework (SemAF) - Part 10: Visual information (VoxML) specifies an annotation language for visual information based on VoxML (visual object concept structure modelling language). The standard defines how to encode semantic knowledge for 3D visualizations of concepts, objects, actions and motions denoted by natural language (NL). It is intended to support robust 3D simulation and visualization of actions taken by human and artificial agents and to integrate with the broader SemAF family for time and spatial annotation.
Key topics and requirements
- VoxML-based annotation scheme: A metamodel and concrete/abstract syntax for representing visual object concept structures (VoxML structures) in annotation formats.
- Semantic building blocks: Formalization of objects, actions as programs, relations, properties (attributes), functions, and examples of voxemes (basic entries in the voxicon).
- Basic semantic assumptions: Concepts such as habitats, affordances (Gibsonian and telic), qualia structures, and minimal embedding space (MES) are used to constrain model-theoretic interpretation.
- Representation & syntax: Guidance on representation of VoxML structures, including concrete syntax for visual information markup (visML) and mapping to feature-structure representations (see ISO 24610-1).
- Semantic interpretation: Specification of how VoxML enriches annotation semantics to enable 3D simulation, linking NL input to geometric and functional parameters.
- Conformance: The VoxML annotation scheme conforms to requirements in ISO 24617-1, ISO 24617-7 and ISO 24617-14 for temporal and spatial semantics.
Applications
- Multimodal simulation and visualization driven by natural language: generate 3D scenes from textual descriptions.
- Human-computer interaction and conversational agents that require situated visual understanding.
- AR/VR content generation and interactive training simulations where semantic behaviors and affordances must be modeled.
- Robotics and embodied AI for planning and simulating object interactions (grasping, moving, using).
- Corpus annotation and dataset creation for NLP, computer vision and multimodal research.
- Tooling for annotation pipelines, visualization engines and semantic parsers that require standardized visual-semantic markup.
Who should use this standard
- Computational linguists, NLP researchers and semantics engineers
- Annotation tool and dataset developers
- AR/VR/3D visualization and simulation platform developers
- Robotics and embodied AI researchers integrating NL-driven behaviors
- Standards bodies and interoperability architects working on multimodal systems
Related standards
Keywords: ISO 24617-10:2024, SemAF, VoxML, visual information, semantic annotation, 3D simulation, voxeme, voxicon, affordance, habitat, natural language visualization.