Overview
ISO/IEC 23092-2:2024 - "Information technology - Genomic information representation - Part 2: Coding of genomic information" specifies how to encode core genomic data for efficient storage, exchange and processing. The standard covers representation of unaligned sequencing reads (including read identifiers and quality values), aligned sequencing reads (including identifiers and quality values), and reference sequences. It defines syntax, semantics, data structures and decoding behavior for coded genomic information.
Keywords: ISO/IEC 23092-2:2024, genomic information representation, coding of genomic information, sequencing reads, reference sequences, genomic data formats
Key topics and technical requirements
- Conventions and syntax: formal conventions, operators, range notation and bit ordering for unambiguous encoding and parsing.
- Data structures: definitions for foundational elements such as data units, raw reference, parameter sets, and access units that group encoded genomic objects.
- Access unit types and dataset types: multiple AU types (e.g., Classes P, N, M, I, HM, U) and dataset_type variants are specified; decoding rules and AU semantics are described.
- Descriptors and sequencing reads: descriptor schemas for positional and mapping information (pos, rcomp, flags, mmpos, mmtype, clips, etc.), and rules for representing paired-end reads and reverse-complement reads.
- Quality values and identifiers: standardized representation for read identifiers and quality (QV) values to preserve sequencing metadata.
- Decoding process: explicit decoding flows and semantics for different dataset and AU types to ensure interoperable implementations.
- Extensibility and parameters: parameter sets and encoding parameter constructs that enable configuration of encoding behavior.
Practical applications and users
ISO/IEC 23092-2:2024 is intended for organizations that produce, store, transmit or analyze large-scale sequencing data:
- Bioinformatics and genomics software developers (encoders, decoders, compression tools)
- Sequencing platform vendors and laboratory informatics teams
- Cloud storage and data platform architects for genomic data lakes
- Standards bodies and implementers building interoperable genomic data pipelines
- Clinical genomics and research centers requiring robust, standardized formats for reads, alignments and references
Practical benefits include improved interoperability, consistent decoding semantics, and support for compact, deterministic encoding of reads and references for downstream analysis, archival and transmission.
Related standards
- Other parts of the ISO/IEC 23092 series (adjacent parts address broader genomic information representation and packaging) complement Part 2 by covering metadata, containerization and related functions. Implementers should review the full series for end-to-end genomic data workflows.
For adoption and implementation, developers should reference ISO/IEC 23092-2:2024 for detailed syntax tables, descriptor definitions and decoding algorithms.