Overview
ISO/IEC 23003-3:2020 - Unified Speech and Audio Coding (USAC) - is an international standard in the MPEG audio technologies family that defines a unified speech and audio codec. It is designed to efficiently code signals containing arbitrary mixes of speech and general audio, delivering perceptually transparent quality at high bitrates and very efficient coding at low bitrates while retaining full audio bandwidth. The standard integrates perceptual audio techniques with a source model of human speech to achieve robust compression across content types.
Key topics and technical requirements
- Codec scope and goals: Single codec for mixed speech/audio content with comparable or better performance than specialized speech or audio codecs.
- Support: Single and multi-channel coding, profiles and levels for interoperability, decoder behaviour and configuration (UsacConfig).
- Perceptual compression techniques incorporated:
- Perceptually shaped quantization noise
- Parametric coding of the upper spectral region (e.g., SBR/eSBR)
- Parametric stereo/stage coding (MPEG Surround)
- Source coding: Integration of a speech production model to improve coding of voiced speech.
- Bitstream and syntax: Detailed payloads, element definitions and buffer requirements for reliable decoding.
- Core tools described: Spectral noiseless coding, quantization, noise filling, TNS (Temporal Noise Shaping), filterbanks and block switching, time-warped filterbank, inter-subband temporal envelope shaping (inter-TES), joint stereo tools, and enhancements (eSBR, MPEG Surround, SAOC).
- Interoperability: Interfaces with MPEG Surround, SAOC, MPEG-D DRC, and backward compatibility considerations (e.g., AAC profiles).
Applications and users
ISO/IEC 23003-3:2020 is relevant for:
- Codec implementers - developers of audio encoders/decoders for software and hardware.
- Streaming services and broadcasters - seeking efficient delivery of mixed speech/music content across bandwidth-constrained networks.
- Mobile and IoT audio device manufacturers - where low-bitrate, full-bandwidth audio is required.
- Telecommunication and conferencing platforms - to improve speech intelligibility and mixed-content efficiency.
- Audio research and standards bodies - for extending or combining USAC with spatial/audio scene tools.
Practical benefits include simplified distribution (one codec for both speech and music), reduced bandwidth costs, and improved quality across diverse content types.
Related standards
- MPEG audio family (AAC, HE-AAC)
- MPEG Surround and SAOC (spatial upmix/scene coding)
- MPEG-D DRC (dynamic range control)
Refer to ISO/IEC documentation for normative references, codec profiles/levels, and full syntax and tool definitions.
Keywords: ISO/IEC 23003-3:2020, Unified Speech and Audio Coding, USAC, MPEG audio technologies, audio codec, speech and audio coding, perceptual coding, SBR, MPEG Surround.