Overview - ISO/IEC 20382-2:2017 (Face-to-face speech translation: System architecture and functional components)
ISO/IEC 20382-2:2017 defines a high-level framework for face-to-face (F2F) speech translation system architecture and functional components. It specifies how speech translation devices, servers and communication protocols should interoperate across multiple languages to support convenient, real‑time spoken language exchange in face‑to‑face situations. The standard focuses on user interface behavior, system decomposition and functional requirements - not on implementing specific speech recognition, machine translation or speech synthesis engines.
Key technical topics and requirements
- Functional components: A typical F2F speech translation flow includes a speech recognition module, a language translation module and a speech synthesizer (TTS), coordinated via a user interface as described in ISO/IEC 20382-1.
- System architectures: The standard defines multiple architecture patterns (embedded devices, remote services, hybrid deployments, single fixed device, multi‑party sessions) and a sequence model for message flow between components.
- Performance and usability:
- Sessions should start naturally and quickly (target not exceeding 2 seconds).
- End‑to‑end translation latency should support real‑time interaction (target not exceeding 2 seconds).
- Support for multi‑user sessions and dynamic addition of participants.
- Privacy and remote services: Remote recognition, translation and synthesis services must protect user privacy.
- Data and format guidance:
- Recognized and synthesized text should use UTF‑8 encoding (RFC 2279).
- Speech modules should accept common speech formats; metadata-based formats (e.g., MIME) are recommended.
- The document does not mandate specific engine data formats - it remains engine‑agnostic.
- Translation strategy: If no direct language pair is available, use an intermediate language chosen for best performance (preferably same language family or similar word order when performance data is absent).
- Naturalness of output: Synthesized speech should preserve speaker attributes where possible (gender, base frequency, speed, prosody) to improve conversational naturalness.
Practical applications and who uses this standard
- Device manufacturers building mobile or fixed speech translation devices for retail, travel, healthcare or public services.
- Service providers and cloud vendors implementing remote speech recognition, translation and synthesis services with interoperable protocols.
- Protocol designers and integrators creating communication workflows between clients and translation servers.
- UX and accessibility designers focused on hands‑free, natural conversational interfaces that reduce language barriers.
- Standards bodies and implementers looking to ensure interoperability, privacy and real‑time performance across multilingual deployments.
Related standards and references
- ISO/IEC 20382‑1 (user interface considerations for F2F speech translation) - related part of the series.
- UTF‑8 encoding guidance referenced from IETF RFC 2279.
- Prepared by ISO/IEC JTC 1/SC 35 (User interfaces).
Keywords: face-to-face speech translation, ISO/IEC 20382-2:2017, system architecture, functional components, speech recognition, machine translation, speech synthesis, real-time translation, interoperability, translation servers, user interface.