speech & audio language models

pre-processing, evaluation, and artifacts in full-duplex speech systems

Speech language models are held back less by architecture than by data and measurement. Multi-turn, full-duplex audio is messy to prepare, and β€œdoes the model sound good” is not an evaluation protocol.

Sommelier (Jung et al., 2026) is a scalable, open pre-processing pipeline for multi-turn audio aimed at full-duplex speech language models β€” accepted to the ACL 2026 Industry Track.

At NAVER Cloud

As a WBL Residency Intern I built an end-to-end evaluation pipeline for Audio LLMs covering both understanding and generation, and implemented the multimodal benchmarking methodology used to validate HyperCLOVA X Omni (NAVER Cloud HyperCLOVA X Team, 2026) against industry-standard metrics for its technical report.

Recurring lesson: for audio, the evaluation harness is as much of a research contribution as the model.

References

2026

  1. ACL
    Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models
    Kyudan Jung, Jihwan Kim, Soyoon Kim, Jeonghoon Kim, Jaegul Choo, and Cheonbok Park
    In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL): Industry Track, 2026
  2. HyperCLOVA X 8B Omni
    NAVER Cloud HyperCLOVA X Team
    arXiv preprint arXiv:2601.01792. Contributed as a Residency Intern , 2026