speech & audio language models
pre-processing, evaluation, and artifacts in full-duplex speech systems
Speech language models are held back less by architecture than by data and measurement. Multi-turn, full-duplex audio is messy to prepare, and βdoes the model sound goodβ is not an evaluation protocol.
Sommelier (Jung et al., 2026) is a scalable, open pre-processing pipeline for multi-turn audio aimed at full-duplex speech language models β accepted to the ACL 2026 Industry Track.
At NAVER Cloud
As a WBL Residency Intern I built an end-to-end evaluation pipeline for Audio LLMs covering both understanding and generation, and implemented the multimodal benchmarking methodology used to validate HyperCLOVA X Omni (NAVER Cloud HyperCLOVA X Team, 2026) against industry-standard metrics for its technical report.
Recurring lesson: for audio, the evaluation harness is as much of a research contribution as the model.