<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://hvvan.github.io/feed.xml" rel="self" type="application/atom+xml"/><link href="https://hvvan.github.io/" rel="alternate" type="text/html" hreflang="en"/><updated>2026-09-08T02:44:40+00:00</updated><id>https://hvvan.github.io/feed.xml</id><title type="html">blank</title><subtitle>Jihwan Kim — integrated M.S./Ph.D. student at KAIST GSAI, advised by Professor Jaegul Choo. Research interests: evaluation of agents, speech and audio processing, and multilingual LLMs. </subtitle><entry><title type="html">why I keep ending up in evaluation</title><link href="https://hvvan.github.io/blog/2026/why-evaluation/" rel="alternate" type="text/html" title="why I keep ending up in evaluation"/><published>2026-07-01T01:00:00+00:00</published><updated>2026-07-01T01:00:00+00:00</updated><id>https://hvvan.github.io/blog/2026/why-evaluation</id><content type="html" xml:base="https://hvvan.github.io/blog/2026/why-evaluation/"><![CDATA[<p>Every project I have worked on started as a modeling question and ended as a measurement question.</p> <p>Build an agent that extracts information from the web, and within a week the real question is no longer “can it extract” but “extract from <em>what</em>, verified <em>how</em>, and would this number still hold tomorrow.” Build an audio model, and the interesting part turns out to be that “understanding” and “generation” need entirely different harnesses, and that most reported numbers quietly cover only one of them.</p> <p>This is not a complaint. It is the reason I find the area worth staying in.</p> <p>A benchmark is a claim about what matters. When you freeze a crawl of the web into a static test set, you are claiming that drift does not matter. When you evaluate a speech model on transcription accuracy alone, you are claiming that everything else about the audio is incidental. Those claims are usually made implicitly, by whoever built the harness, and then inherited by everyone who reports on it afterwards.</p> <p>So I would rather make the claim on purpose.</p> <hr/> <p><em>This is a starter post from when I set the site up — more to come.</em></p>]]></content><author><name></name></author><category term="evaluation"/><category term="llm-agents"/><summary type="html"><![CDATA[a short note on why the harness is the research]]></summary></entry></feed>