why I keep ending up in evaluation
Every project I have worked on started as a modeling question and ended as a measurement question.
Build an agent that extracts information from the web, and within a week the real question is no longer “can it extract” but “extract from what, verified how, and would this number still hold tomorrow.” Build an audio model, and the interesting part turns out to be that “understanding” and “generation” need entirely different harnesses, and that most reported numbers quietly cover only one of them.
This is not a complaint. It is the reason I find the area worth staying in.
A benchmark is a claim about what matters. When you freeze a crawl of the web into a static test set, you are claiming that drift does not matter. When you evaluate a speech model on transcription accuracy alone, you are claiming that everything else about the audio is incidental. Those claims are usually made implicitly, by whoever built the harness, and then inherited by everyone who reports on it afterwards.
So I would rather make the claim on purpose.
This is a starter post from when I set the site up — more to come.