Our research program

Understanding speech holistically

Speech carries far more than words. Our goal is to understand everything the voice encodes, not just the text it resolves to, and to do it interpretably. We pursue this along two lines of work that converge on one ambition: speech health foundation models.

Measuring speech faithfully

Pinning down what actually happened in an utterance: every word, its timing, its speaker, and how it was said.

CrisperWhisper 2.0: verbatim speech recognition that captures what was actually said, every filler, repetition, false start, and vocal sound, controllable, multilingual, and open. You cannot study speech you have already cleaned up.

With the what, the when, the who, and the how in place, labeling stops being the bottleneck: we can annotate real-world speech at scale, at a quality that today only expert annotators reach.

Representing speech interpretably

Compressing speech into representations where each attribute lives in its own place, at its own natural granularity, instead of everything entangled in one stream.

where both lines converge

Speech health foundation models

Faithful measurement gives us annotation of real-world speech at expert quality and scale; factorized representations give us the substrate to learn from it. Together they are the data and the architecture we intend to build speech health foundation models on.