docs
One row per language decides what recognises it. A language without a row of its own follows the Auto-detect row, so out of the box a single model serves everything and you change only what you care about.
The language you speak. It picks the model, and it is what the whole tab reacts to. Auto-detect leaves the choice to the model, at the cost of a little speed and accuracy on short phrases.
Click a language and the models that know it unfold, with accuracy and speed bars, the memory each will really take, and a download arrow for the ones you do not have yet.
The common model: it serves every language without one of its own, and those rows are shown in grey. It also works out the spoken language when no source language is set.
Done by the recognition model itself. Whisper reaches every dialog language; a model that cannot translate says so in amber, and offers to hand the phrase to Whisper for that one dictation.