公開的語音AI基準測試越來越常顯示模型已達到人類水準。然而,這些分數並不總是能反映模型在實際應用中的表現。由於公開基準測試開放且廣泛使用,模型可能會針對這些測試本身進行最佳化。它們的分數可能因此提高,是因為學會了基準測試特有的模式,而非在底層任務上變得更好。

其中一個原因是,傳統基準測試忽略了許多使語音系統在實踐中可靠、自然、符合語境且有效的條件和品質。這就是為什麼我們最近在 Real World VoiceEQ、Open-ASR Leaderboard 和 Far-field ASR Leaderboard 中引入了獨立測試集(held-out sets),旨在衡量更多實際應用中重要的指標。

然而,僅僅擴大衡量範圍並不能解決問題。這種現象有時被稱為「基準測試最佳化」(benchmark optimization)或「基準測試作弊」(benchmaxxing),在機器學習領域常被討論,但在語音辨識中卻難以量化。

我們最新的研究引入了三種測試方法來幫助量化這個問題。我們評估了 11 個廣泛使用的開源 ASR 模型,發現幾個得分最高的系統會複製 VoxPopuli 英文和 LibriSpeech(clean, other)資料集中的基準測試轉錄稿,即使音訊內容與之矛盾、相關詞彙已被靜音,或者音訊同樣支持兩種不同的書寫形式。

在某些情況下,模型似乎不僅依賴所說的內容,還依賴細微的聲學線索來判斷它們正在接受哪個基準測試。因此,它們的分數誇大了它們在更廣泛語音轉錄方面的能力。

參考轉錄稿分歧(VoxPopuli 案例研究)

VoxPopuli 以包含大量轉錄錯誤而聞名(這也是 Artificial Analysis 發布了清理版本的原因)。我們的「共識分歧探測」測試了領先的 ASR 模型遇到這些錯誤時會發生什麼:它們是準確地轉錄音訊內容,還是複製基準測試中不正確的參考轉錄稿?

為了大規模測試這一點,我們使用了一組獨立模型,這些模型因其低音素錯誤率(PER)而被選中。PER 衡量書面轉錄稿與音訊中聲音的匹配程度,使其成為衡量模型忠實轉錄所聽內容的有用指標。這組模型的結果可用於標記模型一致不同意基準測試參考轉錄稿的情況。然後,我們將這些標記案例的樣本與人工註釋進行比較,以驗證更正後的轉錄稿。

例如,一個 VoxPopuli 片段中清晰可聞「Thank you, Mr. President」這句話,但參考轉錄稿卻省略了「Thank you」。我們測試的 11 個模型中有 6 個複製了基準測試的錯誤轉錄稿,給出了「預期」的答案,儘管這與音訊內容相矛盾。

在實際片段中,格式也遵循相同的模式:省略「Thank you」的模型也會複製基準測試的標點符號風格,將「Mr」寫成沒有句點,而包含可聞短語的模型則傾向於將「Mr.」寫成帶有句點。

當我們以新收集的歐盟議會錄音或通用聲音呈現相同內容時,這種行為通常會減弱或消失。在下面的範例中,除了其中一個模型外,所有模型都恢復了對新議會錄音複製音訊的忠實轉錄。這表明模型正在響應聲學線索,這些線索幫助它們識別基準測試的成員身份,從而產生預期的轉錄稿,即使它與音訊內容相矛盾。

此片段的參考轉錄稿為「Mr President, I have another complaint about this procedure, which is that it is not secret.」。下面所有三個片段的音訊內容實際上都說了同樣的話,前面帶有清晰可聞的「Thank you,」,這些複製音是該真實句子的文字轉語音版本,因此在所有三個片段中都能聽到這句禮貌語。

綠色高亮和 ✅ 標記了包含可聞「Thank you」的轉錄稿;紅色高亮和 ❌ 標記了複製基準測試錯誤省略的轉錄稿。所有轉錄稿都是未經任何正規化的原始模型輸出,大小寫和標點符號完全按照生成時的樣子保留,包括某些模型的小寫輸出。

| Model | Real clip | Same-speaker clone | ep-fresh clone |

|---|---|---|---|

| CohereLabs/cohere-transcribe-03-2026 | ❌ Mr President… | ❌ Mr President… | ✅ Thank you, Mr President… |

| nvidia/canary-qwen-2.5b | ❌ Mr President… | ❌ Mr President… | ✅ Thank you Mr. President… |

| ibm-granite/granite-speech-4.1-2b | ❌ mr president… | ❌ mr president… | ✅ thank you mr president… |

| microsoft/Phi-4-multimodal-instruct | ❌ Mr President… | ❌ Mr President… | ❌ Mr President… |

| nvidia/parakeet-tdt-0.6b-v2 | ❌ Mr President… | ✅ Thank you, Mr President… | ✅ Thank you, Mr. President… |

| bosonai/higgs-audio-v3-8b-stt-v2 | ❌ mr president… | ❌ mr president… | ✅ thank you mr president… |

| Qwen/Qwen3-ASR-0.6B-hf | ✅ Thank you, Mr. President… | ✅ Thank you, Mister President… | ✅ Thank you, Mister President… |

| mistralai/Voxtral-Mini-3B-2507 | ✅ Thank you, Mr. President… | ✅ Thank you, Mr. President… | ✅ Thank you, Mr. President… |

| moonshotai/Kimi-Audio-7B-Instruct | ✅ Thank you, mr. President… | ✅ Thank you, Mr. President… | ✅ Thank you, mr. President… |

| openai/whisper-large-v3 | ✅ Thank you, Mr. President… | ✅ Thank you, Mr. President… | ✅ Thank you, Mr. President… |

| moonshine-ai/moonshine-streaming-medium | ✅ thank you mr president… | ✅ thank you mr president… | ✅ thank you mr president… |

| 11 個模型中省略禮貌語(❌)的數量 | 6 | 5 | 1 |

Parakeet 是唯一一個在真實片段上複製基準測試錯誤,但在同說話者複製音上正確轉錄的模型。Phi-4 是唯一一個在 ep-fresh 複製音上仍然省略禮貌語的模型。當我們改用與任何議會錄音無關的通用 TTS 聲音重新合成該句子時,所有 11 個模型都恢復了禮貌語。

結果表明這個問題既普遍又具有意義。我們的研究方法在我們分析的 VoxPopuli 測試片段中,有 40% 標記出潛在的參考錯誤,影響了所有參考詞彙的大約 3%。

表現出基準測試最佳化行為的模型,在 18% 到 30% 的時間裡複製了錯誤的參考轉錄稿。下面的散佈圖比較了 X 軸上的 VoxPopuli 詞錯誤率(WER)與每個模型複製基準測試錯誤參考而非共識修正的頻率。詞錯誤率最低(因此報告的基準測試表現最強)的模型,也最有可能複製這些錯誤。

遮蔽實體檢索

為了進一步探究共識分歧問題,我們故意將測試資料集音訊樣本中的數字靜音,並要求模型轉錄其所聽到的內容。由於音訊中確實沒有這些數字,模型不應輸出任何數字,更不應輸出文本中的確切數字。

其中一些數字是半可預測的(儘管模型預測到的可能性仍然很低),但另一些則相當令人驚訝。下面的片段結合了兩種探測方法,展示了模型如何重現包含不正確數字的參考轉錄稿錯誤,甚至有一個模型在數字被靜音的情況下,自動補齊了一個相對隨機的年份(2011)。在下面每個模型的行中:

  • 帶刪除線的綠色高亮標記了模型正確地沒有複製的參考轉錄稿詞彙(忠於音訊);
  • 帶底線的綠色高亮標記了取代參考轉錄稿錯誤措辭的正確、忠於音訊的插入內容;
  • 紅色高亮(純文本)複製了參考轉錄稿中錯誤且音訊不支持的內容:保留「Mr President」、在音訊說「one thousand six hundred」時寫成「more than 1 amendments」、提供被靜音的年份「2011」,或以「plenary」結尾。

| | Reference | What the audio says |

|---|---|---|

| | Mr President, in the Committee on Budgets, we voted on more than 1 amendments to the 2011 draft budget … voted in the plenary. | In the Committee on Budgets, we voted on more than one thousand six hundred amendments to the ⟨silenced⟩ draft budget … voted in the … |

| CohereLabs/cohere-transcribe-03-2026 | Mr President, in the Committee on Budgets we voted on more than 1 amendments to the 2011 draft budget … voted in the plenary. |

| nvidia/canary-qwen-2.5b | Mr President, in the Committee on Budgets we voted on more than one amendments to the 2011 draft budget … voted in the plenary |

| ibm-granite/granite-speech-4.1-2b | Mr President in the committee on budgets we voted on more than one thousand six hundred amendments to the 2011 draft budget … voted on in the plenary |

| microsoft/Phi-4-multimodal-instruct | Mr President In the Committee on Budgets we voted on more than 1 amendments to the 2011 draft budget … voted on in the plenary. |

| nvidia/parakeet-tdt-0.6b-v2 | Mr President In the Committee on Budgets we voted on more than one amendments to the 2011 draft budget … voted in the Protestants. |

| bosonai/higgs-audio-v3-8b-stt-v2 | Mr President in the committee on budgets we voted on more than one thousand six hundred amendments to the 2011 draft budget … voted in the plenary |

| Qwen/Qwen3-ASR-0.6B-hf | Mr President In the Committee on Budgets, we voted on more than 1,600 amendments to the 2011 draft budget … voted in the plenary |

| mistralai/Voxtral-Mini-3B-2507 | Mr President In the Committee on Budgets, we voted on more than 1,600 amendments to the 2011 draft budget … voted in the plenary |

| moonshotai/Kimi-Audio-7B-Instruct | Mr President Ah in the committee on budgets we voted on more than one thousand six hundred amendments to the 2011 draft budget … voted in the plenary |

| openai/whisper-large-v3 | Mr President In the Committee on Budgets, we voted on more than 1,600 amendments to the 2011 draft budget … voted in the plenary |

| moonshine-ai/moonshine-streaming-medium | Mr President in the committee on budgets we voted on more than one thousand six hundred amendments to the 2011 draft budget … voted in the plenary |

在公開基準測試上,數字的「恢復率」最高,而在獨立測試集或新收集的音訊(下面的 ep-fresh 和 libri-fresh)上則較低。在 LibriSpeech 上,一些基準測試表現最強的模型在大約 30-40% 的範例中複製了被遮蔽的數字,即使該數字本身已被移除。對於幾個模型來說,這種效應在新收集的資料上有所減弱,這表明周圍與基準測試相關的音訊,而不僅僅是文本自動補齊,幫助模型恢復了參考轉錄稿。

拼寫切換

我們的「拼寫切換探測」測試模型是否會複製基準測試參考轉錄稿中使用的確切拼寫,即使音訊中並不明確。拼寫變體是指語義和語音相同,但拼寫方式不同的詞彙(例如 1 與 one、Mr. 與 mister、John 與 Jon、Honor 與 Honour 等)。

理論上,模型應該始終偏好一種拼寫,或者以大致隨機的頻率在它們之間切換。如果模型系統性地切換以匹配每個基準測試參考轉錄稿中的內容,這表明模型正在識別測試預期的拼寫方式。

轉錄稿:「I URGED ON THE BOYS THAT WHATEVER HAPPENED WE SHOULD NOT SHOOT ANY ONE」,使用「any one」的模型:6/11,使用「anyone」的模型:5/11

轉錄稿:「CAMOUFLAGE WAS NOT A WORD THE CAPTAIN OR ANYONE ELSE OF HIS TIME YET UNDERSTOOD」,使用「any one」的模型:2/11,使用「anyone」的模型:9/11

在 LibriSpeech 中,我們測試了一種涉及舊式間距約定的資料集內切換:一些參考轉錄稿使用「any one」,而另一些則使用「anyone」。我們測量了給定變體的最低準確度,我們稱之為「切換率」(switch rate)。如果一個模型只使用一種變體,其切換率將為 0%;而一個模型如果