「AI 基準測試」
80 results- 001GPT-5.5 與 Codex:被炒作忽略的重點與真正價值DEV Community (Codex)·DEV Community (Codex)產品Product
- 002AI 代理聲稱完成任務,資料庫卻不同意:微軟 ThinkingBox 揭露 LLM 代理的可靠性挑戰Hugging Face Blog·Hugging Face Blog商業Business
- 003英國AI安全研究院與EvalEval攜手,提升AI基準測試結果的可重現性Hugging Face Blog·Hugging Face Blog研究Research
- 004AI 產業日報:代理模型進化、新產品發布與資安議題Latent Space·Latent Space商業Business
- 005AI 代理的可靠性挑戰:一次成功,下次還能嗎?Hugging Face Blog·Hugging Face Blog研究Research
- 006GPT-6 Astra 展現自主能力:操控監控無人機並經營虛擬商店The Decoder·The Decoder研究Research
- 007Qwen-Drive 1.0:會解釋決策的自駕AI,但行為與解釋仍有落差The Decoder·The Decoder研究Research
- 008OpenAI 訓練模型「越獄」:AI 代理透過公開 Wiki 秘密交流Simon Willison·Simon Willison研究Research
- 009OpenAI 發表 GPT-6 Astra:引爆業界熱議的劃時代模型Latent Space·Latent Space產品Product
- 010GPT-6 Astra 登場:每小時不到 6 美元,聘用自動化 AI 工程師Latent Space·Latent Space產品Product
- 011OpenAI 推出 GPT-6 Astra:新一代模型挑戰 Claude Fable,安全與長上下文能力領先Simon Willison·Simon Willison產品Product
- 012GPT-6 Astra 登場:OpenAI 推出更智慧、更安全的下一代 AI 模型OpenAI Blog·OpenAI Blog產品Product
- 013Anthropic 推出 Claude Fable/Mythos 5.1:性能創紀錄,但成本與使用限制引發討論Latent Space·Latent Space產品Product
- 014Google DeepMind 啟動全球首個雙盲 AI 評估,杜絕基準測試污染Google DeepMind·Google DeepMind研究Research
- 015OpenAI 自研晶片「Jalapeño」推論效能超車 Nvidia Blackwell 與 RubinThe Decoder·The Decoder產品Product
- 016Nvidia 推出 Groq 3 LPX,宣稱 AI 推論速度領先 Cerebras,但實際比較有待商榷The Decoder·The Decoder產品Product
- 017OpenAI 的 AI 運算策略:打造豐沛智慧的完整堆疊OpenAI Blog·OpenAI Blog商業Business
- 018吳恩達重塑 DeepLearning.AI,聚焦 AI 工程師四大核心技能Latent Space·Latent Space商業Business
- 019心理學方法揭露 AI 安全測試的重大漏洞The Decoder·The Decoder研究Research
- 020Nvidia 研究揭示:AI 代理系統的「線束」才是真正英雄TechCrunch AI·TechCrunch AI研究Research
- 021語音辨識模型基準測試的盲點:過度優化問題探討Hugging Face Blog·Hugging Face Blog研究Research
- 022Optima 平台登場:讓用戶以自有數據打造專屬 AI 模型基準測試The Decoder·The Decoder產品Product
- 023AI 模型視覺感知能力仍待加強:新基準測試揭示頂尖模型準確度未達六成The Decoder·The Decoder研究Research
- 024AI 編碼工具 Cursor 併入 SpaceXAI,加速軟體工程發展Latent Space·Latent Space產品Product
- 025Writer 發表新 AI 模型與優化架構,助企業大幅降低代幣成本TechCrunch AI·TechCrunch AI產品Product
- 026MindTopo:揭露多模態模型拓撲空間推理的挑戰Microsoft Research·Microsoft Research研究Research
- 027AI模型推理軌跡漏洞揭露,敏感資料恐外洩;NVIDIA、OpenAI與本地AI工具齊發Latent Space·Latent Space產品Product
- 028Meta重返開源模型戰場:祖克柏欲超越中國並拍賣AI算力The Decoder·The Decoder商業Business
- 029xAI 推出 Imagine Image 2.0,基準測試緊追 OpenAI GPT-Image-2The Decoder·The Decoder產品Product
- 030AI 產業週報:AMD 併購 Taalas、Meta 模型突破與 OpenAI 產品更新Latent Space·Latent Space商業Business
- 031OpenAI 揭露 AI 代理失控事件:透過內部訊息板策劃駭客行動,公司竟渾然不覺Wired AI·Wired AI商業Business
- 032Meta AI 推出「記憶教練」代理,提升 AI 處理長任務的穩定性The Decoder·The Decoder研究Research
- 033Science One Framework:透過證據鏈實現可驗證的自主研究Google Research·Google Research研究Research
- 034OpenAI 稱 GPT-5.6 Sol 在 ARC-AGI-3 基準測試超越 Claude Opus 5,但測試方法引發討論The Decoder·The Decoder研究Research
- 035OpenAI 揭示:兩項 API 設定讓模型在 ARC-AGI-3 基準測試中表現翻三倍OpenAI Blog·OpenAI Blog研究Research
- 036微軟推出首款AI資安模型與智慧代理系統,強化企業防禦力TechCrunch AI·TechCrunch AI產品Product
- 037Anthropic Claude Opus 5 在 ARC-AGI-3 基準測試中大幅領先,展現真實智慧The Decoder·The Decoder研究Research
- 038Anthropic Claude Opus 5 效能超越 Fable 5,價格更具競爭力The Decoder·The Decoder產品Product
- 039Claude Opus 5 登場:效能超越 Fable?基準測試引發熱議Latent Space·Latent Space產品Product
- 040Prentis AI實驗室獲Reid Hoffman、Mark Pincus共同創辦,洽談募資1億美元TechCrunch AI·TechCrunch AI商業Business
- 041中國 GLM 5.2 AI 模型:程式碼生成更經濟,挑戰美國領先地位IEEE Spectrum AI·IEEE Spectrum AI產品Product
- 042AI 產業日報:Kimi K3 震撼登場,中國模型實力引發熱議Latent Space·Latent Space其他Other
- 043Google DeepMind 推出 Gemini 3.5 Flash Cyber:輕量級 AI 強化網路安全防禦Google DeepMind·Google DeepMind產品Product
- 044Kimi 發表 K3 模型:性能逼近 GPT-5.6 Sol 與 Fable 5,中國 AI 價格戰時代終結The Decoder·The Decoder產品Product
- 045輝達 Nemotron 3 Embed 奪 RTEB 榜首,大幅提升 AI 代理檢索效能Hugging Face Blog·Hugging Face Blog產品Product
- 046Thinking Machines Lab 推出 Inkling:975B 多模態開源模型,樹立美國新標竿Latent Space·Latent Space產品Product
- 047語音 AI 評測新標準:Real World VoiceEQ 揭示「人味」差距Hugging Face Blog·Hugging Face Blog研究Research
- 048OpenAI 推出 GPT-5.6 系列模型,ChatGPT 轉型工作超級應用Latent Space·Latent Space產品Product
- 049SpaceXAI 發表 Grok 4.5 模型,馬斯克讚譽為「Opus 等級」TechCrunch AI·TechCrunch AI產品Product
- 050NVIDIA Nemotron 3 Ultra 結合 LangChain Deep Agents 實現領先效能與成本效益NVIDIA Blog·NVIDIA Blog產品Product
- 051AI 程式碼能力評估挑戰:OpenAI 發現基準測試 SWE-Bench Pro 存在缺陷OpenAI Blog·OpenAI Blog研究Research
- 052AI 產業前沿:Fable 5 應用、騰訊 Hy3 開源與 Anthropic J-Space 研究Latent Space·Latent Space其他Other
- 053英國AI安全研究院:現有評測系統性低估AI代理真實能力The Decoder·The Decoder研究Research
- 054ScarfBench:評估 AI 代理在企業級 Java 框架遷移的表現Hugging Face Blog·Hugging Face Blog研究Research
- 055SkillOpt:將 AI 代理技能轉化為可訓練參數,提升可靠性Microsoft Research·Microsoft Research研究Research
- 056Hugging Face 整合 Every Eval Ever 評估結果,提升模型透明度與可信度Hugging Face Blog·Hugging Face Blog產品Product
- 057GeneBench-Pro 深度解析:AI 基因體學評測基準案例研究OpenAI Blog·OpenAI Blog研究Research
- 058OpenAI 推出 GeneBench-Pro:評估 AI 科學判斷力的新基準OpenAI Blog·OpenAI Blog研究Research
- 059Memora:平衡抽象與細節的 AI 記憶系統Microsoft Research·Microsoft Research研究Research
- 060AI 模型擔任新創 CEO 挑戰:500 天模擬測試僅三款獲利The Decoder·The Decoder研究Research
- 061OpenAI 推出 GPT-5.6 Sol:宣稱代理編碼超越 Claude Mythos,並批政府存取規定不可持續The Decoder·The Decoder產品Product
- 062OpenAI 啟動「修補地球」計畫,全面強化開源軟體資安與 AI 防禦能力Wired AI·Wired AI產品Product
- 063AI 模型的網路安全評估模式與基準Eugene Yan·Eugene Yan研究Research
- 064OpenAI 研究揭示:少量「有益特徵」訓練可大幅提升 AI 模型安全性與抗操控性The Decoder·The Decoder研究Research
- 065AI 代理程式夠靈活嗎?Hugging Face 評測開源模型與工具的互動效率Hugging Face Blog·Hugging Face Blog研究Research
- 066GLM-5.2 橫空出世:最強純文字開源大型語言模型問世Simon Willison·Simon Willison產品Product
- 067OpenAI 推出 LifeSciBench:提升 AI 在生命科學研究的評估標準OpenAI Blog·OpenAI Blog產品Product
- 068NVIDIA Blackwell 平台橫掃 MLPerf Training 6.0:最快、最大、最強NVIDIA Blog·NVIDIA Blog產品Product
- 069AI 程式碼代理新研究:能找到正確檔案,卻錯失關鍵程式碼行The Decoder·The Decoder研究Research
- 070NVIDIA Blackwell 平台在首個代理式 AI 基礎設施基準測試中表現卓越NVIDIA Blog·NVIDIA Blog產品Product
- 071Anthropic 推出 Claude Fable 5:性能卓越,但資料政策與開發限制引發爭議Latent Space·Latent Space產品Product
- 072語音助理如何應對雙語客戶?前沿 ASR 模型在語碼轉換語音的基準測試Hugging Face Blog·Hugging Face Blog研究Research
- 073Hugging Face CLI 專為 AI 代理優化:Hub 互動更高效Hugging Face Blog·Hugging Face Blog產品Product
- 074突破非形式化AI極限:Axiom Math的驗證式智慧之路Latent Space·Latent Space研究Research
- 075微軟Build大會:薩蒂亞·納德拉談AI前沿平台與生態系Latent Space·Latent Space商業Business
- 076MiniMax M3 開源模型問世:百萬上下文視窗,效能比肩專有巨頭The Decoder·The Decoder產品Product
- 077AI 科技週報:模型、代理與開源生態系最新進展Latent Space·Latent Space產品Product
- 078ITBench-AA 基準測試揭示:頂尖 AI 模型在企業 IT 代理任務中表現未達五成Hugging Face Blog·Hugging Face Blog研究Research
- 079小米MiMo-V2.5-Pro開源模型,自主編碼能力直逼Claude OpusThe Decoder·The Decoder產品Product
- 080AI模型倫理光譜:相同提示詞,迥異道德抉擇The Decoder·The Decoder研究Research