主播
节目简介
来源:小宇宙
今天,我们将一起“拆开”AI的大脑,看看做决策的竟然只有8个“员工”?我们还会揭秘AI排行榜的“偏科”陷阱,并为它请来一位完美的“虚拟陪练”和一位全能“秘书”。最后,再用一个简单又奇妙的几何学秘密,看穿AI决策的本质。让我们一同进入AI内部,一探究竟!
00:00:23 AI做决策,到底需要多少“人”帮忙?
00:05:38 AI排行榜的秘密,为什么第一名可能不是你想要的全才?
00:11:01 给AI请个“陪练”,它就能开窍?
00:16:09 给高手配个秘书,怎样才能让他越用越顺手?
00:21:46 AI决策的“保守”秘密
本期介绍的几篇论文:
[CL] Through the Looking Glass: Directly Reading and Writing Transformers
[University of Washington]
https://arxiv.org/abs/2609.10210
---
[CL] What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores
[Stanford University]
https://arxiv.org/abs/2609.09372
---
[LG] World-Time Compute with Verified Code World Models
[Quome, Inc.]
https://arxiv.org/abs/2609.09163
---
[CL] Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding
[Together AI]
https://arxiv.org/abs/2609.09338
---
[LG] Exact-Form Regret for Gradient Descent, Mirror Descent and Follow-the-Regularized-Leader
[MIT]
https://arxiv.org/abs/2609.09466
00:00:23 AI做决策,到底需要多少“人”帮忙?
00:05:38 AI排行榜的秘密,为什么第一名可能不是你想要的全才?
00:11:01 给AI请个“陪练”,它就能开窍?
00:16:09 给高手配个秘书,怎样才能让他越用越顺手?
00:21:46 AI决策的“保守”秘密
本期介绍的几篇论文:
[CL] Through the Looking Glass: Directly Reading and Writing Transformers
[University of Washington]
https://arxiv.org/abs/2609.10210
---
[CL] What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores
[Stanford University]
https://arxiv.org/abs/2609.09372
---
[LG] World-Time Compute with Verified Code World Models
[Quome, Inc.]
https://arxiv.org/abs/2609.09163
---
[CL] Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding
[Together AI]
https://arxiv.org/abs/2609.09338
---
[LG] Exact-Form Regret for Gradient Descent, Mirror Descent and Follow-the-Regularized-Leader
[MIT]
https://arxiv.org/abs/2609.09466