主播
节目简介
来源:小宇宙
今天,我们不聊堆算力的“大力出奇迹”,而是要探索几条让AI变得更智慧、更可靠的巧妙路径。我们会看到,一个简单的“反刍”机制,如何让AI不再健忘;一场内部“辩论赛”,又如何教会它诚实;“专家分工”的智慧,怎样让它在变强的同时还更省钱。最后,我们还会探究AI是如何学会像高手一样“抬头看路”地做决策,甚至在数学领域领悟“功夫在题外”的道理。准备好了吗?让我们一起揭开这些最新论文背后的绝妙构思。
00:00:37 AI的“反刍”,一个让它更聪明的简单魔法
00:05:10 如何让AI变得更聪明,同时还不变坏?
00:09:50 AI 进化新思路,从“大力出奇迹”到“聪明分工”
00:14:55 高手决策的秘密,既要埋头拉车,又要抬头看路
00:20:36 AI做数学,功夫在诗外
本期介绍的几篇论文:
[LG] Recirculation
[Google DeepMind]
https://arxiv.org/abs/2608.17981
---
[LG] Debate Training Reduces Reward Hacking in RLAIF
[Google DeepMind]
https://arxiv.org/abs/2608.17776
---
[CV] MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
[Meta]
https://arxiv.org/abs/2608.17402
---
[LG] Q-Learning With World Models
[Stanford University & Peking University]
https://arxiv.org/abs/2608.17163
---
[AI] The Problem Is the Problem: Towards Scalable Mathematical Discovery
[CMU]
https://arxiv.org/abs/2608.16977
00:00:37 AI的“反刍”,一个让它更聪明的简单魔法
00:05:10 如何让AI变得更聪明,同时还不变坏?
00:09:50 AI 进化新思路,从“大力出奇迹”到“聪明分工”
00:14:55 高手决策的秘密,既要埋头拉车,又要抬头看路
00:20:36 AI做数学,功夫在诗外
本期介绍的几篇论文:
[LG] Recirculation
[Google DeepMind]
https://arxiv.org/abs/2608.17981
---
[LG] Debate Training Reduces Reward Hacking in RLAIF
[Google DeepMind]
https://arxiv.org/abs/2608.17776
---
[CV] MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
[Meta]
https://arxiv.org/abs/2608.17402
---
[LG] Q-Learning With World Models
[Stanford University & Peking University]
https://arxiv.org/abs/2608.17163
---
[AI] The Problem Is the Problem: Towards Scalable Mathematical Discovery
[CMU]
https://arxiv.org/abs/2608.16977