HuggingFace 每日AI论文速递 - 节目列表

2026.09.17 | 科学代码库转智能体环境;上下文机制网络赋能结构化数据智能

2026.09.17 | 科学代码库转智能体环境;上下文机制网络赋能结构化数据智能

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:31] 🧪 ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments(ScienceIDE:将全球科学代码库转化为智能体可学习环境) [01:23] 🧠 LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence(LimiX-2:面向通用结构化数据智能的上下文机制网络) [02:10] 📉 Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening(重新思考PPO中的评论家学习:理解与缓解价值平坦化) [02:56] 🧠 Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents(置信度源于经验:从推理到智能体的经验性置信度估计) [03:36] 🤖 ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks(ProgramDistill:从交互式 Web 应用到可验证的参考引导软件工程任务) [04:21] 🤖 ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models(ActionPiece:重新思考自回归视觉-语言-动作模型的动作标记化) [05:01] 🧠 Agora: Git as Shared Memory for Collective AutoResearch(Agora:将 Git 作为集体自动研究的共享记忆) [05:43] ⚡ VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention(VC-Attention:面向低比特注意力的值平滑与 Softmax 转换) [06:22] 📈 EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents(EvolveTrade:面向自演化LLM交易智能体的经验驱动策略精炼) [07:02] 🎮 Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control(Zing-0.5:迈向具备实时联合动作与文本控制的可玩世界) [07:51] 👀 Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX(注视作为共同基础证据:MapTask与MUNDEX的跨语料库分析) [08:35] 🎯 A Zeroth-Order Paradigm for LLM Preference Alignment(面向LLM偏好对齐的零阶范式) [09:17] 🧬 HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses(HypoEvolve:遗传算法使多智能体大语言模型能够发现科学假设) [10:02] ✋ EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset(EventEgoHands++:基于事件的第一人称3D手部网格重建与真实数据集) [10:55] 🔭 SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization(SpectralShift:通过谱重参数化有效扩展 Gated DeltaNet 的上下文窗口) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

12分钟
26
5天前
2026.09.16 | 持续学习组合提升保留率;游戏AI六类角色待整合

2026.09.16 | 持续学习组合提升保留率;游戏AI六类角色待整合

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:29] 🧠 Continual Learning Mechanisms Compose for Long-Horizon Memorization(持续学习机制组合用于长时程记忆) [01:17] 🎮 AI for Games in the Foundation Model Era(基础模型时代的游戏人工智能) [02:05] 🎙 StepAudio 3 Realtime Technical Report(StepAudio 3 Realtime 技术报告) [02:45] 🎵 StepAudio 3 Music Technical Report(StepAudio 3 Music 技术报告) [03:22] 🤖 ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents(ScienceBuddy:面向交互式科学智能体的递归嵌套式自我改进) [04:09] 🧭 HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness(HarnessVLN:通过智能体框架统一免训练具身导航) [04:52] 🤖 The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement(人类建造的最后一个AI:迈向真正的递归自我改进) [05:40] 🏗 Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?(墙里的另一张蓝图:如何像孩子一样向前沿AI提问?) [06:26] 🧠 Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States(Mind2Dialogue:通过模拟用户心理状态训练人类感知语言模型) [07:10] 🧩 ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement(ModularRSI:模块化且可泛化的递归式执行框架自我改进) [07:54] 🤖 Modality-Autoregressive World-Action Models(模态自回归世界动作模型) [08:38] 📐 Disentangling Representation Evolution in Transformers through Directional Decomposition(通过方向分解解耦 Transformer 中的表征演化) [09:25] 🎥 PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control(PhysStream:具备结构化场景记忆与细粒度运动控制的流式物理基础视频生成) [10:06] 🧪 ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals(ImpossibleRubrics:对作为奖励信号生成的评分标准进行压力测试) [10:55] 🧠 Convergent Emergence of In-Context Learning Across Modalities(跨模态上下文学习的趋同涌现) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

12分钟
56
6天前
2026.09.15 | 智能体超级智能黎明;实时可编辑空间视频生成

2026.09.15 | 智能体超级智能黎明;实时可编辑空间视频生成

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:30] 🤖 Atria Dawn: The Dawn of Agentic Superintelligence(Atria Dawn:智能体超级智能的黎明) [01:13] 🎬 Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation(Vidu S2:实时交互、可编辑与空间视频生成) [01:58] 🤖 ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search(ZGCM-1:一个完全开放且极其高效、面向数学与智能体搜索的基础模型) [02:47] 🧠 Dream-RSI: Recursive Self-Improvement through Evolving Worlds(Dream-RSI:通过演化世界实现递归自我改进) [03:24] 🤖 PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models(PhysBrain 1.5:从视觉语言模型到物理基础模型) [04:06] 🗜 Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction(分组值注意力:通过按需键重建实现高效KV缓存) [04:51] 🎬 LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows(LynnReal-Omni:面向智能体视觉工作流的原生多模态视频生成) [05:45] 🤖 RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments(RSIAgent:新环境中面向递归自我改进的自主探索) [06:30] 🔬 Discovery Foundation Models: Toward Open-Ended Discovery Intelligence(发现基础模型:迈向开放式发现智能) [07:13] 🎬 BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender(BVB:在 Blender 中通过程序化重建对智能体视频理解进行基准测试) [07:56] ⚖ How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus(无损投机解码究竟有多无损?数值精度在 Orthrus 中的作用) [08:39] 🌐 AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video(AlayaVista:从全景状态到透视视频的流式世界建模) [09:26] 🧩 Kaininja: Extending Native 3D Generators to the Part Level(Kaininja:将原生3D生成器扩展至部件级) [10:13] 🧠 Omni-Streaming Thinking(全模态流式思维) [11:00] 🖥 LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents(LLaDA-UI:将块级扩散引入视觉语言 GUI 智能体) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

12分钟
87
1周前
2026.09.11 | 下一概念预测提效;原生统一视觉智能

2026.09.11 | 下一概念预测提效;原生统一视觉智能

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:28] 🧠 NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction(NCP-ArchPreview 技术报告:通过下一概念预测迈向潜在空间语言模型) [01:13] 🧠 SenseNova-U1.5: Towards Native Unified Visual Intelligence(SenseNova-U1.5:迈向原生统一视觉智能) [02:09] 🧱 SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem(SpatialBlock:通过合成积木堆叠问题增强大型视觉语言模型的空间智能) [02:53] 🛡 EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents(EvoSafeHarness:为保障智能体安全而演化模型与领域特定的安全护栏) [03:39] 🧹 Mi-Ripple: Restoring Images Degraded by Iterative AI Editing(Mi-Ripple:修复由迭代式 AI 编辑退化的图像) [04:23] 🗜 X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation(X-AuT:基于跨尺度蒸馏的语音大语言模型渐进式音频编码器压缩) [05:11] 🧠 Memory as Plans: World-Action Modeling with Memory-Grounded Planning(记忆即计划:基于记忆锚定规划的世界-动作建模) [05:54] 🧩 TempCloze: Can Video-LLMs Identify the Missing Middle?(TempCloze:视频大语言模型能否识别缺失的中间片段?) [06:41] 🌊 FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation(FreeFlow:一种用于光流估计的无偏置层次化Transformer) [07:34] 🧩 Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs(递归代码世界模型:通过递归场景程序构建复杂世界) [08:21] 🫀 CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation(CARDEA:基于空间证据的可审计推理,用于端到端冠状动脉造影解读) [09:07] 🚫 Negative Self-Distillation: Learning to Reason by Avoiding Flaws(负向自蒸馏:通过规避缺陷学习推理) [09:53] ✈ DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat(DRG-MAPPO:面向协同空战的分层动态角色图多智能体强化学习) [10:35] 🩻 UniH$^3$: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration(UniH³:统一分层同质性与异质性的一体化医学图像恢复) [11:22] 🎥 World in World: Explore the World with World Models(世界中的世界:用世界模型探索世界) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

12分钟
97
1周前
2026.09.10 | VLM语义动作玩转机器人;可编程世界模型状态可控

2026.09.10 | VLM语义动作玩转机器人;可编程世界模型状态可控

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:28] 🤖 Show-Harness: Just a VLM Agent Can Play Robots(Show-Harness:仅一个VLM智能体即可玩转机器人) [01:11] 🎮 Programmable World Model(可编程世界模型) [01:57] 🤖 AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems(AgentGrad:干预引导的多智能体系统提示优化) [02:40] ⌚ WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data(WearableQA:面向真实世界可穿戴数据的健康推理基准) [03:27] ✅ SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents(SWE-Bench Pro Verified:面向软件工程智能体的可靠基准) [04:11] 🔬 SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?(SAEScientist-Bench:AI智能体能否开展自主SAE可解释性研究?) [04:55] 🔍 Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents(仅凭分数不能证明发现:用于审计 AI 研究代理的发现认证协议) [05:42] 🗣 Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation(Puppeteer:基于物体锚定与姿态感知的共语手势生成) [06:26] ⚗ DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents(DianShi-RxnDB:面向研究人员与AI智能体的、通过全自动流程构建的大规模细粒度有机反应数据平台) [07:09] 🤖 SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators(SyncWorld:视觉校准使世界模型成为零样本模拟器) [07:54] 🏗 $Φ$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?(Φ-Bench:大型语言模型能否工程化支撑自身的基础设施?) [08:39] 🧠 Revisiting Complete Reasoning Traces for Post-Training(重新审视后训练中的完整推理轨迹) [09:31] 🔄 Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning(训练更智能,而非更费力:主动学习中的切换信号引导训练) [10:19] 🎥 Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs(为什么视频仍然如此昂贵?视频与音视频大语言模型中的推理效率机制综述) [11:07] 🤖 Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails(协同演化智能体运行框架与模型:同策略修正帮助较弱模型在模仿失败之处追赶) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

12分钟
94
1周前
2026.09.08 | 离散扩散无损加速;流平衡优化自改进

2026.09.08 | 离散扩散无损加速;流平衡优化自改进

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 7 篇论文如下: [00:31] ⚡ Unlocking Lossless Speedups in LLMs via Discrete Diffusion(通过离散扩散实现大语言模型的无损加速) [01:31] ⚖ FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience(流平衡:基于验证器的在线策略推理经验自我改进方法) [02:29] 🎯 ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation(ENEAS:嵌入引导的神经集成自适应分割) [03:27] 🤖 EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents(具身技能(EmbodiedSkills):编排、训练与部署VLA智能体的统一框架) [04:32] 📉 One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation(一种症状,三个杠杆:同策略自蒸馏的批判性综述) [05:15] 🚦 Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation(先验证再蒸馏:在策略蒸馏中的提示级教师门控) [06:00] 🔧 What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation(还有哪些需要修复?探索经对话生成的产物中修订传播的高性价比测试时计算) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

7分钟
70
2周前
2026.09.07 | 多智能体协作需外部验证;搜索模型应区分自身能力与上下文增益

2026.09.07 | 多智能体协作需外部验证;搜索模型应区分自身能力与上下文增益

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🤖 Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems(双层协调反思:多智能体LLM系统的博弈论方法) [01:27] 🧗 Iris: Climbing to the Search Frontier(Iris:攀登搜索前沿) [02:12] 🤖 Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue(Motion-Omni:面向口语对话的端到端语音与全身动作联合生成) [03:06] 🌍 WorldSculpt: Generating Compositional Worlds from Grounded Videos(WorldSculpt:从锚定视频生成组合式世界) [03:57] 🔍 Enoki: Efficient Multi-Level Hallucination Detection(Enoki:高效的多层级幻觉检测) [04:56] ❓ Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization(优化前先询问:面向交互式优化的动态预建模澄清) [06:04] 🔺 The Attention Triangle in Audio-Video Models(音频-视频模型中的注意力三角) [07:03] ⚡ Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference(不要丢弃Dropout:优化层稀疏性以实现高效的LLM训练与推理) [07:57] 🧠 Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs(思维链的表面之下:大语言模型中推理操作的机制性解读) [08:51] 🔄 RISE: Recursive Improvement via Self-Extrapolating Policy Distillation(RISE:基于自外推策略蒸馏的递归改进) [09:52] 🤖 MaxKernel: Agentic Kernel Generation for TPUs(MaxKernel:面向TPU的智能体内核生成) [10:41] ✂ When Models Edit Too Much: On the Fidelity of Minimal Code Edits(当模型修改过度:关于最小代码编辑的保真度) [11:35] ✂ Group Adaptive Clipping Policy Optimization(分组自适应裁剪策略优化) [12:25] 🧊 Training-Free Speech-Centric Omni Understanding with Frozen VLMs(基于冻结视觉语言模型的免训练、以语音为中心的全模态理解) [13:23] 🤖 $τ^τ$-Bench: An Environment for End-To-End, Realistic Agent Construction(τ^τ-Bench:用于端到端、真实智能体构建的环境) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
95
2周前
【月末特辑】8月最火AI论文 | 循环潜在推理兼顾上下文学习;执行框架扩展提升智能体表现

【月末特辑】8月最火AI论文 | 循环潜在推理兼顾上下文学习;执行框架扩展提升智能体表现

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 10 篇论文如下: [00:41] TOP1(🔥772) | 🧠 BDH-CQ: In-Context Learning with Recurrent Latent Reasoning(BDH-CQ:基于循环潜在推理的上下文学习) [04:16] TOP2(🔥446) | ⚙ StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling(StateM:通过 Harness 扩展在 Terminal-Bench 2.1 上达到 95.3% 原始准确率,或一次 15 美元的前沿运行) [07:08] TOP3(🔥342) | 🔄 Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA(Macaron-V1:迈向具备自我改进和LoRA混合的开放持续学习) [09:53] TOP4(🔥340) | 🕵 HarnessEval-W: Agentifying the Evaluation of Visual Worlds(HarnessEval-W:使视觉世界的评估智能体化) [12:45] TOP5(🔥290) | 📄 Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill(从火花到论文:作为可组合技能的端到端研究论文生成) [16:40] TOP6(🔥282) | 🛡 Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination(我们能防御针对现实世界危机事件的AI生成视频攻击吗?对检测器、生成器与社会传播的系统评估) [19:31] TOP7(🔥274) | 🌍 EnvHarness: Awakening Static Worlds for Agent Learning(EnvHarness:为智能体学习唤醒静态世界) [22:34] TOP8(🔥269) | 🧠 VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning(VBVR-Pro:可扩展且可验证的原生视觉推理套件) [26:15] TOP9(🔥263) | 🧬 OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution(OpenART:通过开放式环境演化扩展智能体红队测试) [29:25] TOP10(🔥252) | 🔁 Recursive Synthesis for Long-Horizon Terminal Tasks(面向长时程终端任务的递归式合成) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

32分钟
97
2周前
2026.09.04 | 终端环境可由轨迹重建;模型后训练需审慎迁移

2026.09.04 | 终端环境可由轨迹重建;模型后训练需审慎迁移

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:33] 🌌 Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments(终端宇宙:将智能体轨迹转化为可扩展的终端环境) [01:31] 🚦 Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training(知道何时不该复用:自主大语言模型后训练中的条件经验迁移) [02:34] 🎨 LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes(LLaDA-Image:以完全开放的训练配方构建强大图像生成器) [03:33] 🧠 LatentPress: Context Compression Beyond Text and Vision(LatentPress:超越文本与视觉的上下文压缩) [04:20] 🧠 Compile by Training: Turning Natural-Language Specifications into Local Neural Functions(通过训练进行编译:将自然语言规范转化为本地神经函数) [05:12] 🎲 Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning(随机注意力:重新思考面向高效推理的KV缓存淘汰) [06:19] ⚡ Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM(为什么门控DeltaNet能在4比特量化下存活:面向混合27B大语言模型循环半部分的NVFP4 W4A4量化) [07:15] 🎯 Rethinking On-Policy Distillation of Large Language Models II: One Training Example(重新思考大语言模型的同策略蒸馏(二):一个训练示例) [08:18] 🌍 Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States(Puffin-World:以原生3D世界状态扩展统一多模态模型) [09:17] 🧩 Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction(Scal3R:学习高效多相对位姿查询以实现可扩展的在线三维重建) [10:09] 🎨 Editable Visual Design(可编辑视觉设计) [11:08] ⏱ The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation(缺失的时间链接:用于脚本驱动的音视频生成的时间上下文路由) [11:58] 🧠 Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding(超越检索:用于流式视频理解的渐进式潜在记忆演化) [12:50] 🧩 CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation(CORE:通过重排序器蒸馏提升多模态大语言模型嵌入的组合推理) [13:47] 🤖 RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests(RealSWE:真实用户请求下编码智能体的组合式评估) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

15分钟
54
2周前
2026.09.03 | 代码库蒸馏成AI技能;长视频世界模型可扩展

2026.09.03 | 代码库蒸馏成AI技能;长视频世界模型可扩展

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:31] 🤖 Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills(从仓库到技能:将GitHub代码库蒸馏为AI4AI技能) [01:26] 🌍 SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models(SolarWM:面向长时程视频世界模型的开放数据与可扩展训练) [02:19] ⏱ EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction(EarlyEval:通过早期结果预测实现更廉价的智能体评估) [03:21] 🤝 It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning(两相匹配:基于强化学习的生成式检索器协同进化) [04:18] 🎯 Language Models Can Control Their Own Attention(语言模型能够控制自身的注意力) [05:10] 🖼 On the Design Fundamentals of Pixel Text Representation Learning(论像素文本表示学习的设计基础) [06:12] 🔍 Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering(超越视觉相似性:面向知识型视觉问答的实体对齐检索) [07:14] 🤖 HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?(HarnessDev:大语言模型能否创建并演化自己的智能体执行框架?) [08:05] 🗜 ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes(ZipTok3D:利用紧凑令牌前缀的高保真三维令牌化) [09:04] 🎯 Cliff: Learning Process Rewards from the First Mistake(Cliff:从第一个错误中学习过程奖励) [10:05] 🎯 Aspire: Can Models Self-Evolve from Vague Goals?(Aspire:模型能否从模糊目标中自我进化?) [10:46] 🎭 Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation(影响导向蒸馏:解决采样词元在线策略蒸馏中的多样性瓶颈) [11:42] 🎮 S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?(S3Gym:大语言模型能否将自我测试与自我评判转化为自我改进?) [12:36] 👀 A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss(一瞥足矣:基于SimLoss的单遍细粒度图像描述) [13:35] 🎙 VibeVoice-ASR-Streaming Technical Report(VibeVoice-ASR-Streaming技术报告) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
99+
2周前

加入我们的 Discord

与播客爱好者一起交流

立即加入

扫描微信二维码

添加微信好友,获取更多播客资讯

微信二维码

播放列表

自动播放下一个

播放列表还是空的

去找些喜欢的节目添加进来吧