Album

HuggingFace 每日AI论文速递

10分钟速读热门AI论文

拨号上网 佚名 科技 · 科技
1.46万 订阅 664 集 5天前
播客简介
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。 📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】 🖼另外还有图文版,可在小红书搜索并关注【AI速递】
节目
【月末特辑】7月最火AI论文 | 虎鲸模型预测世界状态;Kimi K3开源逼近顶尖

【月末特辑】7月最火AI论文 | 虎鲸模型预测世界状态;Kimi K3开源逼近顶尖

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 10 篇论文如下: [00:42] TOP1(🔥473) | 🌍 Orca: The World is in Your Mind(虎鲸:世界在你心中) [03:50] TOP2(🔥438) | 🧠 Kimi K3: Open Frontier Intelligence(Kimi K3:开放前沿智能) [06:31] TOP3(🔥309) | 🧩 Program-as-Weights: A Programming Paradigm for Fuzzy Functions(程序即权重:面向模糊函数的编程范式) [09:09] TOP4(🔥307) | 🌍 ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU(ABot-World-0:在单个桌面GPU上实现无限交互式世界展开) [12:34] TOP5(🔥293) | 🧪 AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis(AskChem:以论断为中心的化学文献综合基础设施) [15:21] TOP6(🔥289) | 🤖 Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents(Qwen-UI-Agent技术报告:迈向下一代以真实世界为中心的基础GUI智能体) [18:00] TOP7(🔥259) | 🧠 Metis: Memory Foundation Model(Metis:记忆基础模型) [22:07] TOP8(🔥232) | 🧭 Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable(驾驭手册:使不断演化的智能体驾驭系统可读、可导航且可编辑) [25:09] TOP9(🔥205) | 🧠 LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget(长稻草:在固定GPU预算下实现超过200万Token的长上下文强化学习) [28:24] TOP10(🔥198) | 🤖 RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model(RynnBrain 1.1:迈向更强大和更通用的具身基础模型) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

31分钟
99+
5天前
2026.07.30 | TurboVLA实现消费级显卡实时操控;CoRT精细化信用分配提升指令遵循。

2026.07.30 | TurboVLA实现消费级显卡实时操控;CoRT精细化信用分配提升指令遵循。

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:33] ⚡ TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM(TurboVLA:在RTX 4090上以32Hz频率运行且显存占用低于1GB的实时视觉-语言-动作模型) [01:36] 🎯 CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization(CoRT:用于令牌级准则引导策略优化的反事实重放) [02:31] 🤖 HumanCLAW: Can Vision-Language Models Act Through a Body?(HumanCLAW:视觉-语言模型能否通过身体行动?) [03:41] 🧬 DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space(DecoEvo:文本空间中求解器与评估生成器技能的解耦协同进化) [04:39] 🧠 CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition(CLBench-V:从基础定位到知识获取的多模态上下文学习评估) [05:32] 🧩 CAST: Game Solvers as Turn-Level Teachers for LLM Agents(CAST:游戏求解器作为LLM智能体的回合级教师) [06:28] 🧠 SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution(技能崛起:面向跨任务技能演化的智能体强化学习) [07:18] 🎮 StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation(StatePlay:状态感知的游戏世界模型用于机制一致的内容生成) [08:05] 📊 OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding(OmegaUse-OfficeVal:基于经济基准评估LLM智能体在长期办公套件任务中的表现) [09:08] 🤖 Can AI agents conduct open-ended AI research? Early evidence from two case studies(AI智能体能否进行开放式的AI研究?来自两个案例研究的早期证据) [10:04] 🛡 GPT-Red: Automated Red Teaming via Self-Play at Scale(GPT-Red:通过大规模自我对弈实现自动化红队测试) [11:04] 🎬 Explicit Layer Modeling for Video Object Insertion and Layer Decomposition(显式层建模用于视频对象插入与层分解) [12:01] 🕵 StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents(StealthBench:衡量自主攻击安全代理的操作隐蔽性) [12:54] 🧠 Memory for Large Language Models(大型语言模型的记忆机制) [13:48] 📜 Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems(为叙述者评级:面向多智能体知识系统中声明级溯源的一种伊斯纳德-里贾尔框架) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

15分钟
73
1周前
2026.07.29 | 高保真数据训练策略,逼近真机效果;相关性动态引导搜索,精准高效检索

2026.07.29 | 高保真数据训练策略,逼近真机效果;相关性动态引导搜索,精准高效检索

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:31] 🤖 HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone(HiFi-UMI:仅从高保真UMI数据学习可部署的操作策略) [01:31] 🔍 A New Role for Relevance: Guiding Corpus Interaction in Agentic Search(相关性的新角色:在智能体搜索中引导语料库交互) [02:20] 🎨 ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition(ReDesign:通过智能体分解从图像中恢复可编辑的设计结构) [03:07] 🧠 Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory(保持铭记:基准测试代理记忆中的内隐关联盲点) [03:56] 🏃 Pass the Baton: Trajectory-Relayed On-Policy Distillation(传递接力棒:轨迹中继的在线策略蒸馏) [04:47] ⚡ Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model(Mage-VL:一种高效的编解码器原生流式多模态基础模型) [05:49] 🧩 CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents(CodeNib:一种为编码智能体提供仓库上下文服务的多视图数据系统) [06:48] 🌍 Wonder: Video World Model Done Better(Wonder:更优的视频世界模型) [07:51] 👁 PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models(感知基准:评估多模态大语言模型中的原子视觉感知能力) [08:44] 🔍 Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking(新颖主张还是似曾相识?重新思考多模态自动事实核查的“无污染”动态评估) [09:40] 🛡 Shieldstral(盾星) [10:41] 🔀 MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities(MODUS:仅解码器的任意模态到任意模态多样化建模) [11:41] ⚡ Parallel Decoding Distillation for Fast Image and Video Generation(并行解码蒸馏:面向快速图像与视频生成的方法) [12:23] 🎬 Visual prompt engineering for video models(视频模型的视觉提示工程) [13:16] 🎯 OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs(OmniDelta:面向全模态大语言模型令牌压缩的技能驱动预算分配方法) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
80
1周前
评价

空空如也

加入我们的 Discord

与播客爱好者一起交流

立即加入

扫描微信二维码

添加微信好友,获取更多播客资讯

微信二维码

播放列表

自动播放下一个

播放列表还是空的

去找些喜欢的节目添加进来吧