HuggingFace 每日AI论文速递 - 节目列表

2026.09.02 | 学生模拟器实现因材施教;自动驾驶模型融合感知规划。

2026.09.02 | 学生模拟器实现因材施教;自动驾驶模型融合感知规划。

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:28] 🎓 StudentSim: Training LLM-based Student Simulators(StudentSim:训练基于大语言模型的学生模拟器) [01:34] 🚗 Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving(Qwen-Drive-1.0:迈向自动驾驶视觉语言基础模型的第一步) [02:46] 🔁 SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers(SMELT:计算匹配的MoE循环Transformer的缩放定律) [03:44] 🤖 UI-Venus-2 Technical Report(UI-Venus-2技术报告) [04:41] 🤖 ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training(ZimaBlue:通过可扩展视频预训练演化可泛化的世界动作模型) [05:40] 🌍 H3-World: Turning Language Understanding into World Control(H3-World:将语言理解转化为世界控制) [06:41] 🔍 Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering(Hi-Q:层次化证据引导的多跳问答查询细化) [07:36] 🚁 Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching(评估多模态大语言模型作为无人机控制的通用视觉-语言-行动智能体:指挥、接近、跟踪与搜索) [08:33] 🔗 Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System(揭示原生统一多模态模型中理解与生成的协同效应:从表征、任务到系统) [09:27] 🛡 Safin-1: Safety from Within through Memory-Native State Evolution(Safin-1:通过记忆原生状态演化实现由内而外的安全) [10:24] 🏢 From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix(从生产流量到后训练:构建覆盖企业请求组合的自托管大语言模型) [11:15] 🩺 DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory(DiagEvo:通过分层错误记忆进行诊断引导的自我进化) [12:07] 🧠 EM^2Mem: Event-Centric Multimodal Memory for Large Language Models(EM²Mem:面向大语言模型的事件中心多模态记忆) [13:09] 🤖 Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement(Harness-of-Harness:多日持续改进的自主软件开发) [14:07] 🛡 Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs(控制流与数据流分离:多智能体大语言模型中的稳定提示优化) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

15分钟
83
2周前
2026.09.01 | 在线策略蒸馏实为自我改进;原生2K音视频联合生成

2026.09.01 | 在线策略蒸馏实为自我改进;原生2K音视频联合生成

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:30] 🔄 Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement(在线策略蒸馏真的在蒸馏吗?从噪声教师到自我改进) [01:24] 🎬 DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution(DreamX-Creator:在2K分辨率下实现原生音视频生成的民主化) [02:13] 🧠 GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling(GenFirst:先生成后重建的稳定端到端潜在生成建模) [03:06] 🧩 Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling(Lucida:面向可组合真实到仿真场景建模的解析、生成与放置) [03:57] ⚖ Normalized Low-Rank Adaptation(归一化低秩自适应) [04:54] 🎯 PaperGym: Rubric-Centered Evolution for Research-Plan Generation(PaperGym:以评分标准为中心的研究计划生成进化) [05:53] ⚡ On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability(论Qwen3.8-Next架构的设计:评估、效率与训练稳定性) [06:51] 🎓 CogEvol: Towards Efficient and Reliable Learning Environment Generation(CogEvol:迈向高效可靠的学习环境生成) [07:50] 🧭 LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation(LightNav-0:激发视觉语言模型空间智能,实现通用具身导航) [08:43] 🧩 SHAPE of Chain-of-Thought in Math Reasoning(数学推理中思维链的SHAPE框架) [09:48] 🧠 Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence(超越人类监督的大规模推理模型扩展:通往超级智能之路) [10:51] 🧩 Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase(超级库智能体:超越单一代码库的多应用联合生成与维护) [11:48] 🤖 Evaluating the Hidden Costs of Personalization in Large Language Models(评估大型语言模型中个性化的隐藏成本) [12:42] 📋 Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents(先评估后改进:面向自动科研智能体的自动评分标准归纳) [13:32] 🎭 Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions(看得见的谎言:具身社会交互中VLM智能体的言语与非言语联合欺骗) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
87
3周前
2026.08.31 | LoopArena揭示循环控制难题;DART-SD拓扑感知优化工具调用。

2026.08.31 | LoopArena揭示循环控制难题;DART-SD拓扑感知优化工具调用。

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:29] 🔁 LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering(LoopArena:将模型作为循环工程的运行时控制器进行基准测试) [01:27] 💎 DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents(DART-SD:多轮工具调用代理的菱形拓扑感知检索与自蒸馏调优) [02:25] 🛠 Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities(智能体工件创建:系统、评估、原则与机遇) [03:25] 🤖 Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models(超越数据规模:面向视觉-语言-动作模型的以表征为中心的持续预训练) [04:42] 🌍 Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning(代码即世界:面向物理推理的可执行世界表示的智能体发现) [05:45] 🔄 J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data(J-Zero:零数据下挑战者—求解者—评判者的统一协同进化) [06:37] 🎥 Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction(重新审视面向长时程流式三维重建的局部上下文) [07:35] 🧠 ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL(ContextPilot:通过细粒度强化学习训练智能体进行主动上下文管理) [08:33] 🧠 LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation(LayerRecall:用于视频生成长时程一致性的状态条件记忆路由器) [09:38] 💰 Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090(Puro-2B:穷实验室在RTX 5090上以不到5090美元训练出的Qwen2-1.5B) [10:27] 🐘 Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge(盲人摸象:探究长尾分歧知识下大语言模型的认识论短视) [11:18] 🛡 StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing(StepGuard:通过可扩展监督与安全效用平衡学习步骤级护栏) [12:12] 🎨 Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents(画你所见:多模态智能体中灵巧视觉工具使用的基准测试) [13:08] ⚡ Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding(在视频中定位一切:重新思考高效生成式时空视频定位) [14:04] 🧠 PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control(PonderPounce:预训练多模态大语言模型作为机器人控制的回合上下文引擎) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

15分钟
97
3周前
2026.08.28 | 视频模型概率校准不足;智能体数据需ACE平衡

2026.08.28 | 视频模型概率校准不足;智能体数据需ACE平衡

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:30] 🎲 PAWBench: How Far Are We from Probabilistically Aligned World Modeling?(PAWBench:我们距离概率对齐的世界建模还有多远?) [01:09] 🤖 What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents(什么造就了好的智能体数据?面向LLM智能体的数据生成的ACE视角) [02:01] 🎯 TTPO: Test-Time Policy Optimization(TTPO:测试时策略优化) [02:49] 🤖 Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report(训练智能体与其运行框架协同进化:TaoLive数字人智能体技术报告) [03:41] 🏙 UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City(UrbanGround:从局部感知到真实尺度城市中的空间能动性) [04:34] 🎮 GameWAM: A World Action Model for Video Games(GameWAM:视频游戏的世界动作模型) [05:24] 🔄 PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents(PILOT在环:面向长时程智能体的实时自我改进) [06:03] 🤖 Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization(Zero-WAM:基于人类视频的上下文世界-动作建模实现开放式任务泛化) [06:55] ⚗ Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher(Self-OPD:无需教师模型的流匹配模型在线策略蒸馏) [07:51] 🧬 Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO(理解面向LLM推理的进化策略:比GRPO更广泛的推理覆盖) [08:44] 🎮 Magpie: Real-Time World Renderer for Interactive Games(Magpie:用于交互式游戏的实时世界渲染器) [09:41] 🧩 Procedura: Agentic 3D Modeling with Procedural Control(Procedura:程序化控制下的智能体3D建模) [10:39] 🎬 Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning(思考镜头:基于智能体推理的一致性多镜头视频编辑) [11:40] 🧠 WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution(WikiSkill:将智能体经验编译为持久知识以促进技能进化) [12:46] 🔗 CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval(CaSKG:面向可扩展智能体技能检索的反事实因果技能图) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
68
3周前
2026.08.27 | 双脑记忆让语音助手实时又准确;科学工作流多数模型难善终

2026.08.27 | 双脑记忆让语音助手实时又准确;科学工作流多数模型难善终

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:31] 🧠 VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction(VoiceMem:面向实时交互的流式双脑记忆) [01:30] 🧪 FrontierChallenge: Evaluating Scientific Workflow Completion(前沿挑战:评估科学工作流完成度) [02:29] 🎬 VGI-Bench: Probing Visual Intelligence in Video Generation Models(VGI-Bench:探究视频生成模型的视觉智能) [03:32] 🚀 WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation(WarpSAC:重思探索与利用,迈向可扩展离策略强化学习的巅峰) [04:27] ⚙ JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution(JIT-Agent:通过即时框架演化实现框架智能的规模化扩展) [05:20] 🧠 VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning(VBVR-Pro:可扩展且可验证的原生视觉推理套件) [06:15] 🎛 D$^3$-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation(D³-MOPD:面向高效多教师蒸馏的自适应动态领域调度) [07:22] 🧠 Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data(下一块推理强化学习真的比SFT更好吗?重新审视无CoT数据下的训练策略) [08:26] 🎯 Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning(Agent-G²:面向智能体强化学习的高斯引导) [09:13] 🎬 Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds(面向持久故事与交互世界的长时程音视频生成) [10:19] ⚖ Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation(Open-MOPD:多教师同策略蒸馏中能力失衡的诊断与修复) [11:26] 🤖 StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models(StreamPI:面向视觉-语言-动作模型的流式多模态时序建模) [12:29] 🎬 Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios(视频-IFBench:评估多模态大语言模型在视频理解场景中的指令跟随能力) [13:30] 🧠 Code World Model: Coding Agent as World Brain(代码世界模型:代码智能体作为世界大脑) [14:25] 🎯 V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning(V-Rubrics:基于评分标准的强化学习实现视觉忠实性) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

15分钟
82
3周前
2026.08.26 | 将标注视为轨迹以加速视频强化学习;微信多模态嵌入模型刷新纪录

2026.08.26 | 将标注视为轨迹以加速视频强化学习;微信多模态嵌入模型刷新纪录

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🎯 Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs(将标注视为轨迹:高效可扩展的视频多模态大语言模型强化学习) [01:25] 🧲 WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report(WeMM-Embedding:微信多模态嵌入技术报告) [02:17] 🔧 AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces(AutoSaddler:基于智能体执行轨迹的自动框架优化与持久更新) [03:17] 🎯 On-Policy Self-Distillation in Diffusion Models(扩散模型中的同策略自蒸馏) [04:02] 🛡 CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild(CyberFactory:利用现实世界实例规模化网络安全能力) [05:00] 🧠 Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses(面向长时程智能体框架的递归经验-工作记忆演化) [05:53] 🎯 Best Practice Critic Optimization(最佳实践评论家优化) [06:45] 🎯 On-policy Distillation with Verifiable Reward(基于可验证奖励的在线策略蒸馏) [07:42] 🕶 From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms(从看见到行动:智能眼镜作为第一人称智能平台) [08:41] 🎮 Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training(Game2World引擎:解锁真实场景游戏视频以训练世界模型) [09:34] 📏 Length-Adaptive Decoding for Masked Diffusion Machine Translation(掩码扩散机器翻译的长度自适应解码) [10:27] 🎬 LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training(LAION-BVD:面向多模态预训练的千万小时开放视频数据集) [11:21] 🔁 Meta$^n$: Recursive Self-Improvement through Emergent Depth(Meta^n:通过涌现深度实现递归自我改进) [12:13] 🤖 CAFE: Self-Improving Search Agents Need Co-Evolving Feedback(CAFE:自我改进的搜索智能体需要协同演化的反馈) [13:14] 🔮 Latent Action as Intention Enables Efficient Future Imagination for World Action Models(以潜在动作作为意图,实现世界动作模型的高效未来想象) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
70
3周前
2026.08.25 | 智能体从会答到会交付;全模态世界开放可进入

2026.08.25 | 智能体从会答到会交付;全模态世界开放可进入

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:30] 🤖 Apodex 1.1: Scaling Agentic Intelligence for Complex Work(Apodex 1.1:扩展智能体智能以应对复杂工作) [01:36] 🎮 EchoWM: Open and Enterable Omnimodal World Models(EchoWM:开放且可进入的全模态世界模型) [02:40] 🛒 TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming(TLive-Omni:面向电商直播的全模态理解模型) [03:22] 🎨 Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision(通过概念缩放与密集监督解锁图像编辑的潜力) [04:20] 📱 MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks(MobilePA-Bench:面向复杂真实世界任务的移动规划智能体基准评测) [05:16] 🤖 Prime Agent: A Self-Improving RLM Harness(Prime Agent:一种自我改进的递归语言模型工具框架) [06:22] 🧊 Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion(Block3D:通过块状扩散的高效文本到三维生成) [07:09] 🚗 RISE: Adaptive Imagination for World Action Models(RISE:面向世界行动模型的自适应想象) [08:19] 📈 Towards a Densing Law for User Representation Learning at Billion-Scale Capacity(面向十亿级容量用户表示学习的稠密化定律) [09:26] ⚖ ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction(ARC:开放式真实世界交互中的公平相对优势比较) [10:27] 🧠 ReWorld: An Interactive World Model with Long-Horizon Memory(ReWorld:一个具有长时记忆的交互式世界模型) [11:29] ⚖ Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization(超越稳定性-探索困境:面向大语言模型策略优化的环境正则化) [12:23] 🎮 GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?(GameXpert-Bench:编码智能体距离专家级游戏开发还有多远?) [13:17] 🧪 One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows(一次成功不等于可靠:Thinkingbox——面向有状态业务工作流的智能体沙盒与基准测试) [14:26] 🎯 Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection(Task-CoEvolve:通过自适应验证任务选择实现高效的工具链优化) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

15分钟
99+
4周前
2026.08.24 | 大模型超参迁移降本;图工程引领系统协作

2026.08.24 | 大模型超参迁移降本;图工程引领系统协作

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:32] 🚀 Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts(让我们一步一步扩展规模:面向大规模混合专家模型的高效超参数迁移) [01:24] 🕸 Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence(大语言模型智能体时代的图工程:从个体智能到系统智能) [02:21] 🤖 OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs(OmniAssistBench:全模态大语言模型的助手式交互基准) [03:17] ⚡ ParaTempo: Efficient Parallel Reasoning via Temporal Confidence(ParaTempo:基于时间置信度的高效并行推理) [04:00] ♾ InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter(InfinityEdit:基于轻量级编辑触发适配器的无限视频编辑) [04:53] 🪙 Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models(每一枚硬币都有两面:论大型语言模型同策略蒸馏中泛化的双重性) [05:37] 🧩 EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking(EviRank:面向多模态图像重排序的结构化相关性证据) [06:46] 🧠 Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs(超越正确性:混合思考多模态大语言模型的响应行为基准测试与对齐) [07:45] 🎨 UniSpace: Unified Visual Representation and Scalable Multimodal Modeling(UniSpace:统一视觉表示与可扩展多模态建模) [08:39] 🛒 Towards Faithful Simulation of Human Shopping Behavior(面向人类购物行为的忠实模拟) [09:35] ⚙ AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale(AgentMercury:你的智能体能够大规模合成可验证的商业场景环境) [10:41] ⚡ Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference(Daedalus-150M:为CPU推理设计的卷积-注意力混合模型) [11:35] 🛡 CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment(CLEAR:面向保持实用性的大语言模型安全对齐的连续潜在适配器路由) [12:42] 📱 Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs(Llama-Mobile:高效2.7比特视觉语言模型量化) [13:32] 🍳 FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth(FlavourBench:用可执行的烹饪真值对前沿语言模型进行排名) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
90
4周前
【周末特辑】8月第4周最火AI论文 | 操控系统升级让AI更准;视频检测难敌生成攻击

【周末特辑】8月第4周最火AI论文 | 操控系统升级让AI更准;视频检测难敌生成攻击

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 5 篇论文如下: [00:43] TOP1(🔥435) | ⚙ StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling(StateM:通过 Harness 扩展在 Terminal-Bench 2.1 上达到 95.3% 原始准确率,或一次 15 美元的前沿运行) [03:30] TOP2(🔥277) | 🛡 Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination(我们能防御针对现实世界危机事件的AI生成视频攻击吗?对检测器、生成器与社会传播的系统评估) [06:28] TOP3(🔥245) | 🌍 EnvHarness: Awakening Static Worlds for Agent Learning(EnvHarness:为智能体学习唤醒静态世界) [09:45] TOP4(🔥167) | 👁 Self-Supervised Visual On-Policy Distillation(自监督视觉同策略蒸馏) [12:21] TOP5(🔥162) | 🧩 Demystifying Agent Skills: Why They Work-Until They Don't(揭秘智能体技能:它们为何有效——直到失效) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

15分钟
96
1个月前
2026.08.21 | 动态环境适配弱点;任务合成保留源意图

2026.08.21 | 动态环境适配弱点;任务合成保留源意图

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:31] 🌍 EnvHarness: Awakening Static Worlds for Agent Learning(EnvHarness:为智能体学习唤醒静态世界) [01:28] 🖥 FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis(FACET:在终端任务合成中保留源意图与可执行状态) [02:29] 🧪 SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?(SWE-bench Science:编码智能体能否解决科学领域的工程任务?) [03:15] 🕺 4DAnyone: Create Anyone in 4D from a Casual Monocular Video(4DAnyone:从随意单目视频创建任意人物的4D形象) [04:11] 👥 WithEveryone: Unified Planning and Identity Grounding for Group Image Generation(WithEveryone:面向群像生成的统一规划与身份锚定) [05:01] 🧠 MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use(MemTrapBench:大语言模型记忆使用中的认知陷阱基准测试) [05:52] 🔄 SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback(SkillEvo:来自多轮交互反馈的自我更新进化梯度) [06:50] 🎮 ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models(ForgeWM:面向少步动作条件视频世界模型的渐进式因果训练) [07:48] 🧩 Repo0: Design-Driven Zero-to-All Code Generation(Repo0:设计驱动的从零到全代码生成) [08:46] ⚡ FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving(FlashPrefill V2:面向长上下文大语言模型服务的块稀疏预填充注意力) [09:41] 🧠 Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization(注入、对齐、恢复:面向无检索文档知识内化的分阶段后训练) [10:42] 🤖 EXIMO: VLM Guided Exploration of VLA Policies(EXIMO:视觉语言模型引导的VLA策略探索) [11:36] 🧠 Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See(用低资源语言思考:SFT构建了什么,RL修复了什么,准确率无法看到什么) [12:18] 🎯 Towards Quantifying Benchmark Optimization in ASR Models(面向ASR模型中基准优化的量化研究) [13:20] 🛡 PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents(PolicyGuide:从守护单一动作到引导策略合规型LLM智能体的整个工作流) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
45
1个月前
2026.08.20 | 闭环进化提升具身智能;验证门控保障工业代码

2026.08.20 | 闭环进化提升具身智能;验证门控保障工业代码

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:34] 🤖 Zetta $ζ$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence(Zetta ζ:面向自进化物理智能的高效闭环具身智能体框架) [01:29] ✅ SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation(SemaPLC:一种基于项目、以验证为门控的PLC代码生成智能体框架) [02:31] 🎯 SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation(SemComp-Bench:视频生成中语义任务完成的基准测试) [03:28] 🔬 OmniScientist: An Omni-Modal Omni-Discipline AI Scientist(全能科学家:一个全模态、全学科的人工智能科学家) [04:26] 🧠 Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL(Co-RL:多智能体强化学习中多样化群体催生无监督推理) [05:13] 🎮 SPADE: Self-Play in Adaptive Synthetic Executable Environments(SPADE:自适应合成可执行环境中的自博弈) [06:08] 🧪 Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis(训练面向单步逆合成的化学合理性感知大语言模型) [07:04] 🎯 Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning(潜在世界模型中的决策度量对齐:诊断与面向MPC规划的动作条件目标) [07:58] 🧬 Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification(训练留下痕迹:面向语言模型谱系验证的中心化残差签名) [08:50] 🔁 Looped Language Models Improve Compositional Tool Calling(循环语言模型提升组合式工具调用能力) [09:36] 🖐 SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation(SoftVTBench:面向可变形物体操作的变形感知视触觉数据集与基准) [10:33] ⚽ FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents(FM-Bench:面向竞争智能体的长时程管理基准) [11:25] ✍ Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion(借助属性引导的体裁扩展,将创意写作扩展到故事中心数据之外) [12:15] 🔥 The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning(越热门越难遗忘:大语言模型遗忘的自适应流行度方法) [13:01] 🔍 Temporal Multi-Signal Fusion for Token-Level Hallucination Detection(面向Token级幻觉检测的时序多信号融合) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
82
1个月前

加入我们的 Discord

与播客爱好者一起交流

立即加入

扫描微信二维码

添加微信好友,获取更多播客资讯

微信二维码

播放列表

自动播放下一个

播放列表还是空的

去找些喜欢的节目添加进来吧