HuggingFace 每日AI论文速递 - 节目列表

2026.07.16 | 行为定位与高效多模态;开源模型追赶闭源系统

2026.07.16 | 行为定位与高效多模态;开源模型追赶闭源系统

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:33] 🧭 Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable(驾驭手册:使不断演化的智能体驾驭系统可读、可导航且可编辑) [01:19] 🖼 Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation(Boogu-Image-0.1:推动开源统一多模态理解与生成) [02:07] 🧠 Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning(环零:将零强化学习扩展至万亿参数以实现涌现推理) [03:08] 📄 OvisOCR2 Technical Report(OvisOCR2 技术报告) [04:12] 🤖 KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill(知深行准-GUIClaw:具备自进化记忆与技能的深度认知、精准执行个人GUI助手) [05:10] 🛡 PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails(政策转变守卫:基准测试与改进策略自适应图像防护机制) [06:02] 🤖 GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch(GigaWorld-Policy-0.5:一种由自动研究驱动的更快更强的世界动作模型) [06:57] 🖼 MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors(MetaView:具有尺度感知隐式几何先验的单目新视角合成) [07:46] ✂ ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation(ShortOPD:通过短到长在线策略蒸馏恢复剪枝后的大语言模型) [08:40] 🎭 Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation(Hallo4D:多模态幻觉缓解实现一致的时空生成) [09:28] 🔬 Registers Matter for Pixel-Space Diffusion Transformers(寄存器对像素空间扩散Transformer至关重要) [10:22] 🤖 Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos(Vinci2:在连续自我中心视频中提供主动辅助) [11:21] 🔍 Tracing Agentic Failure from the Flow of Success(从成功流程中追溯智能体失败) [12:20] 🔍 From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization(从噪声轨迹到根因:面向智能体优化的结构化轨迹分析与因果提取) [13:23] 🧭 AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities(AgentCompass:一种统一的智能体能力评估基础设施) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
99+
3周前
2026.07.15 | SpectraReward:零样本多模态奖励模型;盲点基准:揭示AI的认知盲区

2026.07.15 | SpectraReward:零样本多模态奖励模型;盲点基准:揭示AI的认知盲区

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 10 篇论文如下: [00:31] 🔄 Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation(重新读它:预训练多模态大语言模型是文本到图像生成的零样本奖励模型) [01:35] 🔍 Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models(盲点基准:评估多模态模型中的盲点) [02:30] 📄 SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding(SynthDocBench:面向长上下文视觉文档理解的受控基准) [03:31] 🔍 Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution(先知晓再修复:面向软件问题修复的基于问答的仓库知识获取) [04:20] 🎵 MuScriptor: An Open Model for Multi-Instrument Music Transcription(MuScriptor:面向多乐器音乐转录的开放模型) [05:12] 🔍 Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation(超越可教授的知识边界:在智能体视觉生成中演化知识边界) [06:00] 🎨 Let RGB Be the Language of Vision(让RGB成为视觉的语言) [06:50] 📄 MonkeyOCRv2: A Visual-Text Foundation Model for Document AI(MonkeyOCRv2:面向文档AI的视觉-文本基础模型) [07:50] 🤖 Towards Autonomous and Auditable Medical Imaging Model Development(迈向自主且可审计的医学影像模型开发) [08:40] 🧠 Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms(深度强化学习评估与设计范式的原则性分析) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

9分钟
87
3周前
2026.07.14 | 弱到强泛化提升大模型;双系统架构导航新突破

2026.07.14 | 弱到强泛化提升大模型;双系统架构导航新突破

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 14 篇论文如下: [00:32] 🧠 Weak-to-Strong Generalization via Direct On-Policy Distillation(弱到强泛化:通过直接在线策略蒸馏) [01:22] 🤖 ABot-N1: Toward a General Visual Language Navigation Foundation Model(ABot-N1:迈向通用视觉语言导航基础模型) [02:20] 🤖 ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory(ABot-AgentOS:一种具有终身多模态记忆的通用机器人智能体操作系统) [03:21] 🕺 4D Human-Scene Reconstruction from Low-Overlap Captures(低重叠度捕获下的四维人体场景重建) [04:13] 🧠 LightMem-Ego: Your AI Memory for Everyday Life(轻量记忆自我:你的日常生活AI记忆) [05:15] 🧮 AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification(高级数学推理基准:面向高级数学证明生成与验证的基准套件) [06:11] 🧠 Metacognition in LLMs: Foundations, Progress, and Opportunities(大语言模型中的元认知:基础、进展与机遇) [07:03] 🤖 EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos(EgoSteer:面向从自我中心视频实现可操控灵巧操作的全栈系统) [07:55] 🧩 Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals(代理探索与可复用引导:一种通过代理引导更新信号实现的模块化大语言模型后训练范式) [08:55] 🧠 NeuroCogMap Reveals Cognitive Organization of Large Language Models(神经认知图谱揭示大型语言模型的认知组织) [09:52] 🎬 Motion4Motion: Motion Transfer Across Subjects at Inference(Motion4Motion:推理时跨主体的运动迁移) [10:58] 👗 CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation(CtrlVTON:通过视觉实例提示分割实现可控虚拟试穿) [11:54] 🎭 Latent-Identity Tuning in Text-to-Image Personalization Models(文本到图像个性化模型中的潜在身份调优) [12:52] 🧊 LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow(LATO.2:基于顶点与拓扑流的分解式3D网格生成) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
83
3周前
2026.07.13 | LHTB基准揭示AI长时任务瓶颈;视觉预训练超越文本预训练

2026.07.13 | LHTB基准揭示AI长时任务瓶颈;视觉预训练超越文本预训练

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 13 篇论文如下: [00:33] 🚀 Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading(长时终端基准:在具有密集奖励评分的长期终端任务中测试智能体的极限) [01:25] 👁 Scalable Visual Pretraining for Language Intelligence(面向语言智能的可扩展视觉预训练) [02:22] 🎥 Video Generation Models are General-Purpose Vision Learners(视频生成模型是通用视觉学习器) [03:16] 🎯 Trust Region Policy Distillation(信任区域策略蒸馏) [04:14] 🧠 KronQ: LLM Quantization via Kronecker-Factored Hessian(KronQ:通过Kronecker分解黑塞矩阵实现的大语言模型量化) [05:00] 🌍 PanoWorld: Real-World Panoramic Generation(PanoWorld:真实世界全景生成) [05:56] 🎨 From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models(从RGB生成到密集场读取:利用文本到图像模型进行像素空间密集预测) [06:49] 🎯 Self-Guided Test-Time Training for Long-Context LLMs(自引导测试时训练用于长上下文大语言模型) [07:36] 🚗 Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation(Flow-ERD:面向多样化交通仿真的智能体类型感知流匹配与熵正则化蒸馏) [08:28] 🔊 Phone Segmentation and Recognition through Phonological Activation Mapping(通过音位激活映射进行音素分割与识别) [09:28] 🧠 Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning(探究大语言模型微调中记忆知识为何无法泛化的机制性理解) [10:18] 🤖 A Sovereign, Open-Source Foundation Model for German and English(一个面向德语和英语的主权开源基础模型) [11:13] 🏛 VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery(VaseMuseum:面向古希腊陶器的数字智能博物馆) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

12分钟
99+
3周前
2026.07.10 | Vidu S1实现实时视频对话交互;视频绿洲揭示理解评估缺陷

2026.07.10 | Vidu S1实现实时视频对话交互;视频绿洲揭示理解评估缺陷

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:31] 🎬 Vidu S1: A Real-Time Interactive Video Generation Model(Vidu S1:一种实时交互式视频生成模型) [01:29] 🔍 Video-Oasis: Rethinking Evaluation of Video Understanding(视频绿洲:重新思考视频理解的评估) [02:27] 🔑 Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition(为什么我打不开抽屉?缓解零样本组合动作识别中的对象驱动捷径) [03:21] 🤖 UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks(UniClawBench:面向真实世界任务中主动智能体的通用基准) [04:18] 🎥 LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models(LongE2V:基于视频扩散模型的长时域事件驱动视频重建、预测与帧插值) [05:24] 🧬 Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation(思想具有基因组:科学谱系推理与基于谱系的创意生成基准测试) [06:13] 🌐 Enhancing In-context Panoramic Generation via Geometric-aware Pretraining(通过几何感知预训练增强上下文全景生成) [06:54] 🎬 CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation(CineMobile:面向电影级摄像机运动生成的设备端图像到视频扩散) [07:52] 🎬 OpenCoF: Learning to Reason Through Video Generation(OpenCoF:通过视频生成学习推理) [08:45] 🚀 Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE(Jet-Long:基于动态双焦点旋转位置编码的高效长上下文扩展) [09:39] ⚡ Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing(线性注意力架构:机制、权衡与跨层路由) [10:38] 💊 DrugGen 2: A disease-aware language model for enhancing drug discovery(DrugGen 2:一种疾病感知型语言模型,用于增强药物发现) [11:23] ⚡ Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models(Flash-BoN:扩散模型中推理时缩放的即时草稿生成) [12:08] 🎵 A Quantized Native Runtime for On-Device Semantic Audio Generation(面向设备端语义音频生成的量化原生运行时) [13:01] 🚀 UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma(UP:用于打破探索-稳定性困境的无界正向非对称优化) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
73
4周前
2026.07.09 | 结构可推理,记忆可长存;双突破让AI更懂科学

2026.07.09 | 结构可推理,记忆可长存;双突破让AI更懂科学

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 7 篇论文如下: [00:30] 🔬 Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning(基于深度原生结构推理的精确、跨学科且透明的构效关系理解) [01:28] 🤖 Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation(视觉-语言-动作模型中的双潜记忆用于机器人操作) [02:34] 🤖 Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence(面向具身智能的混合专家视频预训练规模化) [03:28] 🌍 Infinite Worlds with Versatile Interactions(无限世界与多样化交互) [04:32] 🤖 RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies(RoboDojo:统一仿真与真实世界的通用机器人操作策略综合评估基准) [05:34] 🏙 WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence(WildCity:一个用于渲染、仿真和空间智能的真实世界城市规模测试平台) [06:35] 🧩 Teaching LLMs a Low-Resource Language: Enhancing Code Completion in Pharo(教会大语言模型一种低资源语言:提升Pharo语言中的代码补全能力) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

7分钟
99+
4周前
2026.07.08 | AlayaWorld实时生成长视频;RynnWorld-4D提升机器人操作成功率

2026.07.08 | AlayaWorld实时生成长视频;RynnWorld-4D提升机器人操作成功率

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:30] 🎮 AlayaWorld: Long-Horizon and Playable Video World Generation(AlayaWorld:长视界且可游玩的视频世界生成) [01:31] 🤖 RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation(RynnWorld-4D:面向机器人操作的4D具身世界模型) [02:30] 🔍 Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling(分层稀疏注意力机制的正确实现:迈向无限上下文建模) [03:36] 👁 Vision as Unified Multimodal Generation(视觉作为统一多模态生成) [04:25] 🔦 Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory(光之全能:基于长期记忆的智能体视频理解中反思优于推理) [05:25] 🎥 Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning(并行化自回归解码用于全模态密集视频字幕生成) [06:14] 🧠 Gemma 4 Technical Report(Gemma 4 技术报告) [07:11] ⚡ SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe(SkillOpt-Lite:通过一行“氛围”实现更优更快的智能体自我进化) [08:07] ⚡ DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation(DSpark:基于置信度调度的半自回归推测解码) [09:03] 🧠 MentalThink: Shaping Thoughts in Mental SVG World(MentalThink:在思维SVG世界中塑造思想) [09:57] 🤖 RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation(RynnWorld-Teleop:一种用于数字遥操作的动作条件世界模型) [10:58] 🎯 TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training(TurnOPD:使在线知识蒸馏具备回合感知能力以实现高效的长时程智能体训练) [11:58] 🎨 CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration(CanvasAgent:通过可视化工具编排实现复杂图像创建与编辑) [13:05] 🧪 TREK: Distill to Explore, Reinforce to Refine(TREK:通过蒸馏进行探索,通过强化进行精炼) [13:51] 🖼 PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation(PointDiT:面向单目几何估计的像素空间扩散模型) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

15分钟
92
4周前
2026.07.07 | 跨平台智能体学习新范式;科研构思可复用技能提炼

2026.07.07 | 跨平台智能体学习新范式;科研构思可复用技能提炼

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 14 篇论文如下: [00:32] 🤖 UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning(UI-MOPD:面向持续GUI智能体学习的多平台在线策略蒸馏) [01:30] 💡 ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes(ResearchStudio-Idea:基于证据的科研构思技能套件——来自机器学习会议成果) [02:21] 🎨 PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space(PixWorld:在像素空间中统一3D场景生成与重建) [03:11] 🧩 OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers(OmniOpt:现代优化器的分类、几何结构与基准测试) [04:04] 🤖 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation(GigaWorld-1:构建用于机器人策略评估的世界模型路线图) [04:55] 🧩 Vision Pretraining for Dense Spatial Perception(面向密集空间感知的视觉预训练) [05:54] 🤖 EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots(EVA-Client:面向实体机器人上的具身策略的统一数据收集、推理与部署框架) [06:47] 🎥 Wan-Streamer v0.2: Higher Resolution, Same Latency(Wan-Streamer v0.2:更高分辨率,相同延迟) [07:48] 🤖 InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization(InternVLA-A1.5:统一理解、潜在预知与动作以实现组合泛化) [08:46] 🔍 Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval(所有视觉标记都同等重要吗?面向视觉-语言检索的保留对象证据的标记合并方法) [09:59] 🧠 KVpop -- Key-Value Cache Compression with Predictive Online Pruning(KVpop——基于预测性在线剪枝的键值缓存压缩) [10:50] 🧠 dOPSD: On-Policy Self-Distillation for Diffusion Language Models(dOPSD:扩散语言模型的在线自蒸馏方法) [11:41] 🎨 Perceptual Flow Matching for Few-Step Generative Modeling(感知流匹配:用于少步生成建模) [12:37] 🔬 Multi-Turn Agentic Scientific Literature Search via Workflow Induction(多轮交互式科学文献搜索:基于工作流归纳的方法) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
99+
1个月前
2026.07.06 | 策略对齐破解推理盲区;轻量外挂实现实时修正

2026.07.06 | 策略对齐破解推理盲区;轻量外挂实现实时修正

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 9 篇论文如下: [00:35] 🎯 The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning(优化训练策略的幻影:单调推理策略作为大语言模型强化学习的真正目标) [01:34] 🛠 VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon(VLA-Corrector:一种用于自适应动作视界的轻量级检测与校正推理框架) [02:35] 🤖 Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots(Embodied.cpp:面向异构机器人的具身AI模型便携推理运行时) [03:43] 🔢 OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers(OrbitQuant:面向图像和视频扩散Transformer的数据无关量化方法) [04:38] 📊 DataComp-VLM: Improved Open Datasets for Vision-Language Models(DataComp-VLM:改进的视觉语言模型开放数据集) [05:36] 🛡 Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming(保护AI智能体:一种多层次智能体红队测试的统一框架) [06:36] ☁ Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment(面向解译的云去除:基于观测锚定残差流与地理上下文对齐的方法) [07:38] 📄 MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering(多注意力归因:长文档问答中的无训练多模态归因) [08:25] 🔗 AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation(AGE:面向图检索增强生成的自适应掩码图嵌入方法) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

9分钟
90
1个月前
2026.07.03 | 小模型本地化击败大模型;自主策略演化聚焦结构合成

2026.07.03 | 小模型本地化击败大模型;自主策略演化聚焦结构合成

HuggingFace 每日AI论文速递

【赞助商】 OpenClaw快报 每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论 传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43 【目录】 本期的 15 篇论文如下: [00:31] 🧩 Program-as-Weights: A Programming Paradigm for Fuzzy Functions(程序即权重:面向模糊函数的编程范式) [01:24] 🧠 EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments(EvoPolicyGym:在交互环境中评估自主策略演化) [02:24] 🧠 AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents(AgenticSTS:面向长时程LLM智能体的有界内存测试平台) [03:18] 🔍 Morphing into Hybrid Attention Models(变形为混合注意力模型) [04:07] 📊 AgenticDataBench: A Comprehensive Benchmark for Data Agents(AgenticDataBench:面向数据智能体的综合性基准测试) [05:12] ⚡ Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling(多分辨率流匹配:通过分阶段采样的无训练扩散加速) [05:52] 🎬 WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory(世界导演:构建具有持久动态记忆的可控世界模拟器) [06:49] 🏥 Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning(打破失败级联:面向医学多模态推理的步骤感知强化学习) [07:37] 🎨 Optimizing Visual Generative Models via Distribution-wise Rewards(通过分布级奖励优化视觉生成模型) [08:32] 🎯 SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use(SkillCoach:用于评估和增强智能体技能使用的自我演化评分标准) [09:21] 🖐 AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition(AGVBench:面向静脉识别的可靠性导向数据增强基准) [10:16] 🔬 From SRA to Self-Flow: Data Augmentation or Self-Supervision?(从SRA到Self-Flow:数据增强还是自监督?) [11:10] 🧠 Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads(对数几率贡献评分识别非字面检索头) [11:58] 🎯 AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models(AnyGroundBench:面向视觉语言模型中视频定位的专业领域基准) [12:46] 🤔 When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search(搜索代理何时应提问:面向澄清感知的深度搜索基准DiscoBench) 【关注我们】 您还可以在以下平台找到我们,获得播客内容以外更多信息 小红书: AI速递

14分钟
99+
1个月前

加入我们的 Discord

与播客爱好者一起交流

立即加入

扫描微信二维码

添加微信好友,获取更多播客资讯

微信二维码

播放列表

自动播放下一个

播放列表还是空的

去找些喜欢的节目添加进来吧