前沿情报日报 · 2026-09-23

飘飘飘
发布于 2026-09-23 / 1 阅读
0

前沿情报日报 · 2026-09-23

每日自动抓取:arXiv 当日新论文、近半年高被引论文、Hacker News 热门。仅供学习参考。

arXiv 今日新论文(AI / 机器学习 / 量化金融)

1. GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Yiran Wang, Xingyilang Yin, Junfu Pu 等 · 2026-09-21

Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing …

2. Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use
Zixiang Chen, Wenting Zhao, Zhepeng Cen 等 · 2026-09-21

Multi-turn tool-use failures can hinge on a single model call, yet reward variation alone does not reveal which call would benefit from training. When rewards depend on later interactions, their variation can reflect …

3. WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
Wangbo Yu, Kunhao Liu, Wenbo Hu 等 · 2026-09-21

Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a …

4. onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction
Lei Yang, Mengyin Liu, Jia Wang 等 · 2026-09-21

We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its core interaction: while reading a model response, the annotator …

5. DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation
Haoran Yuan, Zekai Wang, Boning Shao 等 · 2026-09-21

Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain …

6. Harness-Zero: Harness Distillation via Agent-as-Harness
Haoran Ye, Yuxing Lu, Haonan Dong 等 · 2026-09-21

Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies …

7. RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
Peng Xia, Rujun Han, Zifeng Wang 等 · 2026-09-21

An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this …

8. DolphinBench: Mapping the Pareto Frontier of Agent Memory
Soumil Rathi, Deshraj Yadav, Taranjeet Singh 等 · 2026-09-21

Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question …

9. Rare Event Estimation via Iterative Unalignment
Hanming Yang, Daksh Mittal, Jing Dong 等 · 2026-09-21

As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic. Safe deployment therefore does not depend on whether these events …

10. Emergent Collusion in Long-Horizon LLM Agent Interaction
Xinrui Shi, Yanzhe Zhang, Diyi Yang 等 · 2026-09-21

LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two …

11. Jev for Scientific Decisions: Evaluating Semantic Choices and Their Consequences
Boyuan Deng, Shuyi Fan, Hongyang Zhang 等 · 2026-09-21

Scientific workflows often require choosing among known relations before a deterministic calculation can proceed. Whether observations share a culture, treatment or reference standard can change the scientific meaning …

12. Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks
Muzhe Wu, Zuchen Li, Xu Wang 等 · 2026-09-21

Visual instructions for physical tasks are typically authored in one context and followed in another, requiring users to translate demonstrated tools, materials, and spatial relationships into their own environment. We …

13. Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization
Filipe Marinho Rocha, Inês Dutra, Vítor Santos Costa 等 · 2026-09-21

A model generalizes outside its training distribution only when it computes a representation structurally equivalent to the generating mechanism, not an approximation fitted to it. Such equivalence is necessary for …

14. Linguistic Features for Interpretable Textual Entailment
David Torres-Moreno, Jorge Hermosillo-Valadez, Asela Reig-Alamillo 等 · 2026-09-21

Despite the success of neural models in natural language processing, their black-box nature limits interpretability and conceals the linguistic phenomena underlying their predictions. We present SLITE, an explainable …

15. Et Tu, Brute? Economic Misalignment in Personal AI Agents
Aman Priyanshu, Supriti Vijay, Brian Jabarian 等 · 2026-09-21

Personal AI agents make recommendations and take actions on people's behalf in high-stakes economic contexts, e.g., buying a flight, choosing health insurance, or selecting a graduate program. The agent is given access …

共 15 篇。

近半年高被引论文

AI / 机器学习方向

论文 被引
A Survey of Large Language Models 1555 次
Polyendocrine metabolic ovarian syndrome, the new name for polycystic ovary syndrome: a multistep global consensus process 232 次
The MetroVolt Data-Center Burner: Direct-DC Campus Power, Q_E Closure, and the Plug Requirement in a Low-Neutron D–³He Tandem Mirror 100 次
Accelerating scientific discovery with Co-Scientist 95 次
Chatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts Architecture 94 次

量化金融方向

论文 被引
A Survey of Large Language Models 1555 次
Exploring Large Language Model‐Based Intelligent Agents: Definitions, Methods, and Prospects 38 次
Large Models for Time Series and Spatio-Temporal Data: A Survey and Outlook 35 次
The Scenario Model Intercomparison Project for CMIP7 (ScenarioMIP-CMIP7) 32 次
Investigating the replicability of the social and behavioural sciences 32 次

Hacker News 热门(科技 / 创业 / 投资风向)

热度 讨论 标题
1069 赞 464 评 MiMo v2.6
729 赞 583 评 I said no and Apple said yes
615 赞 544 评 Claude Opus 5.5
433 赞 334 评 Apple has added persistent 'ads' to iOS, and it's driving users crazy
418 赞 221 评 GPT-6 Sol and Luna
410 赞 310 评 OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005
349 赞 130 评 Can gzip be a language model?
237 赞 121 评 I asked Meta’s Muse for its filesystem and it sent me 6.8GB
220 赞 160 评 AMD's random number generator can't generate a 0?
185 赞 137 评 OpenAI is well positioned to fast-follow Jev
130 赞 51 评 There's a high chance of devices being sold with GrapheneOS preinstalled in 2027
122 赞 41 评 Show HN: Drop – A rootless Linux sandbox with gVisor support
116 赞 25 评 MUNI Heritage Weekend in San Francisco
96 赞 34 评 Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
96 赞 22 评 Solitaire Alone Together

本日报由服务器每日自动抓取生成于凌晨,原文链接均已附在上。

💬 聊天