前沿情报日报 · 2026-09-17

飘飘飘
发布于 2026-09-17 / 3 阅读
0

前沿情报日报 · 2026-09-17

每日自动抓取:arXiv 当日新论文、近半年高被引论文、Hacker News 热门。仅供学习参考。

arXiv 今日新论文(AI / 机器学习 / 量化金融)

1. Objective vs. Search: Decomposing What Makes a Good Tokeniser
Ahmetcan Yavuz, Clara Meister, Tiago Pimentel 等 · 2026-09-16

Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM. These differ along two orthogonal axes: their optimisation objective (compression vs. log-likelihood) and …

2. A Zeroth-Order Paradigm for LLM Preference Alignment
Peter Chen, Xi Chen, Wotao Yin 等 · 2026-09-16

Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates …

3. PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
Sara Pieri, Evangelos Kazakos, Shizhe Chen 等 · 2026-09-16

Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but …

4. Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation
Guanhua Ji, Tianyu Li, Dayoon Suh 等 · 2026-09-16

Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories that lack force information, causing failures in …

5. ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
Hejia Geng, Zesen Huang, Haoyang Li 等 · 2026-09-16

Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge …

6. Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments
João Meneses dos Santos, Arlindo L. Oliveira 等 · 2026-09-16

Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps. We extend SwiftSage, a dual-process agent that …

7. Affora: A Design System for Agent-Friendly Interfaces
Jin Gao 等 · 2026-09-16

Computer-use agents increasingly operate software designed for people, but interfaces often leave actions or task state unclear to machine readers. We present Affora, a design system that supports both readers while …

8. Flag Game: A Toy Model for Mechanistic Swarm Interpretability
Elizabeth Pavlova, Hidenori Tanaka 等 · 2026-09-16

Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs about the world, and mechanistic …

9. Playing log(N)-Questions over Wikipedia Abstracts: Communication Efficiency Between Paired Frontier Models
Peter Potash 等 · 2026-09-16

We evaluate six frontier language models on the two-agent $\log(N)$-Questions game. A questioner sees $N$ Wikipedia lead paragraphs and must identify a secretly chosen target using exactly $\log_2 N$ yes/no questions. …

10. rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference
Kaijun Zhou, Zhiyang Li, Le Chen 等 · 2026-09-16

Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs tractable for current policies. …

11. Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
Leon Bergen, Usha Bhalla, Andrew Lee 等 · 2026-09-16

As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented …

12. Prepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeria
Abbas M. Rabiu, Abdulrazaq A. Zubair, Um-mulkhairi Ibrahim 等 · 2026-09-16

Artificial intelligence (AI) is increasingly integrated into healthcare systems worldwide, yet its successful clinical adoption depends critically on workforce readiness, particularly in low- and middle-income countries …

13. Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation
Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny 等 · 2026-09-16

Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose superficially distinct reports, varying in …

14. Securing quantum error correction against misleading advice from AI agents
A. Barış Özgüler 等 · 2026-09-16

Can an attacker turn influence over an artificial intelligence (AI) adviser into a harmful quantum error-correction update? We identify an ambiguity in passive syndrome records that obstructs recovery selection, then …

15. MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
Luyao Zhu, Xun Wei Yee, Wei Li 等 · 2026-09-16

Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated. In AI-assisted language learning, models must …

共 15 篇。

近半年高被引论文

AI / 机器学习方向

论文 被引
A Survey of Large Language Models 1543 次
Polyendocrine metabolic ovarian syndrome, the new name for polycystic ovary syndrome: a multistep global consensus process 220 次
Sycophantic AI decreases prosocial intentions and promotes dependence 117 次
The MetroVolt Data-Center Burner: Direct-DC Campus Power, Q_E Closure, and the Plug Requirement in a Low-Neutron D–³He Tandem Mirror 100 次
Chatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts Architecture 94 次

量化金融方向

论文 被引
A Survey of Large Language Models 1543 次
Exploring Large Language Model‐Based Intelligent Agents: Definitions, Methods, and Prospects 38 次
Large Models for Time Series and Spatio-Temporal Data: A Survey and Outlook 35 次
The Scenario Model Intercomparison Project for CMIP7 (ScenarioMIP-CMIP7) 32 次
Investigating the replicability of the social and behavioural sciences 30 次

Hacker News 热门(科技 / 创业 / 投资风向)

热度 讨论 标题
2170 赞 243 评 Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
705 赞 287 评 Nvidia announces native GPU programming in Rust
559 赞 118 评 Training a 4B model to produce 81% faster query plans than Postgres
528 赞 239 评 Small programming tricks
440 赞 116 评 Xiaomi Mimo 2.6 live post-training dashboard
410 赞 342 评 AWS says it can't restore some data from mideast facilities struck by Iran
295 赞 64 评 Performance Improvements in .NET 11
245 赞 147 评 Backups Aren't Simple
211 赞 34 评 Reversing Factorio's RNG
209 赞 79 评 The engineering behind the US Strategic Petroleum Reserve
209 赞 33 评 Breaking the 1.58-bit Barrier for Ternary LLMs
194 赞 79 评 Japan's book scene is moving from bookstores to libraries
178 赞 65 评 Keys Not Included: recovering the signing keys for US driver's license barcodes
152 赞 234 评 Anecdotally, programmers dislike "reduce"
139 赞 63 评 OpenSpec – A lightweight and configurable AI spec framework

本日报由服务器每日自动抓取生成于凌晨,原文链接均已附在上。

💬 聊天