跳到正文
返回论文列表
cs.CL提交于 已译

LLM4Impact:融合异构信息预测科学影响力

LLM4Impact: Integrating Heterogeneous Information for Scientific Impact Prediction

Yong Cao · Markus Flicke · Haoyu He · Katrin Renz · Andreas Geiger

中文摘要

预测新发表论文的未来影响力具有挑战性,因为需从发表时可获取的异构证据中推断。现有方法通常依赖单一信息源,或在融合多源信息时未考虑其不同的预测作用。LLM4Impact 是一种证据感知的科学影响力预测方法,学习对异构信息进行表征、融合与校准。该方法结合语义、图结构、大语言模型(LLM)与时序四类表征,通过连续前缀词元将图结构信息注入冻结的 LLM。上下文感知门控机制自适应地为不同证据分配权重,独立的校准模块考虑引用规模的领域与时序差异。进一步构建了一个大规模基准数据集,包含 200 万篇论文、防止数据泄露的 point-in-time 异构自媒体图、时序划分,提供年份级与 month-level 引用预测标签。LLM4Impact 显著超越强语义、图结构与 LLM 基线,在分布内测试集上实现 10.13% 的年份 RMSE 降低,在跨域分布下降低 6.87%。结果表明,证据的价值依赖于上下文:不同论文受益于不同信息源,领域与发表时间影响证据到引用的转化。该发现支持自适应证据选择与上下文条件校准,而非仅仅追求更丰富的表征。论文发表后将释放代码、基准以及交互式 Web 演示。

关键要点

  1. 01问题:论文影响力预测依赖异构证据,现有方法多源融合时未区分各源的预测作用
  2. 02方法:LLM4Impact 融合语义、图、LLM、时序四种表征,通过前缀词元注入图信息,并用上下文门控与校准模块自适应加权
  3. 03数据:构建包含 200 万篇论文的大规模基准,提供 point-in-time 自媒体图与时序划分以防止数据泄露
  4. 04结果:在分布内年份 RMSE 降低 10.13%,跨域分布提升 6.87%,优于语义、图结构与 LLM 基线
  5. 05局限:不同论文受益于不同证据源,领域与时间影响引用转化,需上下文自适应而非单一表征堆叠

解读

尚无解读。

原始英文摘要

arXiv:2610.10138v1 Announce Type: new Abstract: Predicting the future impact of a newly published paper is challenging because it must be inferred from heterogeneous evidence available at publication time. Existing approaches often rely on a single source of information or combine multiple sources without accounting for their different predictive roles. In this paper, we present LLM4Impact, an evidence-aware method for scientific impact prediction that learns to represent, integrate, and calibrate heterogeneous information. LLM4Impact combines semantic, graph, LLM, and temporal representations, and injects graph information into a frozen LLM through continuous prefix tokens. A context aware gating mechanism adaptively weights different evidence, while a separate calibration module accounts for domain and temporal variation in citation scales. We further construct a large-scale benchmark dataset with 2 million papers, leakage-safe point-in-time heterogeneous ego graphs, temporal splits, and both year-level and month-level citation targets. Experiments show that LLM4Impact consistently outperforms strong semantic, graph, and LLM based baselines, with a 10.13% reduction in year RMSE on the in distribution test set and a 6.87% reduction under out-of-domain distribution. Our results reveal that the value of such evidence is context dependent: different papers benefit from different sources, while domain and publication time affect how evidence translates into citations. This finding motivates adaptive evidence selection and context-conditioned calibration rather than simply richer representations. We will release our code, benchmark, and an interactive web demonstration upon publication.

同方向论文 · cs.CL

查看全部 →