跳到正文
返回论文列表
cs.LG提交于 已译

两级 Softmax 采样的正确做法:修正规模不平衡与分散性带来的偏差

Two-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion

Walid Bendada · Guillaume Salha-Galvan

中文摘要

Softmax 分布采样是机器学习中的基础操作,但其与条目数成正比的线性时间代价使大规模精确采样难以实用。两级 Softmax(2LS)采样是一种支持亚线性时间采样的常用替代方案。它先将条目划分为若干簇,再依次采样一个簇和该簇内的一个条目。本研究表明,尽管具有优势,2LS 会引入系统性的不良采样偏差,根源在于对簇的加权忽略了簇间规模不平衡与簇内相似性分散性。提出两种采样方法:规模修正 2LS(S-2LS)和规模与分散性修正 2LS(SD-2LS),可修正上述偏差,在可忽略乃至零计算开销下提供理论上更优的 softmax 近似。在五个大规模数据集上的深入实验验证了所提方法在采样性质上的改进。建议在后续工作中以这两种方法替代标准 2LS。

关键要点

  1. 01两级 softmax(2LS)采样因忽略簇规模不平衡和簇内相似性分散性,会引入系统性采样偏差。
  2. 02提出 S-2LS 与 SD-2LS 两种修正方法,可在可忽略到零额外开销下获得理论更优的 softmax 近似。
  3. 03在五个大规模数据集上验证了所提方法相对标准 2LS 在采样性质上的实际改进。
  4. 04建议在未来工作中以 S-2LS 和 SD-2LS 一致替代标准两级 softmax 采样。

解读

尚无解读。

原始英文摘要

arXiv:2610.10483v1 Announce Type: new Abstract: Sampling from a softmax distribution is a fundamental operation in machine learning, but its linear complexity in the number of items makes exact sampling impractical at scale. Two-level softmax (2LS) sampling is a popular alternative enabling sublinear-time sampling. Assuming items are partitioned into clusters, 2LS first samples a cluster and then an item within it. In this paper, we show that, despite its advantages, 2LS introduces systematic and undesirable sampling biases, which arise from misweighting clusters by ignoring both cluster size imbalance and intra-cluster similarity dispersion. We propose two sampling methods, Size-Corrected 2LS (S-2LS) and Size- and Dispersion-Corrected 2LS (SD-2LS), which correct these biases and provide provably better softmax approximations with negligible to non-existent computational overhead. In-depth experiments on five large-scale datasets validate the improved sampling properties of our methods. We recommend their consistent use in place of standard 2LS in future work.

同方向论文 · cs.LG

查看全部 →