One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs 文章

ArXiv CS.AI2026-05-27NEWSen作者: Di He, Songjun Tu, Keyu Wang, Lu Yin, Shiwei Liu

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs · 相关人物

暂无数据