MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering 文章

ArXiv CS.AI2026-08-10PAPERen作者: Jakub Po\'cwiardowski, Mateusz Modrzejewski

详细信息

来源站点
ArXiv CS.AI
作者
Jakub Po\'cwiardowski, Mateusz Modrzejewski
文章类型
PAPER
语言
en
发布日期
2026-08-10

摘要

arXiv:2608.06638v1 Announce Type: cross Abstract: Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored. We analyze two public text-to-MIDI systems of contrasting design: the purpose-built encoder--decoder text2midi and MIDI-LLM, a Llama~3.2~1B model extended with MIDI tokens using linear probing, the logit and tuned lenses, activation patching and difference-in-means steering. Across these methods, we recover musically meaningful structure and show how architecture shapes its formation and control. Pitch, instrumentation, harmony and texture are linearly decodable in both models. text2midi refines predictions gradually across depth, whereas MIDI-LLM works largely in its inherited textual basis before a sharp late rotation into the musical vocabulary; patching identifies a matching late attenuation of prompt-driven instrument transfer.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据