Loading Open Internet
    Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention [P]