DCAI
← 返回全部动态
arXiv 机器学习规则精选09月24日 12:00

Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders

arXiv:2609.27248v1 Announce Type: new Abstract: Next-token prediction has enabled highly fluent autoregressive language models, but it represents global structure only indirectly through sequential factorization. In contrast, high-fidelity autoencoders have become a standard primitive in image generation, enabling generative models to operate over continuous latent spaces; text lacks a comparably faithful continuous representation. We propose LLMAE, a method for repurposing a pretrained decoder-only language model as a continuous text autoencoder by exposing an intermediate fixed-length latent bottleneck within its internal activations. Insta

阅读 arXiv 机器学习 原文 ↗