Pre-Training Transformers as Energy-Based Cloze Models
Abstract: We introduce Electric, an energy-based cloze model for representation learning over text. Like BERT, it is a conditional generative model of tokens given their contexts. However, Electric does not use masking or output a full distribution over tokens that could occur in a context. Instead, it assigns a scalar energy score to each input token indicating how likely it is given its context. We train Electric using an algorithm based on noise-contrastive estimation and elucidate how this learning objective is closely related to the recently proposed ELECTRA pre-training method. Electric performs well when transferred to downstream tasks and is particularly effective at producing likelihood scores for text: it re-ranks speech recognition n-best lists better than LLMs and much faster than masked LLMs. Furthermore, it offers a clearer and more principled view of what ELECTRA learns during pre-training.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.