BERT-LSH: Reducing Absolute Compute For Attention

2024-04-12Code Available0· sign in to hype

ZeZheng Li, Kingston Yip

Code Available — Be the first to reproduce this paper.

Code

github.com/leo4life2/algoml-final
OfficialIn paperjax★ 3

Abstract

This study introduces a novel BERT-LSH model that incorporates Locality Sensitive Hashing (LSH) to approximate the attention mechanism in the BERT architecture. We examine the computational efficiency and performance of this model compared to a standard baseline BERT model. Our findings reveal that BERT-LSH significantly reduces computational demand for the self-attention layer while unexpectedly outperforming the baseline model in pretraining and fine-tuning tasks. These results suggest that the LSH-based attention mechanism not only offers computational advantages but also may enhance the model's ability to generalize from its training data. For more information, visit our GitHub repository: https://github.com/leo4life2/algoml-final

Tasks

Computational Efficiency

BERT-LSH: Reducing Absolute Compute For Attention

Code

Abstract

Tasks

Reproductions