Replies: 3 comments 7 replies
|
Hi @Coluding |
1 reply
|
Hi again, is your implementation of LocalAttention also sparse-aware such that empty tokens are not represented in memory? I could not find any docs on that. Thanks in advance! |
6 replies
|
Great! Thanks for your quick replies! |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi all :)
I was wondering if dilated local attention like in the Longformer paper is already integrated? I could not find it but thought maybe of the Longformer paper reference in the docstring of the LocalAttention class that there may be something already implemented. If not, is the implementation planned?
Thanks in advance!
All reactions