Hey Mantas!
I was looking at your solution for the assignment 2. Specifically question 1, b) 1) -> When is the gradient zero?
Your answers look good, since they indeed could get a gradient zero. Nonetheless, I believe this won't occur in every case, as the update depends on all the outside words. I think a more accurate answer could be the following three scenarios:
a) Trivial one: All outside word embeddings are zero vectors.
b) Error vector (y^-y) = 0, that is when the predicted conditional probabilities match the true distribution of outside words.
c) Rare one: The error vector, although non-zero, is orthogonal to the subspace spanned by all outside vectors.
In all three scenarios no update will ever occur, as the gradient will always be the zero vector. Happy to continue the discussion!
Alex
Hey Mantas!
I was looking at your solution for the assignment 2. Specifically question 1, b) 1) -> When is the gradient zero?
Your answers look good, since they indeed could get a gradient zero. Nonetheless, I believe this won't occur in every case, as the update depends on all the outside words. I think a more accurate answer could be the following three scenarios:
a) Trivial one: All outside word embeddings are zero vectors.
b) Error vector (y^-y) = 0, that is when the predicted conditional probabilities match the true distribution of outside words.
c) Rare one: The error vector, although non-zero, is orthogonal to the subspace spanned by all outside vectors.
In all three scenarios no update will ever occur, as the gradient will always be the zero vector. Happy to continue the discussion!
Alex