Skip to content

How can I train the parameter in cross-attention? #3

Description

@dankimh

Hello, I tried to fix a dimension issue and train the model, but when I check the gradients of the parameters in the cross-attention layer, there are no gradients for the parameters. Is this expected behavior? Or did I miss something? It seems weird since the paper states that the parameters in cross-attention are learnable.

Activity

  1. jorgelerre commented on Jun 15, 2026

    @jorgelerre

    Hello, I found this same issue as dankimh. According to the paper, prototype learning is performed using cross-attention.
    However, the implementation of CrossAttention internally relies on MaskAttention, which does not include the projection matrices (W_Q), (W_K), and (W_V).
    Therefore, the prototype update module (p2c) contains no learnable parameters.
    As a consequence, the operation cluster_emb = self.p2c(cluster_emb_, x_emb_, x_emb_, mask=mask.transpose(0,1)) effectively behaves as a masked weighted average of channel embeddings.

    Thereby, I would like to make the following questions:

    • Was the intention to use MaskAttentionLayer instead (which actually has learnable weights), or is the current behavior the one used in the experiments reported in the paper?
    • In case of using MaskAttentionLayer, how did you train the resulting network so the gradients would be successfully propagated to its parameters, taking into account that the prototypes are replaced at each step?
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions