Published January 1, 2022 | Version v1
Journal article Open

Learning interpretable word embeddings via bidirectional alignment of dimensions with semantic concepts

  • 1. Ludwig Maximilians Univ Munchen, Ctr Informat & Language Proc CIS, Munich, Germany
  • 2. ASELSAN Res Ctr, Ankara, Turkey

Description

We proposebidirectional impartingorBiImp, a generalized method for aligning embeddingdimensions with concepts during the embedding learning phase. While preserving the semanticstructure of the embedding space, BiImp makes dimensions interpretable, which has a criticalrole in deciphering the black-box behavior of word embeddings. BiImp separately utilizes bothdirections of a vector space dimension: each direction can be assigned to a different concept.This increases the number of concepts that can be represented in the embedding space. Ourexperimental results demonstrate the interpretability of BiImp embeddings without makingcompromises on the semantic task performance. We also use BiImp to reduce gender biasin word embeddings by encoding gender-opposite concepts (e.g., male-female) in a singleembedding dimension. These results highlight the potential of BiImp in reducing biases andstereotypes present in word embeddings. Furthermore, task or domain-specific interpretableword embeddings can be obtained by adjusting the corresponding word groups in embeddingdimensions according to task or domain. As a result, BiImp offers wide liberty in studying wordembeddings without any further effort

Files

bib-083b9afc-2ee1-48ec-8ff4-9e945b802803.txt

Files (228 Bytes)

Name Size Download all
md5:10254c4eed597f569383cab522a3f330
228 Bytes Preview Download