Dergi makalesi Açık Erişim

Grounded Sequence to Sequence Transduction

Specia, Lucia; Barrault, Loic; Caglayan, Ozan; Duarte, Amanda; Elliott, Desmond; Gella, Spandana; Holzenberger, Nils; Lala, Chiraag; Lee, Sun Jae; Libovicky, Jindrich; Madhyastha, Pranava; Metze, Florian; Mulligan, Karl; Ostapenko, Alissa; Palaskar, Shruti; Sanabria, Ramon; Wang, Josiah; Arora, Raman


MARC21 XML

<?xml version='1.0' encoding='UTF-8'?>
<record xmlns="http://www.loc.gov/MARC21/slim">
  <leader>00000nam##2200000uu#4500</leader>
  <datafield tag="909" ind1="C" ind2="O">
    <subfield code="p">user-tubitak-destekli-proje-yayinlari</subfield>
    <subfield code="o">oai:zenodo.org:5905</subfield>
  </datafield>
  <datafield tag="520" ind1=" " ind2=" ">
    <subfield code="a">Speech recognition and machine translation have made major progress over the past decades, providing practical systems to map one language sequence to another. Although multiple modalities such as sound and video are becoming increasingly available, the state-of-the-art systems are inherently unimodal, in the sense that they take a single modality - either speech or text - as input. Evidence from human learning suggests that additional modalities can provide disambiguating signals crucial for many language tasks. In this article, we describe the How2 dataset , a large, open-domain collection of videos with transcriptions and their translations. We then show how this single dataset can be used to develop systems for a variety of language tasks and present a number of models meant as starting points. Across tasks, we find that building multimodal architectures that perform better than their unimodal counterpart remains a challenge. This leaves plenty of room for the exploration of more advanced solutions that fully exploit the multimodal nature of the How2 dataset , and the general direction of multimodal learning with other datasets as well.</subfield>
  </datafield>
  <datafield tag="980" ind1=" " ind2=" ">
    <subfield code="a">publication</subfield>
    <subfield code="b">article</subfield>
  </datafield>
  <datafield tag="540" ind1=" " ind2=" ">
    <subfield code="a">Creative Commons Attribution</subfield>
    <subfield code="u">http://www.opendefinition.org/licenses/cc-by</subfield>
  </datafield>
  <datafield tag="100" ind1=" " ind2=" ">
    <subfield code="a">Specia, Lucia</subfield>
  </datafield>
  <datafield tag="856" ind1="4" ind2=" ">
    <subfield code="z">md5:5d1edae43650c9319c0522fa660ab3b2</subfield>
    <subfield code="s">354</subfield>
    <subfield code="u">https://aperta.ulakbim.gov.trrecord/5905/files/bib-53c6b0fd-f295-4c38-bf91-fea0d6afc31a.txt</subfield>
  </datafield>
  <controlfield tag="005">20210315061805.0</controlfield>
  <datafield tag="260" ind1=" " ind2=" ">
    <subfield code="c">2020-01-01</subfield>
  </datafield>
  <datafield tag="024" ind1=" " ind2=" ">
    <subfield code="a">10.1109/JSTSP.2020.2998415</subfield>
    <subfield code="2">doi</subfield>
  </datafield>
  <datafield tag="542" ind1=" " ind2=" ">
    <subfield code="l">open</subfield>
  </datafield>
  <datafield tag="245" ind1=" " ind2=" ">
    <subfield code="a">Grounded Sequence to Sequence Transduction</subfield>
  </datafield>
  <datafield tag="909" ind1="C" ind2="4">
    <subfield code="v">14</subfield>
    <subfield code="p">IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING</subfield>
    <subfield code="c">577-591</subfield>
    <subfield code="n">3</subfield>
  </datafield>
  <datafield tag="650" ind1="1" ind2="7">
    <subfield code="a">cc-by</subfield>
    <subfield code="2">opendefinition.org</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Barrault, Loic</subfield>
    <subfield code="u">Univ Sheffield, Dept Comp Sci, Sheffield S10 2TG, S Yorkshire, England</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Caglayan, Ozan</subfield>
    <subfield code="u">Imperial Coll London, Dept Comp, London SW7 2BU, England</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Duarte, Amanda</subfield>
    <subfield code="u">Univ Politcn Catalunya, Dept Signal Theory &amp; Commun, Barcelona 08034, Spain</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Elliott, Desmond</subfield>
    <subfield code="u">Univ Copenhagen, Dept Comp Sci, DK-1165 Copenhagen, Denmark</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Gella, Spandana</subfield>
    <subfield code="u">Univ Edinburgh, Inst Language Cognit &amp; Computat, Edinburgh EH8 9YL, Midlothian, Scotland</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Holzenberger, Nils</subfield>
    <subfield code="u">Johns Hopkins Univ, Ctr Language &amp; Speech Proc, Baltimore, MD 21218 USA</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Lala, Chiraag</subfield>
    <subfield code="u">Univ Sheffield, Dept Comp Sci, Sheffield S10 2TG, S Yorkshire, England</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Lee, Sun Jae</subfield>
    <subfield code="u">Univ Penn, Philadelphia, PA 19104 USA</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Libovicky, Jindrich</subfield>
    <subfield code="u">Ludwig Maximilians Univ Munchen, Ctr Informat &amp; Language Proc, D-80333 Munich, Germany</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Madhyastha, Pranava</subfield>
    <subfield code="u">Imperial Coll London, Dept Comp, London SW7 2BU, England</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Metze, Florian</subfield>
    <subfield code="u">Carnegie Mellon Univ, Language Technol Inst, Pittsburgh, PA 15213 USA</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Mulligan, Karl</subfield>
    <subfield code="u">Johns Hopkins Univ, Baltimore, MD 21218 USA</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Ostapenko, Alissa</subfield>
    <subfield code="u">Worcester Polytech Inst, Worcester, MA 01609 USA</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Palaskar, Shruti</subfield>
    <subfield code="u">Carnegie Mellon Univ, Language Technol Inst, Pittsburgh, PA 15213 USA</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Sanabria, Ramon</subfield>
    <subfield code="u">Carnegie Mellon Univ, Language Technol Inst, Pittsburgh, PA 15213 USA</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Wang, Josiah</subfield>
    <subfield code="u">Imperial Coll London, Dept Comp, London SW7 2BU, England</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="a">Arora, Raman</subfield>
    <subfield code="u">Johns Hopkins Univ, Ctr Language &amp; Speech Proc, Baltimore, MD 21218 USA</subfield>
  </datafield>
  <controlfield tag="001">5905</controlfield>
  <datafield tag="980" ind1=" " ind2=" ">
    <subfield code="a">user-tubitak-destekli-proje-yayinlari</subfield>
  </datafield>
</record>
74
8
görüntülenme
indirilme
Görüntülenme 74
İndirme 8
Veri hacmi 2.8 kB
Tekil görüntülenme 70
Tekil indirme 8

Alıntı yap