Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

STAPI: An Automatic Scraper for Extracting Iterative Title-Text Structure from Web Documents

Nan Zhang, Shomir Wilson, Prasenjit Mitra

Publikation: Beitrag in Buch/Bericht/Sammelwerk/KonferenzbandAufsatz in KonferenzbandForschungPeer-Review

Abstract

Formal documents often are organized into sections of text, each with a title, and extracting this structure remains an under-explored aspect of natural language processing. This iterative title-text structure is valuable data for building models for headline generation and section title generation, but there is no corpus that contains web documents annotated with titles and prose texts. Therefore, we propose the first title-text dataset on web documents that incorporates a wide variety of domains to facilitate downstream training. We also introduce STAPI (Section Title And Prose text Identifier), a two-step system for labeling section titles and prose text in HTML documents. To filter out unrelated content like document footers, its first step involves a filter that reads HTML documents and proposes a set of textual candidates. In the second step, a typographic classifier takes the candidates from the filter and categorizes each one into one of the three pre-defined classes (title, prose text, and miscellany). We show that STAPI significantly outperforms two baseline models in terms of title-text identification. We release our dataset along with a web application to facilitate supervised and semi-supervised training in this domain.

OriginalspracheEnglisch
Titel des Sammelwerks2022 Language Resources and Evaluation Conference, LREC 2022
Herausgeber/-innenNicoletta Calzolari, Frederic Bechet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Helene Mazo, Jan Odijk, Stelios Piperidis
Herausgeber (Verlag)European Language Resources Association (ELRA)
Seiten3461-3470
Seitenumfang10
ISBN (elektronisch)9791095546726
PublikationsstatusVeröffentlicht - 2022
Veranstaltung13th International Conference on Language Resources and Evaluation Conference, LREC 2022 - Marseille, Frankreich
Dauer: 20 Juni 202225 Juni 2022

Konferenz

Konferenz13th International Conference on Language Resources and Evaluation Conference, LREC 2022
Land/GebietFrankreich
OrtMarseille
Zeitraum20 Juni 202225 Juni 2022

ASJC Scopus Sachgebiete

  • Sprache und Linguistik
  • Bibliotheks- und Informationswissenschaften
  • Linguistik und Sprache
  • Ausbildung bzw. Denomination

Dieses zitieren