Abstract
Formal documents often are organized into sections of text, each with a title, and extracting this structure remains an under-explored aspect of natural language processing. This iterative title-text structure is valuable data for building models for headline generation and section title generation, but there is no corpus that contains web documents annotated with titles and prose texts. Therefore, we propose the first title-text dataset on web documents that incorporates a wide variety of domains to facilitate downstream training. We also introduce STAPI (Section Title And Prose text Identifier), a two-step system for labeling section titles and prose text in HTML documents. To filter out unrelated content like document footers, its first step involves a filter that reads HTML documents and proposes a set of textual candidates. In the second step, a typographic classifier takes the candidates from the filter and categorizes each one into one of the three pre-defined classes (title, prose text, and miscellany). We show that STAPI significantly outperforms two baseline models in terms of title-text identification. We release our dataset along with a web application to facilitate supervised and semi-supervised training in this domain.
| Originalsprache | Englisch |
|---|---|
| Titel des Sammelwerks | 2022 Language Resources and Evaluation Conference, LREC 2022 |
| Herausgeber/-innen | Nicoletta Calzolari, Frederic Bechet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Helene Mazo, Jan Odijk, Stelios Piperidis |
| Herausgeber (Verlag) | European Language Resources Association (ELRA) |
| Seiten | 3461-3470 |
| Seitenumfang | 10 |
| ISBN (elektronisch) | 9791095546726 |
| Publikationsstatus | Veröffentlicht - 2022 |
| Veranstaltung | 13th International Conference on Language Resources and Evaluation Conference, LREC 2022 - Marseille, Frankreich Dauer: 20 Juni 2022 → 25 Juni 2022 |
Konferenz
| Konferenz | 13th International Conference on Language Resources and Evaluation Conference, LREC 2022 |
|---|---|
| Land/Gebiet | Frankreich |
| Ort | Marseille |
| Zeitraum | 20 Juni 2022 → 25 Juni 2022 |
ASJC Scopus Sachgebiete
- Sprache und Linguistik
- Bibliotheks- und Informationswissenschaften
- Linguistik und Sprache
- Ausbildung bzw. Denomination
Dieses zitieren
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver