Skip to main navigation Skip to search Skip to main content

On the Impact of Dataset Size: A Twitter Classification Case Study

  • Thi Huyen Nguyen
  • , Hoang H. Nguyen
  • , Zahra Ahmadi
  • , Tuan Anh Hoang
  • , Thanh Nam Doan

Research output: Chapter in book/report/conference proceedingConference contributionResearchpeer review

Abstract

The recent advent and evolution of deep learning models and pre-trained embedding techniques have created a breakthrough in supervised learning. Typically, we expect that adding more labeled data improves the predictive performance of supervised models. On the other hand, collecting more labeled data is not an easy task due to several difficulties, such as manual labor costs, data privacy, and computational constraint. Hence, a comprehensive study on the relation between training set size and the classification performance of different methods could be essentially useful in the selection of a learning model for a specific task. However, the literature lacks such a thorough and systematic study. In this paper, we concentrate on this relationship in the context of short, noisy texts from Twitter. We design a systematic mechanism to comprehensively observe the performance improvement of supervised learning models with the increase of data sizes on three well-known Twitter tasks: sentiment analysis, informativeness detection, and information relevance. Besides, we study how significantly better the recent deep learning models are compared to traditional machine learning approaches in the case of various data sizes. Our extensive experiments show (a) recent pre-trained models have overcome big data requirements, (b) a good choice of text representation has more impact than adding more data, and (c) adding more data is not always beneficial in supervised learning.

Original languageEnglish
Title of host publicationWI-IAT '21
Subtitle of host publicationIEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology
PublisherAssociation for Computing Machinery (ACM)
Pages210-217
Number of pages8
ISBN (Electronic)9781450391153
DOIs
Publication statusPublished - 13 Apr 2022
Event2021 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology, WI-IAT 2021 - Virtual, Online, Australia
Duration: 14 Dec 202117 Dec 2021

Publication series

NameACM International Conference Proceeding Series

Conference

Conference2021 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology, WI-IAT 2021
Country/TerritoryAustralia
CityVirtual, Online
Period14 Dec 202117 Dec 2021

Keywords

  • dataset size
  • empirical study
  • extrapolation methods
  • machine learning
  • neural network
  • Twitter classification

ASJC Scopus subject areas

  • Software
  • Human-Computer Interaction
  • Computer Vision and Pattern Recognition
  • Computer Networks and Communications

Cite this