Datasets: A Community Library for Contemporary NLP Datasets
Datasets: A Community Library for Natural Language Processing
Datasets is a community library for contemporary non-linear statistical methods (nlp).
The library aims to standardize end-user interfaces, versioning, and documentation, while providing a lightweight front-end that behaves similarlyfor small datasets as for internet-scale corpora.
The design of the libraryincorporates a distributed, community-driven approach to adding datasets and documenting usage.
The library now includes morethan 650 unique datasets, has more than 250 contributors, and has helped support a variety of novel cross-dataset research projects and shared tasks.
Authors
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major