LogIn / Registration

Articles for sale

Head office address: Moscow, 123317, Moscow City, 8th Floor, Presnenskaya Embankment, 6, bldg. 2
article@123mi.ru

#3834. LanguageCrawl: a generic tool for building language models upon common Crawl

November 2026	publication date
Proposal available till	10-07-2025
4 total number of authors per manuscript	0 $

The title of the journal is available only for the authors who have already paid for

Journal’s subject area:

Language and Linguistics;
Linguistics and Language;
Education;
Library and Information Sciences;

Places in the authors’ list:

place 1	place 2	place 3	place 4
Free	Free	Free	Free
2350 $	1200 $	1050 $	900 $
Contract №3834.1	Contract №3834.2	Contract №3834.3	Contract №3834.4

1 place - free (for sale)
2 place - free (for sale)
3 place - free (for sale)
4 place - free (for sale)

Abstract:
The exponential growth of the internet community has resulted in the production of a vast amount of unstructured data, including web pages, blogs and social media. Such a volume consisting of hundreds of billions of words is unlikely to be analyzed by humans. In this work we introduce the tool LanguageCrawl, which allows Natural Language Processing (NLP) researchers to easily build web-scale corpora using the Common Crawl Archive—an open repository of web crawl information, which contains petabytes of data. We present three use cases in the course of this work: filtering of Polish websites, the construction of n-gram corpora and the training of a continuous skipgram language model with hierarchical softmax. Our tool utilizes effective libraries and design. We strongly believe that our work will facilitate further NLP research, especially in under-resourced languages, in which the lack of appropriately-sized corpora is a serious hindrance to applying data-intensive methods, such as deep neural networks.
Keywords:
Common Crawl; Language Models; N-gram; Polish Web Corpus; Word2Vec

Contacts :

Contact Info

Office

Sign up for a meeting through a call center: help@buy-sell-article.com
,