itWaC: Italian web corpus

The Italian web corpus (itWac) is a language corpus made up of texts collected from the Internet. The corpus consists of 1.5 billion words and was prepared by Marco Baroni. Texts are part-of-speech tagged and lemmatized with the TreeTagger tool. Moreover, users can explore the grammatical and collocational behavior of Italian words as a result of a word sketch grammar prepared Marco Baroni and later updated by Valentina Efrati and Francesca Masini (TRIPLE lab, Roma Tre University).

See the Italian part-of-speech tagset describing POS tags used in the corpus.

The corpus is cleaned and deduplicated. More information about this corpus building can be found in Marco Baroni & Adam Kilgarriff (2006).

Search the itWaC corpus

Sketch Engine offers a range of tools to work with this Italian corpus.

or

A complete set of Sketch Engine tools is available to work with this Italian itWaC corpus to generate:

  • word sketch – Italian collocations categorized by grammatical relations
  • thesaurus – synonyms and similar words for every word
  • word lists – lists of Italian nouns, verbs, adjectives etc. organized by frequency
  • n-grams – frequency list of multi-word units
  • concordance – examples in context

Bibliography

BARONI, Marco; KILGARRIFF, Adam. Large linguistically-processed web corpora for multiple languages. In: Proceedings of the Eleventh Conference of the European Chapter of the Association for Computational Linguistics: Posters & Demonstrations. Association for Computational Linguistics, 2006, pp. 87–90.

Other text corpora in Sketch Engine

Sketch Engine offers 350+ language corpora.

Use Sketch Engine in minutes

Generating collocations, frequency lists, examples in contexts, n-grams or extracting terms is easy with Sketch Engine. Use our Quick Start Guide to learn it in minutes.