Croatian n-grams

Please use the following text to cite this item or export to a predefined format:
Tadić, Marko, 2013, Croatian n-grams, HR-CLARIN, http://hdl.handle.net/20.500.14615/2-31
Date issued
2013
Size
8681475 entries
Language(s)
Description
This resource contains sets of n-grams of different sizes (from 1 to 3) computed from the Croatian National Corpus v2.5. N-grams were computed both from lowercased text and text in original character case. For every size of n above one (i.e. for bigrams and trigrams), n-grams were computed in two ways: taking to account only those appearing within sentence and across sentence boundaries. Regarding the tokenization of the corpus, token is considered to be a continuous sequence of non-whitespace characters. Punctuation markings are treated as separate tokens. Complex punctuations are tokenized as a sequence of simple punctuations. Resource consists of 10 textual files, each computed with different combination of paramaters (i.e. n-gram length, character case, sentence boundaries). Each line in the file represents one unique n-gram and its absolute frequency in the corpus, separated by a tabulator. N-grams are ordered according to their frequency, starting from highest to lowest. The n-grams lists were produced using methodology and tools developed by the CESAR Polish partner IPIPAN.
Acknowledgement
This item isPublicly Available
and licensed under:
 Files in this item
Name
archive.zip
Size
67.8 MB
Format
application/zip
Description
zip
MD5
d761a4acdd4e240abd5f6afc929834e1
Preview
  File Preview