Please use the following text to cite this item or export to a predefined format:
dc.contributor.author | Tadić, Marko |
dc.date.accessioned | 2025-09-17T09:40:17Z |
dc.date.available | 2025-09-17T09:40:17Z |
dc.date.issued | 2013 |
dc.description | This resource contains sets of n-grams of different sizes (from 1 to 3) computed from the Croatian National Corpus v2.5. N-grams were computed both from lowercased text and text in original character case. For every size of n above one (i.e. for bigrams and trigrams), n-grams were computed in two ways: taking to account only those appearing within sentence and across sentence boundaries. Regarding the tokenization of the corpus, token is considered to be a continuous sequence of non-whitespace characters. Punctuation markings are treated as separate tokens. Complex punctuations are tokenized as a sequence of simple punctuations. Resource consists of 10 textual files, each computed with different combination of paramaters (i.e. n-gram length, character case, sentence boundaries). Each line in the file represents one unique n-gram and its absolute frequency in the corpus, separated by a tabulator. N-grams are ordered according to their frequency, starting from highest to lowest. The n-grams lists were produced using methodology and tools developed by the CESAR Polish partner IPIPAN. |
dc.identifier.uri | http://hdl.handle.net/20.500.14615/2-31 |
dc.language.iso | hrv |
dc.publisher | University of Zagreb, Faculty of Humanities and Social Sciences, Department of Information Sciences |
dc.rights | The MIT Licence |
dc.rights.label | PUB |
dc.rights.uri | https://zzl-ffzg.mit-license.org/ |
dc.subject | n-grams |
dc.subject | Croatian language |
dc.subject | Croatian National Corpus |
dc.title | Croatian n-grams |
dc.type | lexicalConceptualResource |
local.contact.person | Marko Tadić marko.tadic@ffzg.hr Faculty of Humanities and Social Sciences, University of Zagreb |
local.files.count | 1 |
local.files.size | 71096144 |
local.has.files | yes |
local.language.name | Croatian |
local.size.info | 8681475 entries |
local.sponsor | euFunds CIP-ICT-PSP-2009-4: 271022 European Commission Central and South-East European Resources |
metashare.ResourceInfo#ContentInfo.detailedType | computationalLexicon |
metashare.ResourceInfo#ContentInfo.mediaType | text |