Datasets#
dacy.datasets.dane#
This includes the DaNE dataset wrapped and read in as a SpaCy corpus.
- dacy.datasets.dane.dane(save_path=None, splits=['train', 'dev', 'test'], redownload=False, n_sents=1, open_unverified_connection=False, **kwargs)[source]#
Reads the DaNE dataset as a spacy Corpus.
- Parameters
save_path (str, optional) – Path to the DaNE dataset If it does not contain the dataset it is downloaded to the folder. Defaults to None corresponding to dacy.where_is_my_dacy() in the datasets subfolder.
splits (list[str], optional) – Which splits of the dataset should be returned. Possible options include “train”, “dev”, “test”, “all”. Defaults to [“train”, “dev”, “test”].
redownload (bool, optional) – Should the dataset be redownloaded. Defaults to False.
n_sents (int, optional) – Number of sentences per document. Only applied if the dataset is downloaded. Defaults to 1.
open_unverified_connection (bool, optional) – Should you download from an unverified connection. Defaults to False.
force_extension (bool, optional) – Set the extension to the doc regardless of whether it already exists. Defaults to False.
- Returns
Returns a SpaCy corpus or a list thereof.
- Return type
Union[list[Corpus], Corpus]
Example
>>> from dacy.datasets import dane >>> train, dev, test = dane(splits=["train", "dev", "test"])