Datasets#

dacy.datasets.dane#

This includes the DaNE dataset wrapped and read in as a SpaCy corpus.

dacy.datasets.dane.dane(save_path=None, splits=['train', 'dev', 'test'], redownload=False, n_sents=1, open_unverified_connection=False, **kwargs)[source]#

Reads the DaNE dataset as a spacy Corpus.

Parameters
  • save_path (str, optional) – Path to the DaNE dataset If it does not contain the dataset it is downloaded to the folder. Defaults to None corresponding to dacy.where_is_my_dacy() in the datasets subfolder.

  • splits (list[str], optional) – Which splits of the dataset should be returned. Possible options include “train”, “dev”, “test”, “all”. Defaults to [“train”, “dev”, “test”].

  • redownload (bool, optional) – Should the dataset be redownloaded. Defaults to False.

  • n_sents (int, optional) – Number of sentences per document. Only applied if the dataset is downloaded. Defaults to 1.

  • open_unverified_connection (bool, optional) – Should you download from an unverified connection. Defaults to False.

  • force_extension (bool, optional) – Set the extension to the doc regardless of whether it already exists. Defaults to False.

Returns

Returns a SpaCy corpus or a list thereof.

Return type

Union[list[Corpus], Corpus]

Example

>>> from dacy.datasets import dane
>>> train, dev, test = dane(splits=["train", "dev", "test"])